Skip to content

Cloud & Infrastructure · AI Infrastructure

AI Data Centers Are Learning to Throttle Themselves for the Grid

Emerald AI raised $150M to make AI data centers shed power on demand. A 96-GPU Nvidia trial cut draw 30% in 30 seconds. What that means for capex plans.

Anurag Verma

Anurag Verma

4 min read

AI Data Centers Are Learning to Throttle Themselves for the Grid

Sponsored

Share

A 96-GPU Nvidia Blackwell Ultra cluster outside London cut its power draw by 30% in 30 seconds during a live trial in 2026, then kept it there for up to 10 hours when asked, without dropping the workloads that couldn’t tolerate a slowdown. The software that did it, Emerald AI’s Conductor platform, just closed a $150 million Series A at a $1.05 billion valuation. The bet behind that valuation is specific: the thing actually limiting how fast new AI capacity comes online isn’t chip supply anymore, it’s whether the grid will let you plug the data center in, and a facility that can prove it will throttle itself gets approved faster than one that won’t.

What Conductor does

Emerald AI, founded in November 2024 by former Biden administration energy official Varun Sivaram, built Conductor to sit between a data center’s AI workloads and the power grid feeding it. It tags jobs by how much delay or slowdown they can tolerate, then acts on that tagging when the grid signals stress: shifting latency-tolerant jobs to run later, moving workloads to a different site entirely, throttling GPU clock frequencies to trim draw without stopping a job outright, or drawing down on-site batteries and backup generators. Latency-critical work stays protected throughout; the flexibility comes from everything else in the facility’s workload mix that doesn’t need to run at full tilt every second.

The joint trial with Nvidia and cloud provider Nebius put numbers on that pitch. Running on a 96-GPU Blackwell Ultra cluster, Conductor met every requested power-reduction target across more than 200 simulated grid events, cut demand by up to 40% in under a minute, and hit a 30% load shed within 30 seconds for the kind of emergency curtailment a grid operator would request during a real stress event.

Why a grid operator cares about any of this

The framing that makes this more than an efficiency story is the interconnection queue. Data centers requesting new grid connections are typically evaluated as though they’ll draw peak power continuously, because that’s the worst case a utility has to plan for. A facility that can credibly commit to cutting draw on request during genuine grid stress changes that worst case, and utilities have direct financial and regulatory reasons to prioritize connections that reduce, rather than add to, peak-demand risk.

That’s the actual product Emerald AI is selling: not just lower power bills for the data center operator, though that’s a side benefit, but a faster path through an interconnection process that has become the real bottleneck on new AI capacity. Grid interconnection queues in several major markets already stretch years out, a timeline that GPU procurement, however constrained, doesn’t come close to matching.

What this means if you’re planning infrastructure spend

Grid-flexible compute is a data-center-operator and utility-scale story first, not something an individual engineering team configures directly. But it’s a signal worth reading if you’re the one budgeting for AI infrastructure, whether that’s reserved GPU capacity or a longer-term colocation commitment:

ConsiderationWhat’s changing
Grid interconnection timelinesFacilities offering demand flexibility get prioritized; flat peak-draw facilities face longer queues
Data center site selectionProviders building flexibility into new sites may bring capacity online faster than competitors who don’t
Workload schedulingLatency-tolerant batch and training jobs become a genuine cost lever, not just a nice-to-have for cost optimization
Provider selectionA provider’s grid relationship and flexibility posture becomes a legitimate diligence question, not just its GPU generation

None of this changes today’s procurement decision on its own. It’s a reason to ask a prospective data center or cloud provider whether they’re building demand flexibility into their power posture, the same way you’d ask about redundancy or interconnect quality, because the operators solving the grid problem now are the ones more likely to have capacity available when the next wave of buildout hits its own interconnection queue.

The bigger pattern

Emerald AI’s raise lands alongside a broader shift in how the industry talks about the AI infrastructure buildout: less about whether enough chips exist, more about whether enough power, in the right place, at the right time, can actually reach them. Nuclear restarts, on-site generation, and now demand-flexible software are all responses to the same constraint from different angles. A company getting funded specifically to make data centers ask for less power, rather than more, is a reasonable indicator of where the bottleneck has actually moved.

The chip shortage headlines from the last two years told a supply story. This one is a permissions story: get the grid connection approved, and the GPUs you’ve already bought can finally run.

Frequently asked questions

What does Emerald AI's Conductor software actually do?
It sits between a data center's workloads and the power grid, tagging AI jobs by how much delay or slowdown they can tolerate, then dynamically adjusting the facility's power draw when the grid is under stress. Depending on the situation, it can delay latency-tolerant jobs, shift workloads to a different data center location, throttle GPU clock frequencies, or draw on-site batteries and backup generators, all while protecting the jobs that were tagged as latency-critical.
How much power reduction has it actually demonstrated?
In a joint trial with Nvidia and cloud provider Nebius on a 96-GPU Blackwell Ultra cluster near London, the Conductor platform reduced power demand by up to 40% in under a minute, shed roughly 30% of the site's load within 30 seconds for emergency curtailment scenarios, and sustained reduced draw for requests lasting up to 10 hours. It met every requested power-reduction target across more than 200 simulated grid events during the trial.
Why would a data center operator want to limit its own power draw?
Because the actual constraint on building new AI capacity right now is often grid interconnection approval, not chip availability. Utilities and grid operators are more willing to fast-track a connection for a facility that can prove it will flex its draw during peak grid stress than one that demands constant, uninterruptible peak power. A data center that can credibly promise flexibility becomes a better grid citizen and, in practice, gets online faster.
Does this replace the need for more power generation?
No, and Emerald AI doesn't claim it does. It's a way to make existing and planned data center capacity more compatible with a grid that's under near-term stress from AI demand growth, buying time and easing interconnection friction while new generation, including the nuclear and renewable buildouts several utilities are pursuing, comes online over a longer timeline. It's a scheduling and coordination layer, not a substitute for more electricity.

Sources

Sponsored

Sponsored

Discussion

Join the conversation.

Comments are powered by GitHub Discussions. Sign in with your GitHub account to leave a comment.

Sponsored