
The grid isn't failing because there's not enough generation capacity in the pipeline. There are 2,600 gigawatts of proposed projects queued for interconnection across the United States alone. The failure is that the tools watching the grid — and processing the queue — stop at the wrong level.
On April 28, 2025, Spain and Portugal lost power for up to 10 hours when 15 GW disconnected in five seconds. The ENTSO-E final report, published in March 2026, rejected the renewable-instability narrative that dominated early headlines. The actual failure chain: sub-synchronous oscillations appeared across the Spanish grid from 9 a.m. onward; transmission operators monitoring 400 kV buses saw voltages at 418 kV — within nominal limits; meanwhile, 220 kV collector substations were hitting 242 kV. No one was watching the collector level in real time. Transformer tap-changers couldn't compensate fast enough. One major generation facility injected reactive power into an already-overvoltage grid rather than absorbing it, triggering a positive feedback loop. The CEOE, Spain's principal employers' organization, estimates the economic cost at €1.6 billion.
Six months earlier, PJM completed its 2027/2028 Base Residual Auction and fell 6,625 MW short of the regional reliability target — the first time in history the entire PJM region, including Fixed Resource Requirement entities, failed to meet the 20% installed reserve margin. The capacity clearing price hit $333.44/MW-day for the third consecutive auction; the total auction bill was $16.4 billion. NRDC and CUB analysis projects $163 billion in cumulative capacity costs between 2028 and 2033. ERCOT's interconnection queue reached 233 GW by Q4 2025 — 77% attributable to data center demand — while only 23 GW of new generation synchronized during the year.
These two events look unrelated. One is a cascading physical failure triggered by reactive power dynamics. The other is a capacity-auction shortfall driven by data center load that outpaced interconnection throughput. But they share a root cause: both are consequences of a monitoring architecture that stops at the wrong level. The Iberian blackout stopped at 400 kV when the failure was happening at 220 kV. The US interconnection crisis stops at the transmission-level study process when the throughput constraint is the absence of queue-level intelligence.
The Gap the Incumbents Aren't Solving

The incumbent grid software ecosystem is large and genuinely capable. GE Vernova's GridOS platform prevented 112 million customer-minutes of interruption for Alabama Power in 2025. Siemens partnered with NVIDIA to run Gridscale X digital twin simulations at 10,000x the speed of traditional approaches. Utilidata secured $60.3 million in Series C funding for its NVIDIA Jetson-based Karman chip, embedded in smart meters and deployed at Portland General Electric and Duquesne Light, bringing edge computing to distribution monitoring.
These are real capabilities. They share a structural limitation: they were designed for transmission-level visibility and transmission-level planning. The Iberian blackout exposed precisely the gap between what these systems watch and what actually failed — the collector level, below the threshold of transmission monitoring.
LineVision, the dominant US provider of Dynamic Line Rating sensors, has documented a 61% capacity increase on AES's 345 kV transmission corridors and a 20-30% increase on National Grid's Syracuse circuits — at roughly 5 to 7% of the cost of traditional reconductoring. The DLR value proposition is proven. What hasn't been solved is the analytics layer that tells a utility which corridor to deploy it on first, and how to integrate that prioritization into interconnection queue planning. The hardware exists. The corridor intelligence doesn't.
The monitoring architecture isn't wrong — it was designed for a grid where power flowed one direction. The grid operating today is bidirectional, low-inertia, and data-center-driven. The same infrastructure now has a systematic blind spot.
On the queue side, McKinsey was contracted to overhaul ERCOT's interconnection process, filing a Batch Study framework for the February 2026 Public Utility Commission discussion. Process reform is necessary. But it doesn't substitute for engineering throughput: with 233 GW of proposed generation at ERCOT and 2,600 GW nationally, the median time to commercial operation is five years, and only 20% of projects that entered the queue between 2000 and 2018 ever reached it. GridLab analysis found that if just 10% of the 107 GW of renewables in PJM's queue had been built for the 2026/2027 delivery year, consumers would have saved $3.5 billion in a single capacity auction. The cost of queue paralysis isn't abstract.
Where Physics-Grounded AI Fits

Argonne National Laboratory unveiled GridMind in March 2026: a multi-agent LLM system built as a reasoning copilot for control room operators, providing scheduling recommendations and outage simulations as part of the DOE's Genesis Mission. It represents the research frontier — explainable recommendations, natural-language access to complex scheduling data. It's also explicitly research-stage, with no utility deployment timeline and no embedded physics constraints, which means each recommendation requires a separate physical feasibility verification step before an operator can act.
The gap between research maturity and production deployment matters for procurement decisions. Particularly against the current regulatory timeline: the EU AI Act classifies AI systems in critical infrastructure management — power grid dispatch, fault detection, real-time control — as high-risk, with a compliance deadline of August 2, 2026, carrying penalties of €15 million or 3% of global turnover. NERC CIP-003-9, extending cybersecurity management to lower-risk BES Cyber Systems, took effect in April 2026. FERC Order 1920 requires utilities to demonstrate they've assessed Grid Enhancing Technologies — including DLR and power flow control devices — before submitting proposals for traditional transmission construction.
Our work on power grid AI, documented at veriprajna.com/solutions/power-grid-ai, starts from a different premise than either the incumbent platforms or the research-stage copilots. The systems that close the observability gap are narrow, physics-grounded, and auditable: sub-transmission voltage analytics operating at the 220 kV collector level, DLR corridor prioritization models that answer the "which 50 km first" question utilities can't currently answer, and interconnection queue screening tools that apply AI-driven clustering to the study backlog without displacing the engineering judgment that governs ISA execution.
None of these replace the existing SCADA and energy management systems. GE Vernova and Siemens own that installed base, and the right architecture works with it — pulling from existing operational data, feeding recommendations into familiar operator workflows, generating the documentation trails that NERC CIP and FERC compliance audits now require.
Why Queue Intelligence and Stability Analytics Converge

The connection between interconnection queue reform and grid stability isn't obvious from a policy perspective, but it's direct from an engineering one. A transmission planning engineer running N-1 contingency studies — evaluating what happens if a given element trips under a given load scenario — is working with PSS/E powerflow case files that contain the same information that queue analytics need: which transmission constraints are binding across the most proposed interconnection positions simultaneously, and where new generation injection creates the most voltage stress.
FERC Order 2023's cluster study mandate was designed to address the restudy cascade problem: when a project withdraws from the queue, every downstream project that was studied assuming the withdrawn project was present must be restudied. That cascade is why 68% of interconnection studies were completed late in 2022, per FERC's own reporting. Queue intelligence — AI systems that flag which projects are most likely to withdraw, where restudy cascades will concentrate, and which transmission constraints are binding across the most queue positions — reduces the restudy burden at the source rather than trying to process it faster after the fact.
That same constraint mapping, combined with DLR corridor analytics, tells planners where DLR deployment creates the most queue-unlocking value — which is precisely the GET assessment that FERC Order 1920 now requires utilities to document. The regulatory mandate and the technical opportunity are the same analysis.
The most expensive grid AI mistake is buying a platform that solves the general problem while missing the specific failure mode that's actually at risk in your footprint.
The Compliance Sprint Is Already Running

For European grid operators, the EU AI Act's August 2026 enforcement deadline isn't a planning horizon — it's nine months away. Conformity assessments for high-risk AI systems in critical infrastructure require documented risk management frameworks, human oversight mechanisms, and technical documentation demonstrating that the system behaves as specified across the range of operating conditions it will encounter. Utilities that haven't begun that process are not in a planning cycle; they're in a sprint.
Post-Iberian-blackout, Red Eléctrica has already authorized 24 renewable facilities for dynamic voltage control under the updated P.O. 7.4 operating procedure, requiring each to demonstrate ±30% reactive power capability. That's a hardware and firmware mandate. The analytics infrastructure to monitor dynamic voltage control behavior at the collector level in real time is the next layer — one that the ENTSO-E report explicitly recommends but that no commercial platform currently provides at scale.
For US operators, the compliance picture is different but equally compressed: NERC CIP-013's supply chain risk management requirements apply to OT-adjacent AI systems on a 6-to-12-month vendor evaluation cycle; CIP-003-9 is already in effect; and FERC Order 1920's GET assessment documentation requirement is active for any utility filing a transmission construction proposal.
The utilities that are ahead of this aren't waiting for a commercial platform to solve it universally. They're building point solutions — narrow, verifiable, auditable — that address specific gaps in their existing monitoring stack while staying within the procurement and compliance frameworks they already operate under.
The Math That Makes This Urgent

The DOE estimates annual US power outage costs at $150 billion. The Iberian blackout alone cost €1.6 billion in a single afternoon. PJM's current capacity auction trajectory projects $163 billion in cumulative costs over five years, driven primarily by the gap between what the interconnection queue can process and what data-center load growth requires.
None of those numbers represent the cost of AI. They represent the cost of the status quo — grids monitored at the wrong voltage level, queues processed at human throughput speeds, transmission capacity underutilized because no one can tell which corridor unlocks the most MW at the lowest cost. At veriprajna.com/solutions/power-grid-ai we've laid out what we've built and where the market is. If you're navigating grid AI procurement, interconnection strategy, or GET assessment requirements under FERC Order 1920, the conversation is worth having — not because there's a universal answer, but because the constraints vary enough by ISO footprint that the specific tradeoffs usually aren't what they look like from a distance.