
The March 2026 ENTSO-E final report on the Iberian blackout landed the same week I was deep in a model tracking interconnection queue restudy cascades for a US client. I wasn't looking for a connection. I found one anyway.
The ENTSO-E investigation concluded that the April 28, 2025 blackout wasn't a renewable-energy failure. It was a monitoring architecture failure: 400 kV transmission readings were nominal while 220 kV collector substations were hitting 242 kV. Transformer tap-changers couldn't compensate fast enough. One facility injected reactive power into an already-overvoltage grid when P.O. 7.4 required it to absorb. Fifteen gigawatts disconnected in five seconds. Sixty million people lost power for up to 10 hours. The CEOE estimated the economic cost at €1.6 billion.
What hit me reading that report was not surprise. It was recognition. The US grid's capacity crisis — PJM falling 6,625 MW short of its reliability target for the first time in history, ERCOT's interconnection queue at 233 GW with only 23 GW of new generation synchronized in 2025 — shares the same structural failure: the monitoring stops at the wrong level. In Spain, it stopped at 400 kV when the failure was happening at 220 kV. In the US, it stops at the transmission-level study process when the real bottleneck is the absence of queue-level intelligence.
That's what I've been building toward at veriprajna.com/solutions/power-grid-ai: not a general grid AI platform, but a set of narrow analytics systems that close the observability gap at the specific levels where it's actually failing.
The Question That Broke My Early Confidence

My assumption when I started working on dynamic line rating analytics was that the deployment bottleneck was hardware costs. It isn't. LineVision has already proven the case: 61% capacity increase on AES's 345 kV corridors, 20-30% increase on National Grid's Syracuse circuits, at roughly 5 to 7% of the cost of traditional reconductoring. The economics are not the problem.
The conversation that corrected my mental model happened with a transmission planning engineer at a mid-size regional utility. They had LineVision sensors approved and budgeted. My question was which corridor they were targeting. The answer surprised me: they'd picked the longest one — not because of any analysis of where new capacity unlocked the most queue value, but because nothing in the process had given them a better basis for choosing.
That's a problem I hadn't framed correctly.
The DLR deployment bottleneck isn't hardware cost — it's the absence of an analytics layer that tells you which corridor to deploy on first.
The DLR prioritization question — which 50 kilometers unlocks the most interconnection queue capacity at the lowest deployment cost — is an AI analytics problem, not a hardware problem. And it matters specifically because FERC Order 1920 now requires utilities to document that they've assessed Grid Enhancing Technologies, including DLR, before proposing traditional transmission construction. That documentation has to be grounded in actual corridor analysis, not in picking the longest line.
Our DLR corridor work starts from the queue side: which transmission constraints are binding across the most interconnection positions simultaneously, and where does DLR deployment reduce those constraints most efficiently. The same constraint mapping feeds the GET assessment documentation that FERC Order 1920 requires. The regulatory mandate and the technical opportunity turned out to be the same analysis — I just had to be told about the "longest line" problem before I could see it.
What the Restudy Cascade Math Actually Looks Like

My early framing of the US interconnection crisis was that it was a process problem — too many projects, too few engineering hours. I built a spreadsheet to model FERC Order 2023's cluster study dynamics and had to revise that framing pretty quickly.
The cascade mechanics work like this: when a project withdraws from the queue, every downstream project that was studied assuming the withdrawn project was in service must be restudied. Under FERC's cluster study framework, those restudies can ripple through dozens of projects in the same transmission zone. The 68% late-completion rate for interconnection studies in 2022 — documented in FERC's own analysis — isn't primarily an engineering-hours problem. It's a restudy-cascade problem. The process generates more work as it proceeds.
The AI application here is queue screening: identifying which projects are most likely to withdraw (based on financial security deadlines, project type, zone congestion profile, and developer history), where restudy cascades would concentrate if they do withdraw, and which transmission constraints are binding across the most queue positions. GridLab's analysis found that if just 10% of the 107 GW of renewables in PJM's queue had been built for the 2026/2027 delivery year, consumers would have saved $3.5 billion in a single capacity auction.
The queue isn't slow because the process is inefficient. It's slow because the intelligence layer that would prevent the restudy cascade from happening doesn't exist yet.
ERCOT is working on it — McKinsey is contracted to overhaul the interconnection process, with a Batch Study framework filed for February 2026 PUC discussion. Process reform is real progress. It doesn't substitute for the predictive analytics that reduce withdrawals and restudies before they happen, though. Those are different problems.
The Vendor Landscape I Had to Map Before We Started Building

I want to be clear about what the incumbents do well, because I spent months evaluating them before building anything.
GE Vernova's GridOS platform prevented 112 million customer-minutes of interruption for Alabama Power in 2025. That's a real operational outcome. Siemens partnered with NVIDIA to run Gridscale X digital twin simulations at 10,000x the speed of traditional approaches — meaningful acceleration for transmission planning studies. Utilidata's Series C at $60.3 million for its NVIDIA Jetson-based Karman chip represents genuinely novel edge computing at the distribution meter level, deployed at Portland General Electric and Duquesne Light.
What none of these platforms address is collector-level voltage monitoring at the 220 kV layer — the exact failure mode the ENTSO-E investigation identified. GE Vernova's SCADA sees what happened at Alabama Power because it was watching the right transmission-level metrics. It doesn't solve what happened in Iberia, because Iberia's failure happened at the collector level that transmission monitoring doesn't reach. I spent a lot of time convincing myself this was a product gap that would close in the next platform release. It isn't. It's an architectural choice, and the incumbents aren't targeting it because their installed base doesn't pressure them to.
Argonne National Lab's GridMind project, unveiled in March 2026, is the most intellectually honest AI approach I've seen for grid operations: a multi-agent LLM copilot for control room operators, explicitly positioned as explainable recommendations rather than autonomous control. It's research-stage with no commercial deployment timeline. The design principle — advisory, not autonomous — is exactly right. The physics gap (LLM recommendations without embedded physical constraints require an additional verification step) is the reason it's still at the research stage.
Why August 2026 Now Runs My Planning Calendar

My planning for the European side of this work runs backward from a single date: August 2, 2026. That's when the EU AI Act's enforcement deadline hits for AI systems in critical infrastructure — power grid dispatch, fault detection, real-time control — with penalties of €15 million or 3% of global turnover for non-compliance. For any European grid operator deploying AI in those workflows, conformity assessment is not a future consideration. It's a current project.
My planning for the European side of this work runs backward from that date. Conformity assessment for high-risk AI systems requires documented risk management frameworks, technical documentation demonstrating specified behavior across the range of operating conditions the system will encounter, and human oversight mechanisms. For collector-level voltage monitoring at the 220 kV layer — the gap the ENTSO-E report exposed — the documentation requirements aren't separate from the architecture. They're part of it. Building a monitoring system that can't produce an explainable audit trail isn't just a product shortcoming; it's a non-compliant product.
Red Eléctrica has already authorized 24 renewable facilities for dynamic voltage control under the updated P.O. 7.4, requiring each to demonstrate ±30% reactive power capability. That hardware and firmware mandate is real. The analytics infrastructure to monitor dynamic voltage control behavior at the collector level in real time is what the post-blackout architecture needs next — and it doesn't exist commercially yet.
For US operators, the timeline is different but equally compressed: NERC CIP-003-9 took effect in April 2026, NERC CIP-013 supply chain risk management applies on a 6-to-12-month vendor evaluation cycle, and FERC Order 1920's GET assessment documentation is active for any utility with a transmission construction proposal in flight. I've found that most utilities with grid AI projects underway have one or two of these compliance threads tracked and one or two they haven't gotten to yet. The intersection of all three is where the architecture decisions get hard.
The Question I Keep Getting That I Can't Answer Cleanly

The hardest question I get from ISOs and utilities isn't about technology. It's about sequence: if we're going to add AI analytics to our grid operations stack, where do we start — queue screening, DLR prioritization, collector-level monitoring, or regulatory compliance documentation?
I don't have a universal answer for that, and I'm suspicious of anyone who does. The answer depends on which constraint is most binding in each ISO footprint — and those differ meaningfully between PJM's capacity auction dynamics, ERCOT's large-load queue concentration, and a European TSO managing post-blackout regulatory requirements under the EU AI Act timeline.
What I can say is that the DOE estimates annual US power outage costs at $150 billion, PJM's capacity auction trajectory projects $163 billion in cumulative costs between 2028 and 2033, and the Iberian blackout cost €1.6 billion in a single afternoon. The cost of the wrong sequence is real. But so is the cost of starting with a comprehensive platform that solves the general problem while missing the specific failure mode in your footprint. The page at veriprajna.com/solutions/power-grid-ai documents what we've built and the constraint mapping that precedes each deployment.
The grid is generating data faster than the tools built to watch it. I don't think the sequence question has a right answer in the abstract. It has a right answer for each grid. That's what I'd want to figure out with you.