
I read the Plano post-mortem about six months after the incident made its way into the trade press. Aclara pushed a firmware update to 88,000 water meters in November 2024. 73,000 went dark. The city hired 20 temporary meter readers. The emergency response cost $765,000 over two years.
What stopped me wasn't the scale. It was a footnote in the firmware release notes describing the test configuration: new batteries, strong RF signal, controlled ambient temperature. Nothing about cells at 60–75% capacity after four years of operation. Nothing about the brownout protection behavior when the flash write current slightly exceeds what a degraded battery can sustain.
The firmware was tested against a fleet that didn't exist in the field. I've since looked at documentation from every major AMI vendor, and the pattern is the same.
Our work on AMI firmware validation and endpoint health scoring is at Veriprajna's smart meter AI page.
My First Assumption Was Wrong About Aclara

My first read of the Plano incident was that it was an Aclara execution problem — a quality control failure specific to that vendor and that release. Then I found the Minneapolis documentation, the Toronto reports, the New York City incidents. All Aclara fleets. All the same failure signature.
Toronto Hydro's story runs through a different mechanism: 470,000 transmitters degrading as meters logged data at 15-minute intervals, burning through NAND flash write cycles faster than the manufacturers' 20-year projections assumed. Consumption readings drifted 2–8%. Initial remediation: $5.6M. Memphis is running a $9M repair fund for 8% systemic fleet failure.
The $15.4M in combined remediation across Plano, Toronto, and Memphis isn't bad luck. It's the result of a testing methodology that every AMI vendor uses and that no utility has a systematic process to supplement.
When I reached the point where I understood this, the question changed from "why did Aclara do this?" to "what does a utility need to do differently before the next firmware push?"
The MDM Dashboard That Explained the Problem

My clearest memory from our early work on this is a utility director pulling up his MDM platform and showing me the fleet health view. Green across the vast majority of endpoints. Communication success rates looked strong. The system was doing exactly what it was designed to do — tracking which meters were communicating, flagging the ones that went silent.
I tried to explain that the signals I was watching weren't in that view. RF signal margin trends over time, not point readings. Battery voltage degradation curves, not current status. Flash memory write-cycle counts. Several of those green-reported meters were approaching the battery-age threshold where the next firmware push would behave differently than the vendor's lab result predicted.
His question — why the head-end wasn't showing any of it — clarified the architecture problem.
It doesn't, because MDM platforms from Itron, Oracle, and SAP were designed for consumption data, billing management, and communication status. Itron's Distributed Intelligence platform covers more than 100 million endpoints and now runs NVIDIA edge-AI inference at the grid edge — for load disaggregation and demand response analytics, not endpoint health scoring by battery-age cohort. Landis+Gyr's Revelo and Sensus/Xylem's Evolve, launched February 2026, pursue similar directions. Vendor analytics optimize for what the vendor sells. A mixed fleet has no vendor-neutral platform watching the pre-failure indicators that cross vendor boundaries.
The NERC Gap I Found in the Coverage

What I noticed when NERC CIP-003-9 took effect on April 1, 2026, was a gap in how the AMI industry covered it. The discussion was almost entirely about grid management systems — SCADA, EMS, access controls for high-impact facilities. The firmware OTA implications for smart meters received almost no attention.
I noticed because I was specifically looking for it. Part 1.2.6 requires documented access-revocation procedures for every vendor with electronic remote access to low-impact BES Cyber Systems — which covers most smart meters. A utility pushing firmware through a vendor's OTA channel needs those procedures in place on day one.
What that means operationally: a firmware deployment that disables a significant portion of the fleet now sits inside a compliance accountability structure that didn't exist a year ago. Were the vendor's remote-access controls documented? Can the utility produce the access-revocation procedures in a post-incident audit? NERC penalties run up to $1M per day.
When I brought this framing to utility legal teams, it landed faster than the operational failure math. The operational risk was understood — Plano was publicly documented. The compliance dimension wasn't.
Why the 29% Number Is the One That Stays With Me

The silent failure rate documented by Electric Energy Online is the number I come back to most: 29% of installed endpoints at some utilities failing silently — not dark enough to trigger communication alerts, but not healthy enough to produce reliable readings. Radio modules degraded. Flash memory corrupted. Batteries below the brownout threshold for cold-weather events.
What stays with me about 29% is that it generates no incident report. Plano was dramatic — 73,000 meters offline, an emergency response, a news cycle, someone filed a post-mortem. The 29% silent failure rate generates billing disputes that erode customer trust over months, and eventually a compliance audit that surfaces what the monitoring platform missed.
The consumption readings in the 2–8% drift range Toronto documented aren't a headline. They're a slow deterioration in the foundation the billing system assumes is solid.
The Questions I Ask Before Recommending Any Firmware Push

Before I recommend that a client authorize an OTA firmware update on a fleet with significant battery age, I need three things answered.
What percentage of the deployed fleet has batteries older than four years, and what is the firmware vendor's documented test configuration for that battery condition? The vendors can usually answer the first part. Most cannot answer the second — their release notes have the footnote I found in Plano.
Whether RF signal margin trending is available by endpoint for the past 90 days matters more than the current reading. A meter with stable strong signal is a different risk than a meter whose margin has been steadily declining for three months. The head-end usually has this data; most utilities have never extracted it for pre-firmware analysis.
Flash memory write-cycle count is the least-tracked variable. For most utilities, it's captured at the meter and never pulled into the analytics stack. That's where NAND flash degradation like Toronto's hides until a meter stops accepting firmware patches entirely.
The pre-deployment staging architecture for addressing all three is at Veriprajna's smart meter AI page.
What I keep coming back to is the vendors. Itron, Landis+Gyr, Sensus are building the intelligence layer for the next generation of AMI — grid-edge inference, demand response, load disaggregation. The meter health problem that caused Plano, Toronto, and Memphis sits in the gap that intelligence layer wasn't designed to close. Until the testing methodology changes, the failure pattern travels with the firmware — to whatever city gets the next update.