A vision model that scores 97 percent in lab validation can collapse to a 14 percent false-reject rate on a 200-ton progressive die press running 40 strokes a minute (Veriprajna WP07 research, 2026). Not because the model got worse. Because the inputs moved. Overhead bay lights throw glare that varies with stroke angle, lubricant pools differently on cold dies than warm ones, and the first 50 parts of every shift run before the press reaches thermal equilibrium. Each of those effects walks the images outside the distribution the model was validated on, and no model, at any accuracy, is trustworthy out there. That is the structural problem with edge AI quality inspection as the industry currently buys it: the expensive failure is an input failure, and a better model cannot see it coming, because the model is the thing being fooled.
We built the Inspection Trust Gate to make that argument runnable instead of rhetorical. It is a runtime trust layer that sits between the vision model and the PLC reject actuator and decides, per part, whether the model's verdict is safe to actuate. You can run it at https://veriprajna.com/demos/edge-ai-manufacturing-inspection. The structural case comes first, though, because most of the money in this category is still going to the wrong layer.
The industry keeps funding the wrong layer
Out of the box, automated optical inspection systems false-reject 5 to 15 percent of good parts; well-tuned systems get under 2 percent (Veriprajna WP07 research, 2026). That gap is, in the words of the same research, "a calibration and data problem, not a model architecture problem." The delivery statistics say the same thing from the other side: 84 percent of system integration projects fail or partially fail, and in a typical inspection deployment the integration work is 60 percent of the project timeline while model training is 15. The hardware is a purchase order. The model was never the project. The project is everything that has to hold when the model meets the physics of a real line.
Two clocks make this urgent rather than academic. Deloitte predicts agentic AI adoption in manufacturing rises from 6 percent to 24 percent in 2026 (Deloitte, 2026). And on August 2, 2026, the EU AI Act's high-risk obligations become fully applicable, with safety-critical quality decisions sitting in Annex III and maximum fines of €35M or 7 percent of global turnover for the most serious, prohibited-practice violations (Veriprajna WP07 research, 2026). More autonomy is arriving on the line at the exact moment regulators begin asking for per-decision evidence.
A layer that knows when not to trust the model
The durable fix is not a better model. It is a layer that knows the model's validated envelope and refuses to actuate outside it. In the Trust Gate, that layer starts with an EnvelopeDetector: a Mahalanobis distance computed in physical-signal space (exposure, contrast, dynamic range, focus, high-frequency detail, thermal colour cast, saturation, glare fraction), fitted on the 220 known-good training images. It asks one question per part: is this image inside the capture conditions the model was validated on? If not, no model output is trusted, however confident it sounds.
Downstream of that check, a deterministic gate written in plain code, outside any model and any LLM, applies thresholds fitted from the demo's own data (auto-pass below 0.948, auto-reject above 1.30 on the calibrated confidence scale) inside a hard 750 millisecond stroke-window budget, and routes every part to AUTO_PASS, AUTO_REJECT, or HOLD for human review. The defect model itself is deliberately a stand-in: a kNN texture detector behind a fixed interface where a plant's NVIDIA Metropolis, Cognex, or custom model would sit. The product is the layer around it, and the layer is the durable part: model accuracy is a moving target that every vendor release resets, while drift-gating, provenance, and governed actuation hold at any model accuracy.
There is an agentic piece, bounded on purpose. When parts pile up in the HOLD queue, a diagnosis agent reads the ranked physical-signal deviations and proposes a root cause, and a critic agent checks that hypothesis against the numeric evidence, downgrading it to manual investigation if the cited signal is not actually the dominant deviation. With no API key configured, the triage degrades to a deterministic fallback, so the gate, the metrics, and the audit run fully offline. The agents advise. The gate has already decided.
Even a perfect defect model is only valid on inputs inside its validated envelope. Drift is an input failure, not a model failure, which is why a runtime trust gate does not age out when a better model ships.
Watch a shift change break the model but not the line
The demo replays a scripted shift over real MVTec AD metal_nut held-out test parts: real photographs of real manufactured parts. Drift is honest image corruption (glare, defocus, thermal cast) applied to those real good photos, and the app labels it as exactly that. At the marker "SHIFT CHANGE 06:00 - cold dies, bay lights on," twelve held-out good parts arrive corrupted. The envelope monitor goes red, and the gate routes every one of them to HOLD for human review instead of firing the actuator. The naive baseline, the identical defect model with no envelope check, false-rejects them.
The drift moment mid-run. A real good part arrives defocused after "SHIFT CHANGE 06:00 - cold dies, bay lights on"; the gate holds it as out of envelope in 26.1 ms while the trace notes the naive AOI would REJECT. The comparison panel at this point reads 5 good parts scrapped on the naive path, 0 on the gated path.
Click into any held part and the gate shows its work. On held part test-good-297 the envelope drill-in reads: this frame is outside the conditions the model was validated on. Mahalanobis distance 30.808 against a threshold of 10.462, and the deviation ranked signal by signal, with brightness at plus 7.6 sigma carrying 87.1 percent of the squared distance. That is not a confidence score. It is a physical measurement of why this specific image cannot be trusted.
The breach explained signal by signal on test-good-297: Mahalanobis distance 30.808 against a threshold of 10.462, with brightness at plus 7.6 sigma carrying 87.1 percent of the total. The panel also states the method: fit on the 220 train-good images only, independent of model quality.
The demo's benchmark run makes the same comparison across whole drift families on the MVTec metal_nut held-out split. The naive baseline false-rejects 95.5 to 100 percent of drifted good parts per family, 98.9 percent on average. The Trust Gate holds 100 percent of them and auto-scraps none. The corruptions are full strength, so the honest claim is the direction, not the decimal: a cleanly validated model collapses once its inputs leave the envelope. Two results underneath that headline matter more. The envelope detector scored AUROC 1.000 on a drift family (underexposure) it was never tuned against, which is the independence check that says the gate is not just grading the axes we built it on. And defective parts injected during the drift were still caught or escalated, 93 out of 93 on the held-out split. Nothing hides under drift.
The shift's other beats follow the same discipline. A gross structural defect auto-rejects, with the simulated actuator logging the would-be reject in 25 ms against the 750 ms budget. And the geometric zone rule, an honest stand-in for production metrology, is one we measured ourselves: it changes the gate's outcome on exactly 1 of 93 held-out defects, and the UI discloses that near-inertness per part, because a trust layer that oversells itself is a contradiction in terms.
By the end of the drift phase the comparison panel reads 12 good parts scrapped on the naive path against 0 on the gated path, with the run's measured false-reject rate going from 57.1 percent to 0. A projection panel then extrapolates what that naive rate would cost if the drift ran uncaught for a full 8-hour shift at 40 strokes a minute: roughly $26.5K. The dollar figure is always labeled a projection, because it is one. The $2.42 per-part scrap cost behind it is a sourced order of magnitude from a published cookie-manufacturer case ($94K in annual savings from an 8.7 percent scrap-waste reduction), not a customer's number (Veriprajna WP07 research, 2026).
The payoff panel at the end of the drift phase: 12 good parts scrapped by the naive baseline, 0 by the gate, 12 held for review, the measured false-reject KPI at 57.1 percent to 0, p99 latency 35.6 ms against the 750 ms budget, and the shift-cost extrapolation labeled as a projection.
The receipt, because August 2026 is coming
A deterministic gate decides what actuates, not a model and not an LLM, and every decision it takes writes a per-part audit lineage record, exportable as JSONL.
Every decided part gets one: part id, station, model id and version, dataset hash, defect confidence, OOD score, the physical signals, the exact rules that fired, latency against the 750 ms budget, the actuation log, the baseline decision, and the risk tag high-risk:quality-gate (EU AI Act Annex III, eff. 2026-08-02). One click exports the shift as inspection_audit.jsonl. We are deliberate about the framing: this is EU-AI-Act-ready lineage, evidence designed to be filable in a high-risk conformity file. It is not a certification, and we do not claim one.
One held part's full lineage record: model metalnut-defect-knn v7, dataset hash ae95b5b533c8, the rules that fired, latency 26.9 ms, the simulated actuation and MES traceability lines, the baseline REJECT the gate overrode to HOLD, and the risk tag high-risk:quality-gate (EU AI Act Annex III, eff. 2026-08-02).
What we are claiming, and what we are not
The Inspection Trust Gate is a demo, not a deployment. The EtherNet/IP to Allen-Bradley ControlLogix reject actuator is a simulated adapter that logs what it would do, the MES sink is a stub, and the camera is a recorded stream of MVTec AD photographs. All of its numbers are measurements on that benchmark's held-out split under full-strength synthetic corruptions, not open-world guarantees. What the demo proves is the mechanism: the input can be checked before the verdict is trusted, deterministically, at stroke speed, with a receipt for every part. That mechanism survives every model upgrade a vendor will ever ship you. The full run, with the shift replay and the exportable audit, is at https://veriprajna.com/demos/edge-ai-manufacturing-inspection.
One question for anyone running vision inspection on a real line: when the bay lights change or the dies run cold, what in your cell notices that the inputs have left the conditions the model was validated on? Is there anything between the model's confidence and the reject actuator, or does the verdict go straight to the PLC? We would genuinely like to hear how your plant handles that seam. The problem is industry-wide, and the good answers will be too.