The fourth verdict I had to add to the gate
I had to add a fourth verdict to the policy gate because Illinois HB 3773, in force since January 1, 2026, bans zip code as a protected-class proxy while EU AI Act Article 10(3) leans on exactly that geographic coverage to show the training data is representative, and one column cannot be both. On the demo's seeded 1,040-candidate synthetic export, zip_region correlates with race at a Cramér's V of 0.3321 (bias-corrected 0.3253), above the 0.2 threshold I set, so one column is a banned proxy in Illinois and evidence of representativeness in Brussels. Mask it and the EU deliverable fails; keep it and Illinois is violated. No single model configuration does both, and I could not find an honest way to render that as a colour.
So the gate emits CONFLICT, and a reconciler agent writes a legal-strategy memo where the certification would have gone: run two deployment configurations (a full geographic mask for Illinois inference, a coarser region feature for EU training data), or choose the regime with the larger exposure and document the accepted risk, noting that EU AI Act high-risk penalties reach the greater of EUR 15M or 3% of global annual turnover. The memo ends by withholding the certification rather than printing one. I wrote that refusal into the product deliberately, and it is the decision I have had to defend most.
The Conflict Register on the synthetic export: the zip-code proxy caught between Illinois HB 3773 and EU AI Act Art. 10(3), with both branches quantified and the certification withheld.
What I stopped letting the model touch
I moved every statistic out of the agent framework and into plain numpy in backend/engine.py, and the rule I have not broken since is that agents advise and code decides. Six jurisdiction agents (Pydantic AI, provider-swappable) narrate each regime's verdict, the reconciler turns a hard conflict into that strategy memo, and an adversarial skeptic tries to refute every pass I claim. None of them computes a number or sets a status. With no provider configured the crew falls back to deterministic templates and the result does not move, because the impact ratios, the proxy correlations, the threshold comparisons and the gate itself are code an auditor can re-derive without me in the room.
I only trusted the coverage result once the agents could no longer touch it. One run over the synthetic export from three simulated tools (a Workday-Spotlight-style scorer, a HireVue-style video round, an Eightfold-style match engine, all fixture adapters over synthetic data rather than integrations) produces six jurisdiction-shaped deliverables, each carrying its own citation, effective date and required format. Across NYC Local Law 144, Colorado SB 24-205, Illinois HB 3773, Texas TRAIGA, California's FEHA ADS amendments and EU AI Act Annex III, the engine evaluates 13 obligations and hands me back 2 PASS, 2 FAIL, 7 NEEDS_PROOF and 2 CONFLICT. The console's headline tiles read 6 jurisdiction deliverables, a lowest impact ratio of 0.65, and 9 items requiring human action.
Six deliverables from one audit run over the synthetic export, each shaped to its own regime, effective date and format. Three different verdicts across six rows, which is the output I was building toward.
Two failing cells, and only one of them survives FDR
I ran the vendors' test before I ran mine, because I wanted to know whether my own reveal was doing real work. The marginal four-fifths check passes on this data: race minimum impact ratio 0.8196, sex minimum 0.8744, both at or above the 0.80 line. That is not a strawman I planted to knock over. It is the real computation on the synthetic export, and it really passes, which is where a vendor self-audit stops. The same data cut by race and sex, which is what LL144 actually requires, puts the Black / Female cell at 44 of 130 advanced (33.85%) against a White / Male reference of 68 of 130 (52.31%), an impact ratio of 0.6471. Hispanic / Female lands at 0.7647.
Two failing cells would have made the better screenshot, and that is exactly where I made the engine harder on itself. Under Benjamini-Hochberg FDR control at alpha 0.05, only Black / Female is statistically robust (q = 0.0185). Hispanic / Female comes back at q = 0.1629, so the engine flags it and refuses to call it significant. Selling a chance finding as a robust one is the same failure as hiding a real one, just pointed the other way, and since this disparity is planted and synthetic, built so the method has something to catch, the only thing worth judging is whether the method separates the two.
The passing marginal rate card and the intersectional grid in the same dialog, both computed on the seeded synthetic export. The tile rounds to 0.65; the engine's exact figure is 0.6471.
The three challenges I let the skeptic keep
I gave the adversarial skeptic permission to refuse my own passes, and the three refusals it kept are separate legal theories that a bias audit never tested. The first is scope: a vendor memo asserting "our scorer is not an AEDT" gets rejected, because under the Mobley v. Workday agent theory a tool that recommends or filters candidates sits inside the decision, so the demo routes that question to human counsel for a scope attestation instead of resolving it in software. The NY State Comptroller found 17 potential LL144 violations in the same 32-company sample where DCWP found one, and DCWP agreed to shift to proactive enforcement (NY State Comptroller, December 2, 2025). Self-classification is not a defense I want an engine accepting on an employer's behalf.
The second refusal is about accessibility, and I did not go looking for it. Across the 432 video-round candidates in the synthetic export, word error rate runs 0.0794 on standard speech and 0.3016 for the 104 candidates with non-standard speech, a 3.8x disparity that LL144 never measures, and it echoes the theory raised in D.K. v. Intuit/HireVue. The third is FCRA: 510 of the synthetic candidates were scored from third-party-scraped data and filtered on a numeric score, the pattern at issue in Kistler v. Eightfold, and if the platform is a consumer reporting agency then every scored candidate is owed an adverse-action notice and a dispute path no matter how fair the model turns out to be. None of those matters is decided. Clarion detects both of the last two and routes them to a named human, and it does not build the accommodation workflow or the candidate dispute portal, which I would rather state plainly than let a queue item imply. Those three challenges live in the Human Proof Queue, a different object from the 9 obligations the gate sent to a human.
Seventeen nodes, and none of them mine to sign
I built the export for somebody else's signature. Every run seals a hash-chained pre-audit package, veriprajna-pre-audit-package v1, 17 nodes, each node carrying its inputs, its computation and its rule citation plus a SHA-256 link to the node before it, so any edit anywhere breaks the chain. It comes out as JSON and as a printable HTML packet whose integrity line reads VERIFIED, and the tamper-evidence is one of the 12 tests in app/tests/test_engine.py that pin the planted ground truth on every run.
I am not the independent auditor, and the packet says so in its own header: the LL144 sign-off belongs to firms like DCI, ORCAA or Secretariat. That distinction matters more than it sounds, given that 4.6% of 391 NYC employers had published a bias audit at all (Cornell / Data & Society / Consumer Reports, FAccT 2024). My job is to get a hiring stack into the state where the independent auditor can sign without rewriting the work first. The walkthrough of the whole run, the conflict memo and the packet included, is in the AI hiring compliance demo.
The printable auditor packet on the seeded synthetic export: 17 hash-chained nodes verified, every verdict carrying the citation it came from, and the sign-off left to the independent auditor.
What I did not expect going in was that the refusal would be the hardest thing to build and the easiest thing to defend. A dashboard that says compliant is, in a courtroom, an exhibit. A memo that says you are knowingly carrying either Illinois exposure or EU exposure on one geographic feature, with the size of each spelled out, is a record. I would rather hand a General Counsel the second one.