
The first time I opened an enrollment database audit for a retail client, I asked to see the first image in their facial recognition watchlist. What came up was a grainy CCTV still — low-res, poor lighting, face partially occluded, from what appeared to be a decade earlier. Nobody in the room knew what images were in their watchlist. The algorithm was fine. The input was not.
I've litigated BIPA cases from the plaintiff's side and the defense side, and now I spend most of my time doing the kind of pre-litigation audit that neither side ran before the incident. The pattern I keep finding isn't a bad algorithm. It's a good algorithm running in the wrong operational scenario, against an enrollment gallery nobody audited, in a jurisdiction the deployment team didn't know they were in.
That's the argument behind Veriprajna's biometric compliance practice. Let me walk through the pieces.
The False Positive Math Nobody Ran at Procurement

When I sit down with a new client and walk through the open-set false positive rate calculation, the room usually goes quiet.
NIST's Face Recognition Vendor Test — the benchmark on every vendor's marketing slide — measures closed-set performance: how well the algorithm finds a match within a gallery where a match is guaranteed to exist. Paravision, NEC, and IDEMIA legitimately rank at the top of that test. The 99% accuracy figure is real in that specific context.
What retail watchlist screening runs is an open-set problem. The question isn't "which enrolled person is this face?" It's "is this person, out of the thousands who walk through our doors today, one of the 200 on our watchlist?"
At 8,000 daily visitors and a 200-person watchlist, 97.5% of scans are against people not in the database. A closed-set algorithm finds a "best match" for every face — and even at a 0.1% false positive rate, that's eight incorrect alerts per day per location. Across 500 stores, that's 4,000 false alerts daily. Employees treat those alerts as facts.
Amazon Rekognition and Microsoft Azure Face both placed indefinite moratoriums on police use of their systems because the math, applied in law enforcement scenarios where a false positive can cost someone their freedom, was producing outcomes their legal teams couldn't defend. IBM exited the market in 2020. The vendors that stayed in market disclaimed the accuracy warranty in their contracts. That disclaimer transfers the open-set deployment liability to the enterprise that signed it.
Most procurement teams don't model what that transfer means until the incident report arrives.
What the Rite Aid Order Meant for the Clients Who Called That Week

I had a few conversations the week the Rite Aid FTC consent order was published. One has stayed with me.
A general counsel called — a different retailer, a different vendor — and her question was specific: they had been collecting gallery images for several years. Did the disgorgement requirement from the Rite Aid order expose them?
What I had to explain is that the answer wasn't in the product specs. FTC model disgorgement doesn't just delete the photographs. It deletes every model weight trained on improperly collected data: the fine-tunings, the algorithmic derivatives, anything built on that gallery. Rite Aid's consent order required a five-year ban on facial recognition and deletion of every model or algorithm derived from the data. The FTC applied the same principle to an edtech company in May 2025. The pattern is explicit: you cannot profit from an algorithm built on improperly collected biometric data, even if the original collection was years before the enforcement action.
The GC's question turned out to be the right one. The answer required going back to the data collection conditions from years earlier, before anyone had thought to model disgorgement as a distinct risk category.
The $136.6M in BIPA settlements in 2025 and the $1.375B Google CUBI settlement in Texas are the visible outcome numbers. Disgorgement risk is the hidden one: the possibility that years of model development get ordered destroyed because the consent mechanics around the training data weren't clean. Clearview AI's $51.75M BIPA settlement included this dimension — not just the data collection, but the models built on it.
The Demographic Column the Vendor Didn't Send You

I started carrying an annotated version of the NIST FRVT demographics appendix after I realized vendors were forwarding clients only one column from the report.
The same document that shows 99% top-tier accuracy also shows that within-group false positive rates vary by up to 7,203x across demographic populations. Children, elderly individuals, and people from specific racial and ethnic groups experience false positive rates orders of magnitude higher than the populations those algorithms were optimized on. That number is not buried — it's in the published NIST FRVT data.
What I kept noticing is that the demographic performance gap isn't obscure — it's in the published NIST report, right next to the accuracy headline. Vendors excerpt the headline. The demographics appendix stays unread.
When the FTC reviewed Rite Aid's deployment, they found stores in plurality-Black and Asian communities generated significantly more false alerts than stores in majority-white communities. Employees were following demographic patterns in their alert responses without knowing the pattern existed. That evidence was a central pillar of the enforcement case.
My practice now is to cross-reference the NIST demographic breakdown against the actual demographic composition of each client's deployment market. The mismatch between where the algorithm performs reliably and where the stores are located is often significant. The audit question isn't whether the vendor ranks well on the NIST leaderboard — it's whether their demographic performance profile matches the population you're actually scanning.
Forty Seconds Per Alert

The review log I remember most clearly came from a client with multiple years of HITL documentation.
Every entry said "confirmed" or "not confirmed" with a timestamp. I added up the review times across a month of alerts. Reviewers were averaging roughly 40 seconds per alert. There was no image quality note. No demographic confidence note. No documentation of what the reviewer considered about the algorithm's reliability for the specific face, age, image conditions, and demographic at issue. I didn't say anything when I saw it. The compliance lead was reading the same data.
That review log is HITL theater. A human touched the alert. Nothing about the human's reasoning was recorded.
Angela Lipps spent 108 days in jail after Fargo police acted on a facial recognition match — she was 1,200 miles from the crime at the time it occurred. The Fargo police chief issued a public apology March 27, 2026. The Harvey Murphy wrongful detention case produced a $10 million lawsuit. The Washington Post documented at least eight similar wrongful arrests. In every case, the match score was treated as evidence and no one documented whether that score was reliable given the image quality, the age gap, or the algorithm's demographic performance on the subject's population.
Regulators are converging on "meaningful, not ceremonial" oversight as the enforcement standard. My working definition of meaningful is: the reviewer recorded what they assessed about the algorithm's reliability for these specific conditions. Not just what action they took. Whether the documentation would survive as a defense if the outcome were litigated.
The City-Level Map Nobody Had Built

I've learned to ask about city-level bans early in every engagement, because I've found out about them mid-engagement often enough that the question is now standard.
Sixteen-plus US municipalities have active prohibitions on retail or government facial recognition — San Francisco, Boston, Oakland, Portland, and others. Amazon Ring launched a "Familiar Faces" feature in December 2025 and it was blocked in Illinois, Texas, and Portland within weeks. That timeline — product launch to regulatory block — is the horizon enterprise legal teams should be planning against.
Most deployments I review have a national compliance posture. BIPA is checked. Texas CUBI is checked. A federal privacy policy is posted. The city-level map is usually missing. One client discovered a Portland location during an engagement that had been framed as a state-level compliance exercise. Backing up to add the city layer took time they hadn't planned for.
The full stack for a multi-state retailer: Illinois BIPA with private right of action, 107 new class actions filed in 2025 and $136.6M in settlements. Texas CUBI at $25,000 per violation under AG enforcement, with a $1.375B Google settlement as the scale precedent. Colorado and Washington added biometric consent and retention requirements in 2025. EU AI Act prohibitions on real-time remote biometric identification apply to any operation with European exposure, with conformity assessments due December 2027 and penalties up to €35 million or 7% of global turnover.
Each of these has different consent mechanics, different enforcement posture, and different penalty structures. A system configuration designed for one doesn't automatically satisfy the others.
The Question That Actually Matters
I've been asked some version of the same question by nearly every client who finds Veriprajna's biometric compliance practice: how accurate is our system?
I've stopped trying to answer that question directly, because the accuracy they already know — the NIST FRVT figure the vendor provided — is measured in a scenario they're not running. What I ask instead is: what happens when the system is wrong? That question surfaces whether there's a documented review process that produces defensible reasoning, whether the enrollment gallery would hold up as a contributing factor in a wrongful-match case, and whether the consent conditions around the gallery's collection were clean enough that an FTC investigation tomorrow couldn't order the model destroyed retroactively.
The Angela Lipps apology in March 2026 and the Rite Aid consent order aren't the worst-case scenarios for enterprise biometric deployments. They're what happens when that question wasn't asked at the right time.