
The dominant failure pattern in enterprise facial recognition is one nobody checks for at procurement. The vendor sells you a 99% accurate system. The vendor's contract disclaims any warranty of accuracy for the specific deployment scenario you're about to use. Your procurement team never ran the open-set math. Four thousand false alerts per day across your store network — concentrated in your highest-traffic demographics — is where that chain of decisions lands.
We've reviewed enough biometric deployments to recognize the shape before the first incident report arrives. The failure is rarely about algorithm quality. It's about the gap between how the algorithm was tested and how it's deployed.
The Closed-Set / Open-Set Mismatch

NIST's Face Recognition Vendor Test — the benchmark on every vendor's marketing slide — measures how well an algorithm finds a probe face within a gallery where a match is known to exist. That's a closed-set problem. The "99% accurate" claim is real in that specific context. Paravision, NEC, and IDEMIA consistently rank in the NIST FRVT top tier. That data is accurate and relevant to the test scenario.
Retail watchlist screening is a different problem. The question isn't "which enrolled person is this?" It's "is this person, out of every individual who walks through any of our 500 store doors today, one of the 200 on our watchlist?"
At a store with 8,000 daily visitors and a 200-person watchlist, 97.5% of scans are against people not enrolled in the system. A closed-set algorithm tries to find the best match for every face — and at that volume, even a 0.1% false positive rate generates eight incorrect alerts per day per location. Across 500 locations, that's 4,000 false alerts daily. Employees, untrained in the system's limitations, act on those alerts as if they were facts.
This is the math the Rite Aid deployment never ran. The FTC's consent order required a five-year ban on facial recognition, destruction of all photographs and videos collected, and deletion of every model or algorithm derived from that data. Amazon Rekognition and Microsoft Azure Face placed indefinite moratoriums on police use of their systems because the open-set deployment math, applied in law enforcement contexts, produced outcomes their legal teams couldn't defend. IBM exited facial recognition entirely in 2020.
FTC model disgorgement — an order to destroy the algorithm itself, not just the data — is now an active enforcement mechanism. It was applied again in May 2025 to an edtech company. The principle is explicit: companies cannot profit from algorithms built on improperly collected biometric data, even if the original collection was years ago.
What the NIST Report's Second Column Shows

The same NIST FRVT report that shows top-tier accuracy also shows something most procurement teams never look at: within-group false positive rates vary by up to 7,203x across demographic populations. Children, elderly individuals, and people from specific racial and ethnic groups experience false positive rates thousands of times higher than the populations those algorithms were optimized on.
The 99% accuracy figure and the 7,203x demographic variance are in the same NIST report. Vendors send you the leaderboard column. The audit starts in the demographics appendix.
When the Rite Aid system was reviewed, the FTC found stores in plurality-Black and Asian communities generated significantly more false alerts than stores in majority-white communities. Employees were systematically following demographic patterns in their alert responses without knowing the pattern existed. That evidence became a central pillar of the enforcement case.
If your vendor selection process didn't include a NIST FRVT demographic breakdown analysis — benchmarked against the specific demographic composition of your customer base across each deployment market — your legal exposure isn't hypothetical. Clearview AI's $51.75M BIPA settlement put a number on it. That wasn't a rounding error.
The Gallery Nobody Audited

The algorithm gets benchmarked. The enrollment database doesn't.
Most retail and financial watchlists are assembled from heterogeneous sources: some controlled headshots, many grainy CCTV stills, some booking photos a decade old. Each low-resolution, poorly lit gallery image increases the probability of a partial match against the thousands of non-enrolled faces flowing through your locations daily. The algorithm didn't fail. The gallery contaminated the input before a single scan ran.
Replacing low-resolution CCTV stills with controlled-lighting headshots and pruning stale entries typically cuts false alert rates by 60–80% without modifying the algorithm. This is the highest-leverage intervention in most deployments we assess — and it's not in any vendor's SLA, because the vendor is responsible for the model, not the data you're running through it.
Enrollment database hygiene is the first thing we assess at Veriprajna's biometric compliance practice. It's also where we find the fastest, most defensible improvements before touching anything else in the stack.
108 Days in Jail, 1,200 Miles From the Crime

Angela Lipps, a grandmother from Tennessee, was arrested in July 2025. She spent 108 days in jail. She was 1,200 miles from the crime scene when it occurred. Fargo police acted on a facial recognition match. Charges were dismissed Christmas Eve 2025. The Fargo police chief issued a public apology March 27, 2026. Civil rights claims are in preparation.
The Harvey Murphy case — wrongful detention following a facial recognition match — produced a $10 million lawsuit. The Washington Post documented at least eight Americans wrongfully arrested following FR matches, with investigators in each case skipping basic verification steps like alibi checks.
In every case, the match score was treated as evidence. Nobody documented whether the match was reliable given the image quality, the age gap between the probe and the gallery image, or the algorithm's demographic performance on the subject's population group.
HITL — human in the loop — is the governance language in vendor contracts, policy documents, and regulator guidelines. What most enterprises actually have is HITL theater: a human who reviewed an alert and acted on it without documenting any assessment of the underlying algorithm's reliability for these specific conditions. Regulators are converging on "meaningful, not ceremonial" oversight as the enforcement standard. A review process that doesn't capture whether the reviewer assessed algorithm confidence limits for the specific demographic and image quality at issue is unlikely to survive that standard.
Your Store Network Is Not One Jurisdiction

Illinois BIPA operates with a private right of action — any affected individual can sue — and imposes $1,000 to $5,000 per violation. There were 107 new class actions filed in 2025 and $136.6M in total settlements that year. Texas CUBI imposes penalties up to $25,000 per violation under AG enforcement, with a $1.375B Google settlement establishing the price-of-scale precedent. Colorado amended its Privacy Act in July 2025 to explicitly cover biometric identifiers. Washington state requires consent before enrollment in a biometric database.
Then there are the city-level bans — 16+ US municipalities with active prohibitions on retail or government facial recognition, including San Francisco, Boston, Oakland, and Portland. Amazon's Ring "Familiar Faces" feature launched in December 2025 and was blocked in Illinois, Texas, and Portland within weeks.
No federal US law currently governs commercial facial recognition. What exists is a patchwork of state statutes, city ordinances, and international regulations with overlapping consent requirements and non-overlapping penalty structures.
The EU AI Act adds another layer for any operation with European exposure: real-time remote biometric identification is prohibited except for specified law enforcement exceptions, with high-risk system conformity assessments due by December 2027 and penalties up to €35 million or 7% of global turnover.
Multi-jurisdiction exposure isn't a compliance complexity problem to be managed uniformly. A retailer with stores in Illinois, Texas, and Colorado is managing three distinct consent mechanics, three enforcement regimes, and three penalty structures simultaneously — often with identical system configurations across all locations, because the vendor didn't differentiate and nobody on the enterprise side knew to ask.
What an Independent Audit Actually Changes

The vendor contract created the liability gap. Nobody inside the vendor relationship closes it.
An independent audit begins where the vendor SLA stops. We start with NIST FRVT demographic performance — not whether your vendor ranks highly on the leaderboard, but whether that ranking holds for the specific demographic composition of your customer base in each deployment market. Then the enrollment database: not whether the gallery is populated, but whether the image quality would hold up as a contributing factor if a wrongful-match outcome were litigated. The jurisdictional map follows — consent mechanics, retention schedules, and HITL documentation aligned to each specific jurisdiction in your deployment footprint, not mapped once at the national level and assumed to hold everywhere.
The model disgorgement question tends to be the one nobody has run. Not whether the algorithm is good, but whether the consent conditions under which you collected your gallery were clean enough that an FTC investigation tomorrow couldn't order the model destroyed retroactively.
The audit is a pre-litigation posture review — run before opposing counsel does the same assessment and files a brief using your answers.
Veriprajna's biometric compliance practice was built for the gap the Big 4 treat as a privacy footnote: the specific intersection of NIST benchmarking, open-set deployment math, enrollment database hygiene, and multi-jurisdiction exposure mapping. Most general privacy audits don't go there. The liability does.
The biometric compliance question many teams are deferring tends to become significantly more expensive after the first incident than before it.