
The exclusion isn't happening because your procurement AI was designed to discriminate. It's happening because it was designed to be efficient—and efficiency, applied to historical spend data, locks in every advantage the incumbent supplier already has.
That's the finding sitting inside every major source-to-pay platform right now. SAP Ariba, Coupa, GEP, and Ivalua have all made significant AI investments in supplier scoring. None of them publish fairness metrics. For federal contractors with FAR Part 19 subcontracting obligations, that absence isn't a vendor gap—it's a compliance liability sitting inside your next contracting officer desk review.
We've been studying how procurement AI produces disparate impact, and the mechanism is consistent enough across platforms that it's worth documenting publicly.
The Confidence Weight Is the Problem

Consider a sourcing event for any commoditized category where your S2P platform's AI scores suppliers on delivery performance, quality, financial stability, and price. A large incumbent with 4,200 historical transactions and a 97.2% on-time delivery rate generates a confidence-weighted delivery score of roughly 24 out of 25. A certified minority-owned supplier with 180 historical transactions and a 98.1% on-time delivery rate generates a score closer to 17 out of 25.
The minority supplier has the better delivery rate. The algorithm rewards the incumbent anyway—because confidence weighting equates "more historical data" with "more reliable." The supplier with fewer transactions gets penalized not for their actual performance, but for not having yet accumulated the transaction history that only comes from winning contracts.
The pattern extends to quality metrics, where audit frequency correlates with contract volume, and to financial stability assessments, where revenue size functions as a proxy for risk tolerance. By the time price competitiveness is evaluated, the gap is typically insurmountable.
The scoring isn't biased by intent. It's biased by architecture. Training on historical spend data means the algorithm inherits every advantage the incumbent accrued over the prior decade.
The exclusion is also self-reinforcing. Suppliers who score lower receive fewer contracts, which means fewer transactions, which means lower confidence scores in the next cycle. The algorithm doesn't need any human intent to perpetuate the pattern—it perpetuates itself.
What the Four-Fifths Test Actually Finds

The EEOC's four-fifths rule, codified at 29 CFR 1607.4, provides that any group's selection rate must reach at least 80% of the highest-selected group's rate. Designed for employment screening, the same statistical test applies with equal force to supplier selection—and OFCCP's April 2024 AI guidance extended the adverse impact analysis framework explicitly to automated screening tools used by federal contractors.
The math is not abstract. If your S2P platform's AI advances 60% of non-diverse suppliers past the scoring threshold, it must advance at least 48% of MBE-certified suppliers. In volume-weighted scoring systems, MBE selection rates frequently come in near 22%. That's a disparity ratio of 0.37—well below the 0.80 threshold, and firmly in prima facie adverse impact territory.
What that means operationally: when a contracting officer reviews your subcontracting plan and flags a shortfall against your FAR Part 19 goals, their next question will be what your algorithm's selection rates look like across supplier categories. If your vendor doesn't provide that data—and none of the four major platforms currently do—you are answering that question with a gap where the documentation should be.
Why the Platform Won't Fix This for You

SAP Ariba's Joule Bid Analysis Agent, launched commercially in Q1 2026, is built to serve tens of thousands of customers with AI-powered award recommendations. Coupa's Navi Supplier Discovery Agent runs 100+ AI tools across a community of buyers whose network transaction data feeds the scoring signals. GEP SMART and Ivalua have followed similar agentic paths—full source-to-pay automation built on the same kind of volume-signal scoring.
The platform's economics don't support per-customer fairness configuration. Adding disparate impact constraints specific to your subcontracting percentage goals, your supplier categories, and your regulatory jurisdiction would mean maintaining a different model per customer. That's not how platform pricing works. The platform gives you speed and coverage. The fairness layer is yours to build and yours to document.
Supplier diversity discovery tools—Supplier.io with its 20M+ supplier database, Tealbook with 5M+ verified suppliers, Fairmarkit's AI-powered RFP matching—address a different part of the problem. They help build the pool of qualified diverse suppliers who can enter a sourcing event. They don't audit whether your scoring algorithm treats those suppliers equitably once they're in it. The pool is not the problem. What happens after the sourcing event opens is.
The Regulatory Pinch

Federal contractors are currently operating under two simultaneously binding requirements that appear contradictory.
EO 14319 (July 2025) prohibits federal agencies from procuring AI with "ideological biases or social agendas," language the order explicitly ties to DEI design. FAR Part 19 simultaneously requires prime contractors to maintain subcontracting plans with specific percentage goals for small business, veteran-owned, service-disabled veteran-owned, HUBZone, small disadvantaged, and women-owned subcontractors—goals enforced through contracting officer reviews and contract award decisions. Both are active.
The way through this apparent contradiction is statistical. EO 14319 prohibits DEI programs in AI. It doesn't prohibit—and can't coherently prohibit—mathematical proof that your algorithm selects suppliers equitably across categories. A four-fifths rule analysis applied to your AI's selection rates produces documentation that satisfies FAR Part 19 oversight without invoking the DEI framing the executive order targets. Illinois already allocates up to 20% of technical evaluation points to supplier diversity; demonstrating algorithmic equity is increasingly the mechanism those points turn on.
The GSA's draft GSAR 552.239-7001 AI clause (March 2026), currently in comment period, adds disclosure and use-rights requirements that move in the same direction. The compliance burden is increasing, not stabilizing.
What Auditable Fairness Looks Like in Practice

Veriprajna's AI Procurement Fairness & Supplier Diversity Compliance work is vendor-agnostic: we connect to SAP Ariba, Coupa, GEP, or Ivalua through their APIs and testing interfaces, pull the supplier scoring outputs from representative sourcing events, and run the four-fifths analysis across supplier categories. The output is a disparate impact report that documents actual selection-rate ratios for each category against the 0.80 threshold—the kind of documentation that answers the contracting officer's question directly.
Where the analysis finds disparity below that threshold, we trace the scoring factors driving it. In most configurations it's confidence weighting in delivery or quality metrics, but the specific locus varies by platform and sourcing category. The remediation isn't removing the AI—it's adding fairness constraints calibrated to your subcontracting goals before the AI recommends an award.
The 49% of procurement teams still running pilots rather than deploying—as documented in ProcureAbility's 2026 CPO Report—are largely stuck there because they can't defend moving to full deployment without a credible answer to the compliance question. That answer requires data the vendor doesn't provide by default and documentation that doesn't generate itself. The organizations moving from 49% to the 4% who have reached meaningful deployment aren't doing it with better AI. They're doing it with better audit trails.
If your team is working through what this looks like under your specific contracting structure—FAR Part 19 goals, OFCCP oversight, or state-level supplier diversity mandates—we'd be interested in comparing what we're seeing. The disparity ratios across platform configurations are consistent enough to be worth mapping at an industry level, and organizations willing to share their scoring data help make that picture more precise.