A procurement team can exclude protected labels from its scoring formula and still leave a consequential question unanswered: why did a supplier miss the shortlist? Transaction history, audit volume and revenue can influence the result even when the formula never reads a supplier's diversity status.
For a leader approving an autonomous award, the ranking alone is incomplete evidence. The useful next step is to inspect how those assumptions affect eligibility, then decide whether the proposed award needs human review. We demonstrate that distinction in MeritLens, our Procurement Fairness Firewall, using synthetic suppliers and simulated awards.
Stronger delivery, lower score
In the synthetic Industrial Fasteners example, Bridgepoint Components has a 98.1% on-time delivery rate. Apex Fastener Industries has 97.2%. Bridgepoint also has a slightly stronger quality-pass rate. Yet the demo's simulated platform scorer gives Bridgepoint 70.9 out of 100, below the shortlist floor of 80. Apex scores 91.9 and becomes the original recommendation.
That result does not establish that Bridgepoint deserves the contract. Delivery is one consideration, and the financial-health input still favors Apex. It does establish a reason to inspect what the score measures before treating exclusion as a complete assessment of the supplier.
The original formula gives delivery, quality, financial and price factors equal weight. Its delivery and quality scores depend on how much history supports the raw rates. Fewer transactions or audits pull those rates closer to an 80% prior. Its financial score uses revenue rather than the separate financial-health input.
Bridgepoint has 180 transactions and nine audits, compared with Apex's 4,200 transactions and 140 audits. A stronger raw rate therefore receives less credit in the original score. Revenue adds another size-related assumption. Keeping diversity labels outside the formula does not remove those mechanisms.
Make the alternative explicit
MeritLens compares that result with a specified alternative: give every supplier full credit for its raw delivery and quality rates, score financial health directly, and keep price and factor weights unchanged. The comparison asks how much of the gap depends on the original history and revenue assumptions.
In this fixture, the 21.0-point gap between Apex and Bridgepoint contains 14.0 points attributed to those proxy assumptions, about two-thirds of the gap. Another 7.0 points remain under the alternative formula. The factor split shows why an explanation matters: history weighting changes the delivery and quality comparison, while actual financial health continues to favor Apex.
The synthetic fasteners comparison separates the score gap under the alternative formula from the residual attributed to history, audit volume and revenue assumptions.
The residual is conditional on that alternative. It is a formula comparison, not causal proof of discrimination. A procurement reviewer can reasonably question whether raw rates from a smaller history deserve full credit, or whether revenue captures a risk that financial health alone misses. The point of making the assumptions explicit is to give that disagreement something concrete to examine.
A new shortlist is not a new award
Under the proxy-neutral preview, Bridgepoint reaches 81.0 and enters the shortlist. Delta Fasteners also crosses the floor. Apex scores 88.0 and still ranks first. The changed result concerns who receives consideration, not who is entitled to win.
Bridgepoint and Delta become eligible in the preview. Apex remains first, and the saved simulated award still requires human review.
The preview does not overwrite the platform recommendation or the saved decision. This separation preserves a question that a convenient rescore could otherwise obscure: should the organization approve the proposed award under the original scoring assumptions?
For this event, the configured gate answers BLOCK. The category diagnostic finds a selection disparity, the original top scorer has the fixture's non-diverse label, and diverse challengers cross the shortlist floor only under the alternative formula. All three conditions are required by the policy.
The category evidence also has a defined scope. Across pooled synthetic supplier-event rows, suppliers with the minority-owned business label were selected in 7 of 26 cases, compared with 25 of 37 for the non-diverse baseline. Their selection-rate ratio is approximately 0.40, below the configured 0.80 floor. These counts include repeated suppliers across events; they are not a field study or a legal determination.
BLOCK holds the simulated award for human review. It does not show a completed review, a real contract interception or a legal certification. The gate combines evidence under a stated policy; it does not settle every procurement judgment.
Keep the explanation subordinate to the decision
The code saves the gate decision and audit entry before advisory factor review. Optional AI explanations can label and challenge the interpretation, but cannot change the scores or award policy result. In the paired video, previously completed model explanations are reused for unchanged evidence. Fresh scoring still computes the deterministic gate.
We think that order is useful for an award checkpoint. A fluent explanation should help a reviewer interrogate the evidence without becoming a second, less visible source of decision authority. The saved evidence package keeps the decision, category counts, gap comparison and factor review together in JSON and a printable report. This is a local demonstration, with no live procurement connector or real award execution.
The MeritLens breakdown shows the example in more detail. For procurement leaders evaluating an AI scoring workflow, our recommendation is to require an explanation of changed eligibility before approving an award: which assumptions excluded the supplier, what the alternative preserves, and who has authority to resolve the remaining risk.