MeritLens · Procurement Fairness Firewall
A better delivery rate can still leave a supplier below the shortlist floor. We show how history and revenue assumptions shape that result, then preserve the explanation alongside a simulated award's review decision.
70.9 → 81.0
Bridgepoint reaches preview eligibility
Synthetic fasteners fixture; original → alternative score
14.0 of 21.0
Gap points attributed to proxy assumptions
Synthetic Apex versus Bridgepoint formula comparison
BLOCK
Saved simulated award decision
Human review required; no completed review represented
Synthetic suppliers and simulated awards. The video and evidence explain configured demo behavior; this page is an explainer.
In the synthetic Industrial Fasteners case, Bridgepoint Components delivers on time at 98.1%, compared with Apex Fastener Industries at 97.2%. Yet Bridgepoint scores 70.9 against the 80-point shortlist floor, while Apex leads at 91.9.
The formula gives a supplier with limited transactions and audits less confidence in its delivery and quality rates. It also uses revenue for financial scoring. Excluding certification labels from a score does not remove those history and size assumptions.
A procurement reviewer needs to see which assumptions change eligibility, which differences remain under an alternative formula, and why the award policy holds or allows the recommendation. Better delivery alone does not settle the supplier's financial risk or entitlement to an award.
The simulated scorer gives delivery, quality, financial and price equal 25% weights. Its proxy-neutral alternative uses raw delivery and quality rates and actual financial health, keeping price and weights unchanged. The Oaxaca-Blinder-style decomposition splits the original gap into the gap remaining under that alternative and a residual attributed to the changed assumptions.
The configured four-fifths diagnostic compares each reported class's selection rate with the non-diverse baseline, using a 0.80 ratio floor. It pools synthetic supplier-event rows within the category, including repeated suppliers. Small groups cannot drive the verdict; unavailable category evidence yields ABSTAIN.
BLOCK requires all three conditions: category adverse impact, a non-diverse original top scorer, and a diverse challenger below the original floor that reaches it in the alternative formula. ABSTAIN also holds the simulated award for review. Other cases return ALLOW under this policy. Code saves the gate and audit entry before optional advisory factor review; advisory prose cannot change the decision or numerical evidence.
The alternative formula is an explicit assumption test, not independent causal proof. The 0.80 floor is a configured demo diagnostic, not a procurement-law mandate.
Follow one synthetic sourcing event from unevaluated inputs to an explained hold. The contrasting cases then show why changing shortlist eligibility, passing a category diagnostic and permitting an award are separate outcomes. All supplier names, certifications, histories and awards shown here belong to the local demonstration.
The workspace initially shows supplier inputs with scores and shortlist decisions pending. Inspecting Bridgepoint exposes delivery, quality, transaction history, audit history and financial inputs before evaluation. That distinction matters: a supplier label or an apparently strong delivery percentage cannot explain the score without the formula that consumes it.

| Synthetic input | Apex | Bridgepoint |
|---|---|---|
| On-time delivery | 97.2% | 98.1% |
| Quality-pass rate | 96.5% | 97.2% |
| Transactions | 4,200 | 180 |
| Audits | 140 | 9 |
| Revenue | USD 820 million | USD 60 million |
The simulated scorer weights delivery, quality, financial and price factors at 25% each. Transaction and audit history affect how much confidence it places in the raw delivery and quality rates; revenue affects the financial score. Those choices can penalize a smaller supplier even though certification labels are excluded from scoring. A buyer therefore needs to assess the justification for those assumptions, rather than infer neutrality from omitted labels.
After evaluation, Apex is the original top scorer at 91.9. Bridgepoint scores 70.9 and Delta scores 70.5, below the configured shortlist floor of 80. Bridgepoint has the better delivery rate, but that single measure does not establish superiority on financial health, price or every other legitimate criterion.

BLOCK requires all three configured conditions: a category adverse-impact finding, a non-diverse original top scorer, and a diverse challenger that is below the original floor but reaches it in the alternative formula. The fasteners event meets that combination. The decision is saved before advisory factor review begins, preserving the numerical basis and the authority of the hold.
The category diagnostic pools supplier-event rows, including repeated suppliers. Minority-owned business rows have 7 selections out of 26 considered, compared with 25 out of 37 for the non-diverse baseline. The selection-rate ratio is about 0.40, below the configured 0.80 diagnostic floor. The category history metadata of 312 is a separate sample guard; it is not the denominator of that comparison.

| Changed assumption | Proxy contribution to gap |
|---|---|
| Delivery confidence from transaction history | 3.01 points |
| Quality confidence from audit history | 2.50 points |
| Revenue-based financial scoring | 8.47 points |
| Price scoring, unchanged | 0.00 points |
The Oaxaca-Blinder-style decomposition compares explicit formulas. The alternative uses raw delivery and quality rates and actual financial health, keeping price and weights unchanged. About two-thirds of the original gap is attributed to the assumptions changed in that comparison; the remaining gap still favors Apex. This residual is conditional on the chosen alternative, rather than independent causal proof of discrimination.
| Supplier | Original score | Alternative score | Preview result |
|---|---|---|---|
| Apex Fastener Industries | 91.9 | 88.0 | Remains top scorer |
| Bridgepoint Components | 70.9 | 81.0 | Becomes eligible |
| Delta Fasteners | 70.5 | 82.9 | Becomes eligible |

That preview gives procurement reviewers a focused question: would these suppliers have been excluded under scoring assumptions the organization can defend? It does not choose Bridgepoint as the winner, execute an award, rescore the recorded category assessment or replace the original scorer. A human reviewer would still need to consider the procurement context and decide how the identified assumptions should be handled.
In the Packaging Materials example, the original recommendation is Beacon at 91.7, the category diagnostic is PASS and the configured gate returns ALLOW. Yet Ironclad, a synthetic HUBZone supplier, moves from 72.4 to 81.4 under the alternative formula. A policy permit can therefore coexist with a proxy-sensitive shortlist exclusion.

The overlapping diverse aggregate has 8 selections out of 12 considered rows, compared with 18 out of 38 for the non-diverse baseline. Its ratio is about 1.41. Individual minority-owned, HUBZone and 8(a) groups each have only four considered rows and remain unscored. Category PASS therefore does not establish that every class was assessed, and ALLOW does not establish the absence of proxies or disparity.
The Facilities and MRO event has declared category history of 140, below the configured minimum of 200. Its category assessment and gate both return ABSTAIN, so the simulated autonomous award is held for review. Scores and a decomposition can still be calculated; neither turns an unavailable category assessment into clearance.

Other guards cover a reference baseline with fewer than five considered rows or no selections. Individual groups below five are unscored, while groups with five through nine rows are informational and cannot drive the category verdict. The minimum history of 200 and these group guards are configured boundaries, not a guarantee that a real procurement dataset is statistically adequate.
The fixed suite contains ten authored raw-input cases with expected outcomes: three BLOCK, four ALLOW and three ABSTAIN. The observed outcomes match all ten expected labels. Controls cover history, audit and revenue assumptions, established certified suppliers, price-only differences, performance-based exclusion, sparse reference data and a reference baseline with no selections.

One boundary case permits assessment at history exactly 200 and returns ALLOW; another at 199 returns ABSTAIN. These controls help inspect whether the implementation follows its stated policy at consequential boundaries. They do not establish procurement-field accuracy, calibration or independent validation. The benchmark runs without a model call and does not append sourcing audit entries.
The synthetic portfolio contains 52 events across six categories and 342 supplier-event rows. Those rows are repeated participation records, not 342 unique suppliers. Fifteen events are marked for autonomous awards; recomputing the policy yields eight BLOCK, one ABSTAIN and six ALLOW among them. Nine of the fifteen would therefore be held for review under this configuration.

At category level, the scan reports four ADVERSE_IMPACT findings, one PASS and one ABSTAIN. The scan helps identify where a reviewer would inspect the evidence next. It is separate from the ten-case control suite: one summarizes a seeded portfolio, while the other checks expected implementation behavior. Neither supplies evidence about actual vendor populations.
After the evaluation and advisory factor review finish, Download JSON and Print report include the saved decision, audit entry, category counts, decomposition and factor review. Optional Proxy Classifier and Adversarial Challenger advice adds labels and rationales after the gate is saved. Numeric evidence remains grounded in code; advisory prose cannot authorize the award. Complete saved advisory responses can be reused for unchanged inputs, and unavailable providers fall back to deterministic review.
For a production assessment, the next questions concern which history measures are defensible for the contract, whether category comparisons represent the intended supplier population, and who has authority to resolve a hold. The demo makes those questions inspectable. Its local JSON audit persistence is not immutable third-party custody, and legal-named export fields are configured mappings rather than legal certification or independent assurance.
A score, an assumption test and a policy decision answer different procurement questions. We keep those questions separate so a review does not inherit more certainty than the evidence provides.
| Evidence | Useful for | Limit |
|---|---|---|
| Original score | Understanding the simulated scorer's recommendation | Certification labels are excluded from scoring; history and revenue assumptions remain |
| Proxy-neutral preview | Testing which suppliers become eligible under an explicit alternative | Does not execute an award or independently prove discrimination |
| Category diagnostic and gate | Applying the configured BLOCK, ALLOW or ABSTAIN policy | ALLOW is a policy outcome, not proof that every class passed or no proxy remains |
| Saved evidence export | Reviewing the recorded decision with its numerical basis | Local JSON persistence and template mappings, not immutable custody or legal certification |
MeritLens uses synthetic suppliers and simulated awards, with no live procurement-platform connectors or real contract execution. Its exports include legal-named configured template mappings, not legal certification or independent assurance. The counterfactual does not overwrite the recommendation or saved gate, and the fixed synthetic controls do not establish procurement-field accuracy. Production data access, operational persistence and human-review authority need their own design and validation.
In the synthetic fasteners example, Bridgepoint has a 98.1% raw delivery rate versus Apex's 97.2%, yet its original score is 70.9 against an 80-point shortlist floor. The simulated formula weights delivery and quality confidence by transaction and audit history, and scores the financial factor from revenue. Better delivery alone does not establish superiority on every legitimate criterion.
Remove size proxy changes a comparison preview, not the recorded recommendation or gate decision. Bridgepoint reaches 81.0 and Delta reaches 82.9, making both shortlist-eligible under that alternative formula. Apex remains the top scorer, and the synthetic award stays blocked for human review.
BLOCK requires a category adverse-impact finding, an original top scorer labeled non-diverse, and a diverse challenger that crosses the shortlist floor only under the counterfactual. An unavailable category assessment produces ABSTAIN and also holds the simulated award for review. All other cases return ALLOW under this configured policy, which does not establish an absence of proxies or disparity.
The demo abstains when declared category history is below 200, or when the non-diverse reference baseline has fewer than five considered supplier-event rows or no selections. Individual groups below five are unscored; groups with five through nine rows are informational and cannot drive the category verdict. These are configured sample guards, not a guarantee of statistical adequacy for a real procurement dataset.
This demo consumes a local synthetic supplier dataset and simulates autonomous awards. It has no live procurement-platform connector or real contract execution. A production project would need to establish data access, scoring semantics, persistence and review authority for the target environment.
The gap explanation compares explicit formulas under chosen assumptions; it does not independently establish real-world discrimination. Legal-named export fields are configured template mappings, not legal certification or independent assurance. The blocked fasteners export retains the original adverse-impact finding, and an abstention leaves that assessment unassessed.
Code saves the policy decision and audit entry before optional factor advice begins. The advisory reviewers provide labels, explanations and challenges, while validated numerical evidence remains grounded in the deterministic calculations. Saved complete advice can be reused for unchanged inputs, and unavailable providers fall back to deterministic review; neither path gives advisory prose authority to change the gate.
Explore related research for broader context on this demonstration.
Discuss the evidence your procurement checkpoint needs.
We help teams assess supplier-scoring assumptions and design a review workflow around their data, policy and decision authority.