E-commerce AI accuracy | Vouchmark
One shopping answer can contain a correct detail and a wrong specification. Vouchmark demonstrates separate catalog decisions for each claim: a synthetic TV answer changes HDR10+ to HDR10 while retaining its supported refresh rate.
32/32
True-labeled claims preserved
Fixed labeled synthetic evaluation
60/60
Claim verdicts match fixture labels
Already-decomposed synthetic claim cases
47/47
Pass/correct records carry references
Fixture reference presence, not authentication
These results test local decision rules on authored fixtures. They do not measure claim extraction, unseen shopper traffic or production performance.
The synthetic Vega TV draft says it supports HDR10+ and a 120Hz refresh rate. The local catalog supports 120Hz, but names HDR10. Accepting the whole draft leaves the wrong standard in place; rejecting it discards a supported detail.
The review question is more precise: which assertion matches the record, which contradicts it, and which cannot be established? That distinction gives catalog and product teams a useful basis for deciding what an assistant may display.
Vouchmark wraps an assistant draft. The prepared scenarios supply atomic claims; a separate free-form endpoint supports optional model extraction or a limited pattern fallback. The video uses prepared scenarios and does not establish extraction accuracy.
The television input supplies one HDR claim and one refresh-rate claim. The local JSON catalog holds attribute values, confidence labels and authored reference strings. No external source is fetched.
The deterministic gate recognizes HDR10 and HDR10+ as distinct standards. It corrects the mismatch and passes the refresh-rate match. Its normalization, curated matching and attribute rules have finite coverage.
Claim-level serving text forms the displayed answer, unless a configured safety rule overrides it. A JSON or printable HTML receipt records what was asserted, what the catalog said and why the decision changed.
Unknown attributes yield abstention; missing products or attributes yield outside-coverage. Inferred catalog values can pass the same matching path as verified values, so the receipt confidence label matters. A separate deterministic completeness rule catches selected configured implications; it does not inspect every possible meaning in prose.
All products, records, drafts and evaluation cases shown here are synthetic. The prepared scenarios supply their atomic claims. UI labels such as Verified describe catalog matching in this demonstration; they do not establish independent external verification.
The shopper asks whether the synthetic Vega television supports HDR10+ and 120Hz. The prepared draft answers both questions confidently. The catalog supports the refresh rate but records a different HDR standard, so the answer needs a selective correction.
Prepared assistant draft
Yes, the Vega 65" OLED supports HDR10+ and a 120Hz refresh rate.

The HDR assertion and refresh-rate assertion arrive as separate supplied claims. Each is compared with its own product attribute, rather than giving the entire sentence one approval label. The catalog values below have verified confidence labels in the fixture; those labels are authored metadata.
| Attribute | Draft assertion | Catalog value | Decision | Effect on the answer |
|---|---|---|---|---|
| HDR standard | HDR10+ | HDR10 | Correct | Replace the contradicted standard |
| Refresh rate | 120Hz | 120Hz | Pass | Keep the supported detail |

The gate produces the following displayed answer. The HDR correction names both the recorded value and the rejected assertion; the refresh-rate statement survives. Matching the catalog is the demonstrated check, while the quality of that catalog remains a separate responsibility.
Displayed answer
Vega 65" OLED TV: HDR is HDR10, not HDR10+ (corrected). Vega 65" OLED TV: Refresh Rate = 120Hz (verified).
The JSON and printable HTML receipt retain the original draft, displayed answer and two decision rows, including asserted values, catalog values, confidence labels and fixture references. A reviewer can check that the correction did not silently discard the supported refresh rate. The timestamp and hash-derived identifier identify this generated record; they do not make it signed, immutable or durably archived.

The synthetic Summit Ridge jacket draft promises full waterproofing, while its waterproof-rating attribute is unknown. The gate abstains on that promise. It does not conclude that the jacket is not waterproof or invent a rating. A catalog team would need suitable product evidence before making the stronger claim.

The synthetic Nimbus laptop draft calls it a fantastic gaming laptop and states that it has 16GB RAM. Its supplied claim list contains only RAM. A deterministic completeness rule maps the gaming language to a dedicated-GPU assertion, then compares that added assertion with the catalog's integrated Intel Arc graphics. RAM passes; the configured GPU implication is corrected. This illustrates one finite interpretation rule, not complete semantic extraction or a judgment about every game.

In the synthetic MagCalm scenario, the query asks about a blood thinner. A configured regex screen overrides the product answer with a refusal and text naming a licensed pharmacist as the proposed review destination. The drug-interactions field is also unknown. This demonstrates the decision to withhold the product answer; no professional handoff is executed and no medical recommendation is validated.

The fixed evaluation sends 60 already-decomposed claim fixtures directly to the gate and seven separate query fixtures to the screen. It tests local rules against authored expected labels. It does not test model extraction, unseen shopper traffic, source authentication or production performance.
| Measurement | Result | What was counted |
|---|---|---|
| Claim verdict matches | 60/60 | 32 pass, 15 correct, 10 abstain and 3 outside-coverage fixture labels |
| True-claim preservation | 32/32 | True-labeled claims that remain pass |
| Seeded false-claim interception | 28/28 | Non-pass-labeled claims given correct, abstain or outside-coverage; corrected values can still be served |
| Reference presence | 47/47 | Pass/correct records carrying a nonempty fixture reference string |
| Query-screen matches | 7/7 | Five configured triggers and two ordinary product queries |

The ledger combines different test units. Its total of 67 does not mean 67 product claims or 67 specialist escalations. Reference-string presence explains where a fixture decision points; it does not authenticate an external manufacturer source.
These are design choices for reviewing an assistant answer, rather than claims about competing products. Retrieving a relevant catalog record and checking a specific assertion against it are different steps.
| Review approach | What the TV case reveals | Remaining question |
|---|---|---|
| Accept the entire draft | Retains both HDR10+ and 120Hz | Which assertion contradicts the record? |
| Reject the entire draft | Removes the supported refresh-rate detail too | Which useful claim could be preserved? |
| Apply claim-level catalog checks | Corrects HDR and preserves the refresh rate | Were the relevant claims extracted, and is the catalog sound? |
This is a demonstration of local rules on synthetic inputs, not a store deployment. There are no implemented store, PIM, checkout, payment or specialist connectors, and the retail controls do not execute transactions. Configured medical screening shows a refusal and proposed pharmacist destination; it does not contact a professional.
The demo does not guarantee complete claim extraction, correct catalog data, exhaustive implication checks or performance on unseen answers. Receipts are inspectable exports, not signed records or a durable audit archive. A deployment needs independent validation of those boundaries.
Vouchmark demonstrates separate decisions for individual claims. In the synthetic Vega TV example, it corrects HDR10+ to the catalog value HDR10 and preserves the supported 120Hz refresh-rate statement.
An attribute marked unknown produces abstention; a missing product or attribute produces outside-coverage. In the synthetic rain-jacket example, the displayed answer declines to confirm full waterproofing because the rating is unknown.
The catalog verdicts and current completeness rule are deterministic Python rules. Only free-form claim extraction can optionally use a model; the prepared examples supply their claims and do not measure model extraction.
The demo reads a local synthetic catalog and policy records. Store, PIM, checkout and specialist integrations are not implemented; a deployment would require separate integration and validation work.
All 60 already-decomposed claim cases match their authored expected verdicts in the fixed labeled synthetic evaluation. The seven query-screen cases are a separate test unit; neither result measures extraction completeness, unseen shopper traffic or production performance.
The JSON and printable HTML audit receipts retain the input draft, displayed answer and per-claim assertion, catalog value, decision, confidence label and reference. References are authored fixture metadata; the receipt is not a digital signature, immutable archive or durable audit store.
Explore related research for broader context on this demonstration.
Start with the catalog, the answer and the decision boundary.
We can discuss the checks, missing-data rules and evaluation your product answers need. The demonstration is a starting point for that design conversation.