E-commerce AI accuracy | Vouchmark

Fix the HDR claim. Keep the 120Hz answer.

One shopping answer can contain a correct detail and a wrong specification. Vouchmark demonstrates separate catalog decisions for each claim: a synthetic TV answer changes HDR10+ to HDR10 while retaining its supported refresh rate.

32/32

True-labeled claims preserved

Fixed labeled synthetic evaluation

60/60

Claim verdicts match fixture labels

Already-decomposed synthetic claim cases

47/47

Pass/correct records carry references

Fixture reference presence, not authentication

These results test local decision rules on authored fixtures. They do not measure claim extraction, unseen shopper traffic or production performance.

An answer-level yes can hide a claim-level mistake

The synthetic Vega TV draft says it supports HDR10+ and a 120Hz refresh rate. The local catalog supports 120Hz, but names HDR10. Accepting the whole draft leaves the wrong standard in place; rejecting it discards a supported detail.

The review question is more precise: which assertion matches the record, which contradicts it, and which cannot be established? That distinction gives catalog and product teams a useful basis for deciding what an assistant may display.

How the claim decision works

Vouchmark wraps an assistant draft. The prepared scenarios supply atomic claims; a separate free-form endpoint supports optional model extraction or a limited pattern fallback. The video uses prepared scenarios and does not establish extraction accuracy.

  1. 1. Identify the assertion and catalog attribute

    The television input supplies one HDR claim and one refresh-rate claim. The local JSON catalog holds attribute values, confidence labels and authored reference strings. No external source is fetched.

  2. 2. Compare with explicit rules

    The deterministic gate recognizes HDR10 and HDR10+ as distinct standards. It corrects the mismatch and passes the refresh-rate match. Its normalization, curated matching and attribute rules have finite coverage.

  3. 3. Compose the answer and retain the decision

    Claim-level serving text forms the displayed answer, unless a configured safety rule overrides it. A JSON or printable HTML receipt records what was asserted, what the catalog said and why the decision changed.

Unknown attributes yield abstention; missing products or attributes yield outside-coverage. Inferred catalog values can pass the same matching path as verified values, so the receipt confidence label matters. A separate deterministic completeness rule catches selected configured implications; it does not inspect every possible meaning in prose.

Follow the answer from draft to catalog decision

All products, records, drafts and evaluation cases shown here are synthetic. The prepared scenarios supply their atomic claims. UI labels such as Verified describe catalog matching in this demonstration; they do not establish independent external verification.

Worked example: one TV answer, two different decisions

The shopper asks whether the synthetic Vega television supports HDR10+ and 120Hz. The prepared draft answers both questions confidently. The catalog supports the refresh rate but records a different HDR standard, so the answer needs a selective correction.

Prepared assistant draft

Yes, the Vega 65" OLED supports HDR10+ and a 120Hz refresh rate.

Synthetic Vega TV draft beside the displayed HDR10 correction and preserved 120Hz claim.
The synthetic Vega answer keeps 120Hz and corrects HDR10+ to HDR10. Store prices, reviews and controls are illustrative. Open the image for full-size inspection.

1. Compare each assertion with its catalog attribute

The HDR assertion and refresh-rate assertion arrive as separate supplied claims. Each is compared with its own product attribute, rather than giving the entire sentence one approval label. The catalog values below have verified confidence labels in the fixture; those labels are authored metadata.

Claim decisions in the synthetic television scenario
AttributeDraft assertionCatalog valueDecisionEffect on the answer
HDR standardHDR10+HDR10CorrectReplace the contradicted standard
Refresh rate120Hz120HzPassKeep the supported detail
HDR claim dialog showing HDR10+ asserted, HDR10 in the synthetic catalog and a corrected verdict.
The comparison exposes the asserted value, catalog value and authored fixture reference. The reference is not a retrieved manufacturer document. Open the image for full-size inspection.

2. Compose the answer from those individual decisions

The gate produces the following displayed answer. The HDR correction names both the recorded value and the rejected assertion; the refresh-rate statement survives. Matching the catalog is the demonstrated check, while the quality of that catalog remains a separate responsibility.

Displayed answer

Vega 65" OLED TV: HDR is HDR10, not HDR10+ (corrected). Vega 65" OLED TV: Refresh Rate = 120Hz (verified).

3. Inspect what changed in the receipt

The JSON and printable HTML receipt retain the original draft, displayed answer and two decision rows, including asserted values, catalog values, confidence labels and fixture references. A reviewer can check that the correction did not silently discard the supported refresh rate. The timestamp and hash-derived identifier identify this generated record; they do not make it signed, immutable or durably archived.

Generated Vouchmark television receipt with the original draft, served answer and two claim-decision rows.
The actual generated receipt preserves both decisions and their fixture references. Its timestamp and identifier do not establish a signature or immutable custody. Open the image for full-size inspection.

An unknown rating needs uncertainty, not a replacement value

The synthetic Summit Ridge jacket draft promises full waterproofing, while its waterproof-rating attribute is unknown. The gate abstains on that promise. It does not conclude that the jacket is not waterproof or invent a rating. A catalog team would need suitable product evidence before making the stronger claim.

Synthetic Summit Ridge jacket answer declining to confirm fully waterproof because the catalog rating is unknown.
A separate synthetic jacket case abstains on an unknown waterproof rating. The displayed 0/0 provenance percentage is not coverage evidence. Open the image for full-size inspection.

A configured implication can reveal an omitted claim

The synthetic Nimbus laptop draft calls it a fantastic gaming laptop and states that it has 16GB RAM. Its supplied claim list contains only RAM. A deterministic completeness rule maps the gaming language to a dedicated-GPU assertion, then compares that added assertion with the catalog's integrated Intel Arc graphics. RAM passes; the configured GPU implication is corrected. This illustrates one finite interpretation rule, not complete semantic extraction or a judgment about every game.

Synthetic Nimbus laptop showing the supported 16GB RAM claim and a corrected dedicated-GPU implication.
The configured completeness rule adds a dedicated-GPU assertion from gaming language. The synthetic catalog lists integrated Intel Arc graphics; 16GB RAM is preserved. This rule is not a universal gaming-suitability test. Open the image for full-size inspection.

A product record cannot establish a drug-interaction answer

In the synthetic MagCalm scenario, the query asks about a blood thinner. A configured regex screen overrides the product answer with a refusal and text naming a licensed pharmacist as the proposed review destination. The drug-interactions field is also unknown. This demonstrates the decision to withhold the product answer; no professional handoff is executed and no medical recommendation is validated.

Synthetic MagCalm drug-interaction query with the product answer refused and a simulated pharmacist destination.
The configured regex screen refuses this synthetic drug-interaction query. The connecting and escalation wording is simulated: no pharmacist is contacted. Its 0/0 provenance display is not coverage evidence. Open the image for full-size inspection.

Read the benchmark with its test units attached

The fixed evaluation sends 60 already-decomposed claim fixtures directly to the gate and seven separate query fixtures to the screen. It tests local rules against authored expected labels. It does not test model extraction, unseen shopper traffic, source authentication or production performance.

Fixed labeled synthetic evaluation
MeasurementResultWhat was counted
Claim verdict matches60/6032 pass, 15 correct, 10 abstain and 3 outside-coverage fixture labels
True-claim preservation32/32True-labeled claims that remain pass
Seeded false-claim interception28/28Non-pass-labeled claims given correct, abstain or outside-coverage; corrected values can still be served
Reference presence47/47Pass/correct records carrying a nonempty fixture reference string
Query-screen matches7/7Five configured triggers and two ordinary product queries
Completed fixed synthetic evaluation showing 60 claim verdict matches, 47 referenced served records and 32 preserved true claims.
The completed ledger combines 60 claim cases with seven query-screen cases. The Sony-named rows are synthetic records, not validated manufacturer specifications. Open the image for full-size inspection.

The ledger combines different test units. Its total of 67 does not mean 67 product claims or 67 specialist escalations. Reference-string presence explains where a fixture decision points; it does not authenticate an external manufacturer source.

What each review approach establishes

These are design choices for reviewing an assistant answer, rather than claims about competing products. Retrieving a relevant catalog record and checking a specific assertion against it are different steps.

Review approachWhat the TV case revealsRemaining question
Accept the entire draftRetains both HDR10+ and 120HzWhich assertion contradicts the record?
Reject the entire draftRemoves the supported refresh-rate detail tooWhich useful claim could be preserved?
Apply claim-level catalog checksCorrects HDR and preserves the refresh rateWere the relevant claims extracted, and is the catalog sound?

What this demo does not do

This is a demonstration of local rules on synthetic inputs, not a store deployment. There are no implemented store, PIM, checkout, payment or specialist connectors, and the retail controls do not execute transactions. Configured medical screening shows a refusal and proposed pharmacist destination; it does not contact a professional.

The demo does not guarantee complete claim extraction, correct catalog data, exhaustive implication checks or performance on unseen answers. Receipts are inspectable exports, not signed records or a durable audit archive. A deployment needs independent validation of those boundaries.

Questions from a product and catalog review

Can this fix a wrong product detail without blocking the whole answer?

Vouchmark demonstrates separate decisions for individual claims. In the synthetic Vega TV example, it corrects HDR10+ to the catalog value HDR10 and preserves the supported 120Hz refresh-rate statement.

What happens when our catalog is missing the information?

An attribute marked unknown produces abstention; a missing product or attribute produces outside-coverage. In the synthetic rain-jacket example, the displayed answer declines to confirm full waterproofing because the rating is unknown.

Is this just another model checking the first model?

The catalog verdicts and current completeness rule are deterministic Python rules. Only free-form claim extraction can optionally use a model; the prepared examples supply their claims and do not measure model extraction.

Does it connect to our Shopify store or PIM?

The demo reads a local synthetic catalog and policy records. Store, PIM, checkout and specialist integrations are not implemented; a deployment would require separate integration and validation work.

What does the accuracy result actually test?

All 60 already-decomposed claim cases match their authored expected verdicts in the fixed labeled synthetic evaluation. The seven query-screen cases are a separate test unit; neither result measures extraction completeness, unseen shopper traffic or production performance.

Can we see why it changed an answer?

The JSON and printable HTML audit receipts retain the input draft, displayed answer and per-claim assertion, catalog value, decision, confidence label and reference. References are authored fixture metadata; the receipt is not a digital signature, immutable archive or durable audit store.

Technical Research

Explore related research for broader context on this demonstration.

Define what your shopping assistant can support

Start with the catalog, the answer and the decision boundary.

We can discuss the checks, missing-data rules and evaluation your product answers need. The demonstration is a starting point for that design conversation.

Assess the answer boundary

  • ✓ Inspect catalog attributes and unknowns
  • ✓ Map checkable claims in draft answers
  • ✓ Define correction and abstention rules
  • ✓ Separate gate tests from extraction tests

Plan an implementation

  • ✓ Design catalog and policy integration
  • ✓ Specify claim-comparison coverage
  • ✓ Define receipts and retention needs
  • ✓ Plan validation on representative traffic