An AI shopping answer can get the refresh rate right and the display standard wrong in the same sentence. Approving that sentence preserves the mistake. Rejecting it discards a useful specification.
I prefer a verification boundary that makes those decisions separately. For an e-commerce team, the practical question is which claim to preserve, which to replace, and which the available product data cannot establish.
Our Vouchmark demo makes that distinction inspectable using synthetic products, prepared assistant drafts and a synthetic catalog. Its store interface is illustrative; it does not operate a checkout or publish answers to a real retailer.
A correction should preserve what the catalog supports
The synthetic Vega television draft says the product supports HDR10+ and a 120Hz refresh rate. Its catalog specifies HDR10 and 120Hz. The Python gate corrects the standard while retaining the refresh rate in the displayed answer.
The synthetic television case separates a contradicted display standard from a supported refresh-rate claim. Listing details and ratings are demo fixtures.
That is a small example with a consequential design choice. A blanket rejection would give the shopper less information than the catalog can support. A blanket approval would let an incorrect standard borrow credibility from the correct detail beside it.
The catalog team can review a narrower decision: does the replacement value match the product record, and did the useful claim survive? The engineering team can check whether the displayed answer actually reflects those two dispositions. A single approved/rejected label for the whole draft would hide both questions.
Vouchmark also exposes the asserted value, catalog value and reference behind the correction. Here, the reference is authored fixture metadata, rather than an authenticated manufacturer document. That distinction matters: explaining a comparison does not establish that the catalog itself is correct.
The comparison explains why the gate changed the claim. Its reference identifies the synthetic catalog record; it does not authenticate an external source.
Missing evidence needs a different next step
The synthetic rain-jacket case has a different problem. The draft calls the jacket fully waterproof, while its waterproof-rating field is unknown. The gate abstains from that promise. It does not replace it with a claim that the jacket is not waterproof.
I want that distinction preserved because it changes the work a catalog team should do. A contradiction calls for checking the asserted value against the recorded specification. An unknown rating calls for obtaining suitable product evidence before making the stronger promise. Treating both as a generic error would encourage a correction where the evidence only supports uncertainty.
For a retailer evaluating this design, I would ask the team to trace both cases through its proposed answer policy. Where would the correct refresh rate remain visible? What wording would replace the wrong display standard? What would the shopper see while the waterproof rating remains unresolved? Those are concrete output decisions that an overall answer score cannot supply.
A reliable gate still needs claims to reach it
The boundary has a second dependency: the system must identify the claims it needs to check. Vouchmark's prepared scenarios supply their atomic claims, so the television result demonstrates catalog decisions on those supplied assertions. It does not demonstrate complete extraction from arbitrary shopping prose.
The fixed labeled synthetic evaluation has the same boundary. All 60 already-decomposed claim cases match their expected verdicts. That result tests the gate against local fixture labels; it does not measure model extraction, unseen shopper traffic or the correctness of external catalog sources.
I would therefore keep two acceptance questions separate when evaluating a verification layer. First, does it make the intended decision when given a claim and the relevant product record? Second, does it consistently find the material assertions in the answers the retailer actually plans to serve? Passing the first test cannot answer the second. An omitted assertion can escape scrutiny even when every submitted claim is handled correctly.
The Vouchmark explainer shows the worked comparison and its audit receipt. The receipt keeps the original draft, displayed answer and claim-level decisions together for inspection; it is not a signed or durable audit archive.
For me, the useful standard is a verification layer that preserves the supported detail, makes uncertainty explicit and lets a reviewer see the basis for each change. Evaluating that standard means examining both the decisions the gate makes and the claims that never reach it.