A shopping answer can get one specification right and another wrong in the same sentence. If the review decision applies only to the whole answer, it hides a useful distinction: which detail needs correction, and which detail should stay?
For e-commerce and product leaders, that distinction belongs in the acceptance criteria for an AI answer. A check that rejects an unsupported promise should also demonstrate that it preserves the supported information beside it. Otherwise, the shopper loses an answer the catalog could have supplied.
We use Vouchmark to demonstrate this decision at the level of individual claims. Its television example is synthetic, including the product, catalog values and prepared draft. It gives us a concrete way to examine what a correction should do before making claims about a production shopping assistant.
One sentence, two decisions
The prepared draft says a fictional Vega 65-inch OLED TV supports HDR10+ and a 120Hz refresh rate. The synthetic catalog lists HDR10 and 120Hz.
The refresh-rate assertion matches. The HDR assertion does not. The extra plus sign names a different standard in the demo's comparison rules, so accepting the whole sentence would retain a contradiction. Rejecting the whole sentence would discard the supported refresh rate.
Vouchmark changes the HDR claim to HDR10 and keeps 120Hz in the displayed answer. That is the central design choice: correction should be selective enough to preserve the useful part of an answer.
In this synthetic listing, the HDR claim is corrected while the refresh rate stays. The store controls and listing labels illustrate the decision; no store publishing or transaction takes place.
The product name and numbers make the example inspectable, but the acceptance criterion travels beyond this television: a reviewer should be able to identify the assertion that changed, the value that replaced it, and the supported assertion that survived. An answer-level approval badge alone does not expose those relationships.
This also changes how we would evaluate a proposed verification layer. Seed a draft with a contradiction beside a catalog-supported detail. Then inspect the resulting text, rather than checking only whether the system raised an alert. A useful correction must appear in the answer, and preservation must be visible there too.
The catalog value must be open to inspection
The next question is why the system made that correction. In Vouchmark, a claim detail pairs the asserted HDR10+ value with the catalog's HDR10 value and a corrected verdict. The Python gate applies finite comparison rules outside the optional model-extraction layer.
The detail exposes the disagreement behind the correction. Its confidence label and spec-sheet reference are authored fixture metadata, not independent manufacturer verification.
That boundary matters. A catalog check establishes a relationship between an assertion and the available catalog record. It does not establish that the record itself is correct. In this demo, the reference is a readable string supplied with the synthetic attribute, rather than a retrieved manufacturer document.
For a catalog team assessing this design, those are separate review questions. Does the answer match the record? Is the record an acceptable basis for the answer? The first question has a demonstrated rule decision here. The second requires scrutiny of the catalog's evidence and maintenance outside this fixture.
The word “verified” in an interface also deserves inspection. Vouchmark allows matching attributes marked inferred through the same path as those marked verified. Its serving text can call either match verified, while the receipt retains the underlying confidence label. We would therefore assess the stored evidence classification, rather than assuming that the display label means every passing value was independently confirmed.
Keep the original claim beside the correction
A corrected sentence explains the result. A receipt helps someone examine how that result relates to the input.
Vouchmark's readable audit receipt keeps the original draft, displayed answer and per-claim decisions together. In the television example, the HDR row records the contradiction and replacement value. The refresh-rate row records the supported match. Both rows retain their fixture references.
The receipt preserves the before-and-after relationship for both claims. Its references remain synthetic metadata; the timestamp and receipt ID do not provide a signature or durable audit archive.
For an internal review, that arrangement allows a focused conversation: was the assertion captured correctly, did the comparator use the intended attribute, and did the displayed wording reflect the decision? A reviewer can disagree with a rule or a source value while still seeing precisely what the system did.
The receipt is an inspection aid. The demo returns JSON and printable HTML; it does not implement a persistent append-only record store. That distinction should remain explicit when a team decides what additional operational controls it needs.
Preservation is one test, extraction is another
The television example uses supplied claims. It demonstrates the catalog decision once those claims are available; it does not show that a model found every assertion in arbitrary prose.
The same separation applies to Vouchmark's fixed labeled synthetic evaluation. All 32 claims labeled as expected passes remain passes. Those cases include ordinary true claims and curated paraphrases. The harness calls the claim verifier directly, so this preservation result does not measure extraction completeness or performance on real shopper traffic.
We see two useful acceptance checks here: inspect whether a known contradiction is corrected without losing a supported detail, and separately establish whether the system finds the claims that need checking. A strong result on the first cannot substitute for evidence on the second.
The Vouchmark breakdown includes the video and worked examples. When assessing this kind of layer, ask for the original assertion, the catalog value and the final wording together. That is enough to inspect the demonstrated correction, and specific enough to reveal which questions still require evidence.