
When an AI Shopping Answer Is Only Partly Right
An AI shopping answer can get a product's refresh rate right while promising a standard the product does not support. Accepting the whole answer preserves the error. Rejecting it throws away the useful detail. I prefer a verification boundary that can make those decisions separately and explain each one.
That preference has a practical consequence for commerce teams: an answer should earn trust at the level of its checkable assertions. A fluent sentence is a convenient way to communicate, but it bundles claims that may have different relationships to the catalog. The review task is to separate those relationships before composing the answer a shopper sees.
Our Vouchmark demo makes this boundary inspectable using synthetic products, prepared answers and local catalog records. It wraps supplied drafts rather than operating a shopping assistant or a store. Its most useful example is a television answer that needs a small correction, followed by a jacket answer for which a correction would be unjustified.
Preserve what the catalog supports
In the synthetic television case, a draft says the Vega OLED supports HDR10+ and a 120Hz refresh rate. Those are two distinct specifications. HDR10 and HDR10+ are different standards, so the extra character cannot be treated as a harmless wording variation in this comparison.
The fixture catalog contains HDR10 and 120Hz. The Python gate corrects the first assertion and passes the second. The displayed answer therefore retains the refresh rate while making the standard correction explicit.
This is a more useful outcome than replacing the draft with a generic refusal. The catalog can answer the question; it just cannot support both parts of the proposed answer. It is also more accountable than asking a model for a better-sounding rewrite without preserving which assertion failed. The correction has a specific basis: the claimed standard differs from the catalog value.

For a team assessing this kind of design, preservation deserves its own test. A system that suppresses every answer could avoid serving many unsupported details while contributing little product information. The relevant question is whether it can remove an unsupported assertion without discarding adjacent claims that the available record supports.
Vouchmark's prepared scenarios supply the individual claims. They do not demonstrate that a model found every assertion in the original prose. The television example establishes the behavior of this catalog decision, with the decomposition already provided.
An unknown attribute cannot supply a correction
The synthetic Summit Ridge jacket exposes a different problem. Its draft promises that the jacket is fully waterproof. The catalog's waterproof-rating field is unknown, with no manufacturer rating on file in the fixture.
The gate abstains. It says the available record cannot support the waterproof claim. It does not replace the draft with an assertion that the jacket is not waterproof.
That distinction matters because the catalog is silent on the disputed property. An absent rating cannot justify either side of the claim. If the system turns every unsupported assertion into its opposite, it manufactures a new answer under the appearance of correction.
There is a tempting alternative: use a nearby attribute to keep the answer helpful. The jacket record includes an inferred water-resistance description associated with a coating. But water resistance and a confirmed waterproof rating are different claims. This prepared scenario does not serve the inferred description as a substitute for the missing rating.
I want that separation to survive the pressure to provide a neat answer. Under these conditions, a qualified statement is preferable because it identifies the unresolved property without silently changing the shopper's question. The cost is real: the shopper still lacks the rating they asked for. In a production design, resolving that gap would require a suitable source or a separate review process; the demo does not obtain either.
A useful catalog review therefore distinguishes two repair jobs. A contradiction asks the team to reconcile competing values. An unknown asks the team to establish a missing fact. Sending both to a queue labeled “wrong answer” conceals what work remains and what the assistant can safely say in the meantime.
A catalog match has an authority boundary
Claim-level decisions make the answer easier to inspect. They do not make the catalog infallible.
The source label visible in the television dialog identifies an authored fixture reference. It is useful for following the demonstrated decision, but it is not evidence that Vouchmark retrieved and authenticated a manufacturer document. The demo also permits matching attributes marked inferred to pass through the same comparison path as those marked verified. Its serving text can call such a match “verified” while the receipt retains the inferred label.
This exposes a policy choice that a commerce team has to make explicitly. If inferred attributes are acceptable for a particular answer, the displayed qualification should preserve that status. If the business requires direct confirmation for that property, matching an inferred catalog value should not be enough to serve it as confirmed. Vouchmark demonstrates the comparison mechanics; it does not settle that production evidence policy.
I favor keeping the evidence status visible at the point where the claim is used. Otherwise an inference can acquire apparent authority merely by passing through another component. A correct comparison and a trustworthy source are related requirements, but one cannot stand in for the other.
The same reasoning limits what an audit receipt can establish. Vouchmark's readable record keeps the draft, displayed answer, claim decisions and references together. That supports inspection of why the answer changed. Its timestamp and hash-derived identifier do not create a signed, immutable or independently held record. A production team that needs those properties has additional work to do beyond exporting the explanation.
Check the gate and the coverage separately
The demo's fixed labeled synthetic evaluation produces the expected verdict for all 60 already-decomposed claim cases. That result is a check of the finite rules against their supplied labels. It does not measure how completely a shopping assistant's free-form answer can be decomposed, or establish performance on real shopper traffic.
An ordinary hypothetical makes the distinction clear. Suppose a draft says a laptop has the listed amount of memory and is ideal for editing video all day. A verifier might receive the memory claim and correctly pass it. The suitability and endurance promises could remain outside the extracted set. Perfect decisions on the submitted claim would leave the consequential parts of the answer unexamined.
Vouchmark has a narrow deterministic completeness rule that illustrates one way to surface an omission. In its synthetic gaming-laptop scenario, the supplied claims contain only memory. A configured rule adds a dedicated-graphics assertion, which the catalog then contradicts. That mapping is a chosen demo interpretation of the wording, not a universal definition of gaming suitability or exhaustive detection of implied claims.
The implication for evaluation is specific. A team needs evidence about the decisions made on supplied claims, and separate evidence about which assertions reach that decision layer. Testing the full draft against an independently prepared claim set would help examine the latter. It would also expose whether apparently helpful language adds promises the comparison rules cannot represent.
There is a tradeoff here. Restricting served text to catalog-grounded templates can narrow what the system has to inspect, at the cost of expressive range. Allowing richer prose can improve the explanation while increasing the set of explicit and implied promises that need coverage. A clean gate benchmark cannot choose between those designs on its own.
The Vouchmark explainer shows the synthetic examples and decision record. For a commerce team, the more important review artifact is the boundary around the proposed answer: which assertions were identified, what evidence each relied on, and what remained unresolved.
Here is the founder walkthrough of these catalog decisions in Vouchmark.
I judge a verification layer by whether it can preserve an answer's useful detail while making those unresolved parts visible. Approval of a complete answer should require a reason to believe its material assertions reached the gate, together with a policy for the quality of evidence behind them. Until then, an accurate verdict is a result about a submitted claim, and the remaining answer still needs an account.

