An AI accuracy review can uncover a clear conflict and still leave important work unfinished. A website makes a numerical promise. A supplied test reports a lower result. Identifying that gap gives the reviewer a reason to challenge the wording, but it does not establish whether the test is sound or what a replacement claim could responsibly say.
We use that distinction in ClaimLens, our demonstration of AI claim substantiation. The useful output is a reviewable relationship between a sentence, its supplied support and the reason for the decision. Keeping that relationship visible helps marketing, engineering and compliance work on the same unresolved question.
What the gap establishes
The example is Nimbus Capital AI, a synthetic company with synthetic disclosures and technical records. Its website claims 98% accuracy for AI document and content detection. The supplied test record reports 53% on a mixed human and AI text sample.
ClaimLens links the sentence to that test record and marks the claim contradicted. The difference is 45 percentage points, which exceeds the demo's configured tolerance of five percentage points. This is a comparison rule chosen for the demonstration. It is not a statistical confidence test or a legally sufficient boundary.
The evidence panel preserves the website wording alongside the supplied result and the rule that produced the decision. A reviewer can see which assertion is in question and which record creates the conflict.
The synthetic content-detection claim conflicts with its supplied test. The visible legal labels are illustrative references, not findings about legal applicability.
The scope of that decision matters. Under the configured comparison, the supplied record contradicts the advertised figure. That gives the team a concrete issue to investigate. It does not establish that every possible performance claim about this system is false, or that the supplied result is an independently verified measurement.
Why changing 98% to 53% would leave a question open
The lower number is tempting because it appears to offer an immediate correction. But replacing one percentage with another still leaves the wording dependent on the test behind it.
In this fixture, the record describes a mixed human and AI text sample. That description is part of what a reviewer needs to carry forward. A result from a particular sample does not, by itself, explain how broadly the result can be described. Copy that drops the sample context can make a wider promise than the record establishes.
For a real review, we would ask the evidence owner to explain what was tested, how the sample was assembled, which system version produced the result and how accuracy was calculated. Those are proposed review questions, not extra checks this demo performs. They help distinguish a disagreement between two numbers from confidence in the measurement itself.
This also changes the handoff to marketing. The next task is to determine which wording the validated evidence supports. Merely substituting the test percentage would skip that determination. Engineering needs to explain the measurement; the person approving the claim needs to understand the promise a reader will take from it.
Keep the reason available for follow-up
We designed the demonstration around supplied records. It operates on the bundled Nimbus dataset and does not offer general document upload or paste. Its policy checks execute during the recording; the AI advisory replies shown alongside them are cached responses. Neither the comparison rule nor those replies independently authenticates the evidence or validates the test methodology.
The review record therefore needs to preserve enough context for a person to continue the work. ClaimLens's completed-run JSON export retains the claim, source, verdict, deciding rule, rationale and evidence identifiers. That makes the numerical conflict traceable to the supplied record. The export references evidence; it does not contain a complete evidence archive.
The ClaimLens explainer shows the wider workflow. Its decisions are configured demo outcomes, not legal clearance. For this example, the important boundary is narrower: the comparison identifies a conflict, while validation of the underlying test remains a human review task.
A useful next step is to ask for an evidence-backed wording proposal that names the tested scope. The reviewer should be able to explain both why the original percentage was challenged and why the proposed sentence is supported. Until the second explanation exists, correcting the number alone has not completed the review.