Conceptual drive-thru worker comparing order slips showing three burgers and one burger before confirmation.
Artificial IntelligenceProduct ManagementSoftware EngineeringUX

What Should a Human Confirm in a Voice AI Order?

Ashutosh SinghalAshutosh SinghalAugust 4, 20266 min read

In a synthetic voice ordering example, a repeated speech token becomes three burgers. The system offers a plausible correction: change the quantity to one. For a product team, the difficult question is what happens between that suggestion and permission to submit the order. A helpful-looking replacement has not established what the customer wanted.

Disclosure: The examples below are synthetic orders in Veriprajna's Drive-Thru Order Firewall, which validates structured vendor output before a simulated kitchen display. It does not recognize real audio or send orders to a restaurant's point-of-sale system.

I want human confirmation to resolve a specific uncertainty. That requires separating three judgments: what the customer intended, whether that order is allowed under the restaurant's policy, and whether the exact proposed order satisfies the checks. Combining those judgments into one approval button makes it difficult to know what an approval means.

A correction is a hypothesis about intent

The repeated-token fixture contains a disfluent transcript, three identical raw tokens and a quantity of three burgers. Its supplied confidence score is 0.71, below the configured review threshold of 0.85. Repetition and low-confidence rules therefore place the order on HOLD. The repetition rule suggests one burger.

One is a reasonable candidate to put in front of an operator. It is not proof of the intended quantity. The repeated token might explain the structured output, but a rule that notices repetition cannot ask the customer what they meant. Automatically replacing three with one would trade a questionable interpretation for an unconfirmed interpretation.

Synthetic order detail showing confidence 0.71 against threshold 0.85, cached model advice, and a suggested change from three burgers to one
The synthetic fixture proposes one burger; it does not establish customer intent. The visible approval control only acknowledges a click in this demo, and the background timer measures rules plus gate only.

Rejecting the order has a different cost. It treats an interpretation that needs clarification as if no acceptable order can be recovered. I prefer a hold at this point because it preserves the original evidence and leaves the intended quantity open. In a production design, confirmation should ask about the quantity itself, rather than ask an operator to endorse the system's general confidence.

That is a design position, not a completed workflow in this app. The demo's Approve Correction and Escalate buttons change their labels and disable themselves. They do not resubmit, release, record a human action or change the receipt. A production team would still need to build the conversation and the state change that makes its answer consequential.

An unusual request may be accurately understood

Interpretation is only one reason to pause. Another synthetic fixture contains 18,000 water cups with a supplied confidence score of 0.97 and a zero customer menu price. The quantity exceeds the configured water cap of eight, so the gate holds it. Neither the high input score nor the zero price answers whether that quantity may proceed. The score is an input from the fixture, not a calibrated probability of permission.

An extreme quantity makes this distinction easy to see. The harder product decision is an unusual quantity that a customer actually wants. Consider a hypothetical group order that exceeds a restaurant's ordinary automatic limit. If the customer confirms the number, interpretation may be settled while permission remains unresolved. Reducing it to the usual cap would change the request. Rejecting it outright might discard legitimate demand.

I prefer using an exception threshold to trigger review when the policy permits an exception. The operator then needs to decide whether the restaurant can accept the confirmed order, possibly through a separately authorized route. A threshold designed to limit automatic submission should not quietly become a rule that rewrites customer intent.

This costs attention. A looser automatic threshold allows more unusual orders through; a tighter one creates more review work. The demo cannot choose that balance for a restaurant. Its caps come from seeded synthetic order history, and a distribution describes what appeared in that history. It does not establish a real location's capacity or how frequently a legitimate group order will occur. Before adopting such a policy, a team would need evidence about acceptable exceptions and the practical burden of confirming them.

Confirmation must apply to the exact next order

Even a confirmed correction can remain invalid. The high-volume fixture starts with 40 fries and 40 sodas. The quantity rule proposes reducing fries to four, but leaves 40 sodas unchanged. The soda cap is six. Agreeing that four fries is the right replacement would therefore leave another quantity violation in the proposed order.

Synthetic order detail comparing 40 fries and 40 sodas with a suggestion of four fries and 40 sodas, above approval and escalation buttons
The suggestion changes only fries. Forty sodas remain, so the proposed order has not been shown to pass. The UI's “Rate limit” checks total units in one order, not requests over time; its actual boundary is 44 units.

This is why I keep confirmation and validation distinct. Customer confirmation addresses intent. An authorized operator's exception decision addresses policy. Checking the complete proposed order addresses whether the next transaction meets the applicable rules. None of those answers can safely be inferred from the others.

For a production workflow, I would require a confirmed proposal to pass through validation again before submission. If an operator can override a policy, that authority should be explicit and attached to the particular rule and order being accepted. A generic approval should not erase unrelated failures. These are requirements for a future implementation, not capabilities demonstrated by the current correction controls.

Advice can help without granting permission

The explanatory model note has a useful but narrower role: it can make the reason for a hold easier to read. In the accompanying founder recording, that note uses cached real advisory responses. It is not fresh inference on every replay. The engine decides first using deterministic rules; the subsequent advisory text cannot change the decision. The advisory call is synchronous within processing, so this separation of authority does not establish that model work adds no request latency.

The full order-validation breakdown shows this boundary and the synthetic examples. They support an inspectable design argument, not evidence that the seeded policy or unfinished operator workflow is ready for deployment.

Here is the founder recording of the order-review example.

For me, the decisive product review is to follow the proposed order all the way to its next permitted action. Who confirmed its meaning? Who may accept an exception? What checked the complete result? Until those answers refer to the same exact order, a reassuring correction and an approval label leave the transaction unfinished.

Related Research

Also Published On

Build Your AI with Confidence.

Partner with a team that has deep experience in building the next generation of enterprise AI. Let us help you design, build, and deploy an AI strategy you can trust.

Veriprajna Deep Tech Consultancy specializes in building safety-critical AI systems for healthcare, finance, and regulatory domains. Our architectures are validated against established protocols with comprehensive compliance documentation.