Drive-Thru Order Firewall
In our synthetic drive-thru example, 18,000 free water cups arrive with vendor confidence 0.97. The quantity cap is eight. The gate holds the order before simulated kitchen submission.
7 min 39 sec walkthrough. Synthetic vendor JSON and simulated point of sale (POS); cached real Codex advisory responses.
18,000
Water cups held
One synthetic order
8
Configured water quantity cap
Saved synthetic order profile
0.97
Vendor confidence input
A score, not a calibrated probability
We separate the interpretation of an order from authority to submit it. Building True Intelligence.
A confidence score describes the vendor's interpretation. It does not answer whether the restaurant permits that quantity. In the water fixture, the menu total is $0.00, so a price-only check has no reason to object. The quantity check does: 18,000 exceeds the saved cap of eight.
That distinction gives an operations team a useful review question: which restaurant rule grants submission permission, and where can an operator inspect the reason it was withheld? The demo preserves the incoming order and shows the deciding evidence rather than treating a confident-looking interpretation as authorization.
The local engine normalizes structured vendor JSON, evaluates eight deterministic checks and applies a policy gate. The checks cover item quantity, observed modifiers, price, daypart, total units in one order, repeated tokens, low vendor confidence and configured injection patterns. The saved historical profile comes from 5,000 seeded synthetic orders; it is not a restaurant chain's operating history.
No rule fires. The engine permits submission to the simulated display.
A non-injection rule fires. Submission stays withheld for confirmation.
The configured injection rule fires. The engine refuses simulated submission.
The water order fires both the item quantity cap and the total-units check. The latter has a 44-unit boundary for a single order. Its UI label says Rate limit, but it does not measure orders across sessions or a time window.
For flagged orders, the advisory note follows the gate and cannot change its decision. This recording replays cached responses from the real configured Codex model. PASS orders skip model adjudication. The displayed timer covers rules plus gate only; model work is synchronous within the full processing request, and the timer excludes that work and delivery.
These retained frames come from the actual local demonstration. Orders, lane imagery and the kitchen display are synthetic or simulated; vendor and menu brand labels are fixture styling, not evidence of integrations, customers or endorsements.
The drawer exposes the incoming 18,000 cups, the cap of eight and the single-order unit boundary. The suggested quantity is eight. That proposal comes from rule evidence and remains separate from the HOLD decision.

Two fries and one burger pass validation in the normal fixture. A receipt is retained for PASS as well as for exceptions, so the review surface does not depend on a model-generated explanation.

Repeated raw tokens produce three burgers in the synthetic interpretation. Repetition and vendor confidence 0.71 below the configured 0.85 threshold produce HOLD, with a one-burger suggestion. Confirmation remains necessary: the engine has not established what the customer intended.

The synthetic order asks for bacon on an ice cream cone. The saved observed modifier set contains chocolate dip and sprinkles, but not bacon. The combination check therefore produces HOLD and proposes removing the modifier. That is a reason to ask for confirmation, not evidence that the combination is physically impossible or the customer is acting maliciously.

This distinction affects the review design. A production policy would need an authoritative menu and a way for an operator to confirm a legitimate exception. The demonstrated profile comes from 5,000 seeded synthetic orders, not operational history from a restaurant chain.
The 260-nugget fixture exceeds its per-item quantity cap of 20. Its $117 menu total also exceeds the configured price boundary of $116.76, and its 260 units exceed the single-order boundary of 44. Three checks agree that the order should wait; none of these ordinary policy exceptions alone produces BLOCK.

The price rule uses the greater of three times the saved historical total statistic and $100: max(3 × $38.92, $100) = $116.76. The total-units rule uses max(2 × 22, 40) = 44. The smaller historical statistics shown in the drawer are inputs to those formulas, not the final firing boundaries. Both boundaries are configured demo policy, not calibrated limits for an operating restaurant.
A breakfast burrito requested at 11:15 is held because this fixture has a 10:30 breakfast cutoff. The item and price can be understood while the request falls outside the configured serving window. The removal suggestion exposes that conflict; it does not confirm what substitute the customer would accept.

This example checks one configured breakfast window against the incoming fixture time. It does not establish live inventory, store-specific schedules, timezone handling or an integrated menu service. Those would need separate design and validation before a real submission path relies on them.
One spicy chicken sandwich has vendor confidence 0.62, below the configured 0.85 threshold. Its ordinary quantity does not remove the uncertainty, so the engine returns HOLD. Unlike the repeated-token example, this case isolates low confidence without requiring a quantity correction.

The appropriate next question is whether the interpreted item matches the request. The demonstration routes that uncertainty for review; it does not diagnose speech, assess an acoustic recording or prove that this threshold gives acceptable production error rates.
The instruction-bearing transcript asks to ignore previous instructions and includes 500 nuggets. The injection pattern fires and produces BLOCK. Quantity, price, unit volume and low confidence also fire, but only the injection rule changes this outcome from HOLD to BLOCK. A finite pattern set cannot establish exhaustive injection resistance.

The receipt retains the order, all eight rule evaluations, decision, suggested corrections and advisory text. The unchanged water receipt verifies through the real local endpoint; changing HOLD to PASS while keeping the original signature fails.


HMAC-SHA256 uses the same shared secret for signing and verification. The default key is public demo material, so anyone who knows it can re-sign a changed body. This demonstrates a bounded local integrity check, not independent custody, immutable storage or a record of completed human action.
The high-volume fixture contains 40 fries and 40 sodas, for 80 total units above the 44-unit boundary. The correction algorithm changes the single worst relative quantity offender: fries fall to their cap of four, but sodas remain at 40. The soda quantity still exceeds its own cap of six. A visually smaller order is therefore not evidence that the whole proposed order would pass.

| Order state | Fries | Sodas | Authority |
|---|---|---|---|
| Incoming order | 40 | 40 | HOLD; simulated submission withheld |
| Suggested edit | 4 | 40 | Not resubmitted or revalidated |
| Per-item caps | 4 | 6 | Saved synthetic-profile boundaries |
Approve Correction and Escalate change their labels and disable themselves. They do not log human action, resubmit, revalidate, release a HOLD, change the receipt or send an order to a real point-of-sale system. A production handoff would need confirmed customer intent, a new validation decision on the complete revised order, and a recorded action before granting submission authority.
On the saved labelled set of 43 synthetic orders, the engine produces 35 PASS, 7 HOLD and 1 BLOCK. All eight fixtures labelled for review or blocking are intercepted; none of the 35 normal fixtures is falsely held. The comparison below uses two simple local code baselines on those same fixtures.

Read the counters with their scope. The displayed $1,251 is a rounded illustrative item-cost estimate of $1,250.80 across four selected withheld fixtures, not measured waste reduction or realized savings. The recorded timer covers only rules plus gate, excluding model work, receipt signing, network and delivery; it is not end-to-end latency. The model call is synchronous within the full request even though its advice cannot change the gate.
On small screens, scroll the comparison table horizontally.
| Local decision approach | Review/block fixtures intercepted | What it checks |
|---|---|---|
| Drive-Thru Order Firewall | 8 of 8 | Eight checks plus the PASS/HOLD/BLOCK gate |
| Quantity over 100 baseline | 3 of 8 | Holds if any raw line quantity exceeds 100 |
| Always-PASS baseline | 0 of 8 | Permits every fixture |
This result establishes tested behavior on a finite labelled stream. It does not estimate field accuracy, production false holds or another vendor's performance. The report is returned by the local evaluation endpoint; there is no visible benchmark scoreboard or OFF toggle in the dashboard.
It does not recognize audio, ingest a real vendor feed, connect to a real POS or complete human review. Thresholds have not been validated for an operating restaurant. No customer deployment, measured savings or production service-level result is demonstrated.
The on-screen waste counter totals illustrative synthetic item costs for selected withheld orders, not realized savings. The timer measures only rules and gate. We recommend testing representative local menus and order traffic, confirming the operator handoff and validating the POS submission boundary before a production design relies on this approach.
Drive-Thru Order Firewall demonstrates a validation layer for a vendor's structured order output. It processes synthetic JSON before a simulated point-of-sale and kitchen display; it does not capture audio, recognize speech or connect to a real vendor.
Any fired rule other than the injection rule produces HOLD and withholds simulated submission. Quantity, price, availability, unfamiliar modifiers, repeated tokens, low confidence and total units can trigger review; the total-units check measures one order, not traffic over time.
The advisory model cannot change the deterministic gate's decision in this engine path. The recording uses cached real Codex advisory responses after the gate; it does not perform fresh inference on every replay.
Approve Correction and Escalate only change their button labels and disable themselves in this demo. They do not release a HOLD, resubmit an order, log human action or write to a real point-of-sale system.
Local HMAC-SHA256 verification checks that a receipt body matches its signature under the same shared secret. Changing the decision without re-signing fails verification; the public demo key permits re-signing by anyone who knows it, so this is not independent custody or immutable storage.
The evaluation uses 43 fixed synthetic orders: 35 PASS, 7 HOLD and 1 BLOCK. All eight labelled review or block fixtures are intercepted with no false holds among the 35 normal fixtures; the simple baselines are local code comparisons, not vendor or restaurant measurements.
Discuss the rules and review path your operation needs.
We can help assess where vendor interpretation becomes transaction authority and design a validation approach for your menu and point-of-sale workflow.
Explore related research for broader context on this demonstration.