Drive-Thru Order Firewall

A confident order still needs permission to proceed.

In our synthetic drive-thru example, 18,000 free water cups arrive with vendor confidence 0.97. The quantity cap is eight. The gate holds the order before simulated kitchen submission.

7 min 39 sec walkthrough. Synthetic vendor JSON and simulated point of sale (POS); cached real Codex advisory responses.

18,000

Water cups held

One synthetic order

8

Configured water quantity cap

Saved synthetic order profile

0.97

Vendor confidence input

A score, not a calibrated probability

We separate the interpretation of an order from authority to submit it. Building True Intelligence.

The price can be right while the quantity needs review

A confidence score describes the vendor's interpretation. It does not answer whether the restaurant permits that quantity. In the water fixture, the menu total is $0.00, so a price-only check has no reason to object. The quantity check does: 18,000 exceeds the saved cap of eight.

That distinction gives an operations team a useful review question: which restaurant rule grants submission permission, and where can an operator inspect the reason it was withheld? The demo preserves the incoming order and shows the deciding evidence rather than treating a confident-looking interpretation as authorization.

Rules decide permission; advice explains the exception

The local engine normalizes structured vendor JSON, evaluates eight deterministic checks and applies a policy gate. The checks cover item quantity, observed modifiers, price, daypart, total units in one order, repeated tokens, low vendor confidence and configured injection patterns. The saved historical profile comes from 5,000 seeded synthetic orders; it is not a restaurant chain's operating history.

PASS

No rule fires. The engine permits submission to the simulated display.

HOLD

A non-injection rule fires. Submission stays withheld for confirmation.

BLOCK

The configured injection rule fires. The engine refuses simulated submission.

The water order fires both the item quantity cap and the total-units check. The latter has a 44-unit boundary for a single order. Its UI label says Rate limit, but it does not measure orders across sessions or a time window.

For flagged orders, the advisory note follows the gate and cannot change its decision. This recording replays cached responses from the real configured Codex model. PASS orders skip model adjudication. The displayed timer covers rules plus gate only; model work is synchronous within the full processing request, and the timer excludes that work and delivery.

Follow the order from interpretation to permission

These retained frames come from the actual local demonstration. Orders, lane imagery and the kitchen display are synthetic or simulated; vendor and menu brand labels are fixture styling, not evidence of integrations, customers or endorsements.

The water order waits, even at zero price

The drawer exposes the incoming 18,000 cups, the cap of eight and the single-order unit boundary. The suggested quantity is eight. That proposal comes from rule evidence and remains separate from the HOLD decision.

Synthetic water order held with 18,000 cups, quantity cap 8 and total-unit hard cap 44
The water drawer shows a zero menu total, two fired rules and a suggested quantity of eight. Its operator explanation is a cached real model response. Open full-size evidence

Ordinary orders still pass

Two fries and one burger pass validation in the normal fixture. A receipt is retained for PASS as well as for exceptions, so the review surface does not depend on a model-generated explanation.

Normal synthetic order with two fries and one burger shows passed validation and a receipt
A normal fixture passes without model adjudication. The kitchen-routing message refers to the simulated display. Open full-size evidence

Uncertain meaning deserves confirmation

Repeated raw tokens produce three burgers in the synthetic interpretation. Repetition and vendor confidence 0.71 below the configured 0.85 threshold produce HOLD, with a one-burger suggestion. Confirmation remains necessary: the engine has not established what the customer intended.

Synthetic repeated-token order with three burgers is held with repetition and low-confidence evidence and a one-burger suggestion
The repeated-token fixture is held for confirmation. The proposed quantity of one does not establish the customer's intent. Open full-size evidence

An unfamiliar modifier is a question, not an attack

The synthetic order asks for bacon on an ice cream cone. The saved observed modifier set contains chocolate dip and sprinkles, but not bacon. The combination check therefore produces HOLD and proposes removing the modifier. That is a reason to ask for confirmation, not evidence that the combination is physically impossible or the customer is acting maliciously.

Synthetic ice cream modifier evidence shows bacon absent from observed chocolate dip and sprinkles, with a removal suggestion
The rule exposes the historical absence and the proposed edit. The original order remains held; the suggestion does not establish a complete or correct restaurant menu. Open full-size evidence

This distinction affects the review design. A production policy would need an authoritative menu and a way for an operator to confirm a legitimate exception. The demonstrated profile comes from 5,000 seeded synthetic orders, not operational history from a restaurant chain.

Quantity, price and total units are separate checks

The 260-nugget fixture exceeds its per-item quantity cap of 20. Its $117 menu total also exceeds the configured price boundary of $116.76, and its 260 units exceed the single-order boundary of 44. Three checks agree that the order should wait; none of these ordinary policy exceptions alone produces BLOCK.

Synthetic 260-nugget order shows HOLD with quantity cap 20, price hard cap 116.76 and total-unit hard cap 44
The drawer preserves the incoming quantity and each fired rule. The cached advisory note explains the hold, but it is not the authority that decides it. Open full-size evidence

The price rule uses the greater of three times the saved historical total statistic and $100: max(3 × $38.92, $100) = $116.76. The total-units rule uses max(2 × 22, 40) = 44. The smaller historical statistics shown in the drawer are inputs to those formulas, not the final firing boundaries. Both boundaries are configured demo policy, not calibrated limits for an operating restaurant.

Recognized words still need an availability check

A breakfast burrito requested at 11:15 is held because this fixture has a 10:30 breakfast cutoff. The item and price can be understood while the request falls outside the configured serving window. The removal suggestion exposes that conflict; it does not confirm what substitute the customer would accept.

Synthetic breakfast burrito at 11:15 is held after the configured 10:30 cutoff
Availability is a transaction rule separate from recognition. The shown Approve Correction and Escalate buttons are presentation-only acknowledgements, not a completed operator workflow. Open full-size evidence

This example checks one configured breakfast window against the incoming fixture time. It does not establish live inventory, store-specific schedules, timezone handling or an integrated menu service. Those would need separate design and validation before a real submission path relies on them.

Low confidence can hold an otherwise ordinary order

One spicy chicken sandwich has vendor confidence 0.62, below the configured 0.85 threshold. Its ordinary quantity does not remove the uncertainty, so the engine returns HOLD. Unlike the repeated-token example, this case isolates low confidence without requiring a quantity correction.

One synthetic spicy chicken sandwich is held because confidence 0.62 is below 0.85
The drawer shows the vendor score and the configured threshold. That score is an input to the policy, not a calibrated probability of customer intent. Open full-size evidence

The appropriate next question is whether the interpreted item matches the request. The demonstration routes that uncertainty for review; it does not diagnose speech, assess an acoustic recording or prove that this threshold gives acceptable production error rates.

A configured attack signal has a different outcome

The instruction-bearing transcript asks to ignore previous instructions and includes 500 nuggets. The injection pattern fires and produces BLOCK. Quantity, price, unit volume and low confidence also fire, but only the injection rule changes this outcome from HOLD to BLOCK. A finite pattern set cannot establish exhaustive injection resistance.

Synthetic instruction-bearing transcript and 500 nuggets show BLOCK with the matched injection pattern
The configured injection pattern produces BLOCK. This is evidence for one tested pattern, not exhaustive attack resistance. Open full-size evidence

A receipt checks integrity within a stated boundary

The receipt retains the order, all eight rule evaluations, decision, suggested corrections and advisory text. The unchanged water receipt verifies through the real local endpoint; changing HOLD to PASS while keeping the original signature fails.

Water receipt displays Valid untampered after verification at the local endpoint
The unchanged water receipt verifies under the shared demo secret. The approval and escalation labels visible above are cosmetic acknowledgements. Open full-size evidence
Local receipt verification displays Tamper detected after changing the decision and retaining the original signature
Changing HOLD to PASS while retaining the old signature fails local verification. Someone who knows the public demo key can create a new signature. Open full-size evidence

HMAC-SHA256 uses the same shared secret for signing and verification. The default key is public demo material, so anyone who knows it can re-sign a changed body. This demonstrates a bounded local integrity check, not independent custody, immutable storage or a record of completed human action.

A suggestion is not a released order

The high-volume fixture contains 40 fries and 40 sodas, for 80 total units above the 44-unit boundary. The correction algorithm changes the single worst relative quantity offender: fries fall to their cap of four, but sodas remain at 40. The soda quantity still exceeds its own cap of six. A visually smaller order is therefore not evidence that the whole proposed order would pass.

Synthetic high-volume correction changes 40 fries and 40 sodas into 4 fries and 40 sodas while the original order remains held
Only fries change in the proposal. Forty sodas remain, so a suggested correction must not be read as an approved or fully revalidated order. Open full-size evidence
Same high-volume fixture: a proposal does not change the saved decision.
Order stateFriesSodasAuthority
Incoming order4040HOLD; simulated submission withheld
Suggested edit440Not resubmitted or revalidated
Per-item caps46Saved synthetic-profile boundaries

Approve Correction and Escalate change their labels and disable themselves. They do not log human action, resubmit, revalidate, release a HOLD, change the receipt or send an order to a real point-of-sale system. A production handoff would need confirmed customer intent, a new validation decision on the complete revised order, and a recorded action before granting submission authority.

What the fixed evaluation establishes

On the saved labelled set of 43 synthetic orders, the engine produces 35 PASS, 7 HOLD and 1 BLOCK. All eight fixtures labelled for review or blocking are intercepted; none of the 35 normal fixtures is falsely held. The comparison below uses two simple local code baselines on those same fixtures.

Completed synthetic stream shows 35 sent to the simulated kitchen, 7 held and 1 blocked
The completed replay keeps review holds distinct from the single BLOCK. Its 81% auto-approval display is rounded from 35 of 43 synthetic orders. Open full-size evidence

Read the counters with their scope. The displayed $1,251 is a rounded illustrative item-cost estimate of $1,250.80 across four selected withheld fixtures, not measured waste reduction or realized savings. The recorded timer covers only rules plus gate, excluding model work, receipt signing, network and delivery; it is not end-to-end latency. The model call is synchronous within the full request even though its advice cannot change the gate.

On small screens, scroll the comparison table horizontally.

Same fixed synthetic set, same eight review/block fixtures
Local decision approachReview/block fixtures interceptedWhat it checks
Drive-Thru Order Firewall8 of 8Eight checks plus the PASS/HOLD/BLOCK gate
Quantity over 100 baseline3 of 8Holds if any raw line quantity exceeds 100
Always-PASS baseline0 of 8Permits every fixture

This result establishes tested behavior on a finite labelled stream. It does not estimate field accuracy, production false holds or another vendor's performance. The report is returned by the local evaluation endpoint; there is no visible benchmark scoreboard or OFF toggle in the dashboard.

What this demo does NOT do

It does not recognize audio, ingest a real vendor feed, connect to a real POS or complete human review. Thresholds have not been validated for an operating restaurant. No customer deployment, measured savings or production service-level result is demonstrated.

The on-screen waste counter totals illustrative synthetic item costs for selected withheld orders, not realized savings. The timer measures only rules and gate. We recommend testing representative local menus and order traffic, confirming the operator handoff and validating the POS submission boundary before a production design relies on this approach.

Questions restaurant technology teams ask

Does this replace our drive-thru voice AI vendor?

Drive-Thru Order Firewall demonstrates a validation layer for a vendor's structured order output. It processes synthetic JSON before a simulated point-of-sale and kitchen display; it does not capture audio, recognize speech or connect to a real vendor.

What makes an order wait for human confirmation?

Any fired rule other than the injection rule produces HOLD and withholds simulated submission. Quantity, price, availability, unfamiliar modifiers, repeated tokens, low confidence and total units can trigger review; the total-units check measures one order, not traffic over time.

Can the AI approve an order that fails a rule?

The advisory model cannot change the deterministic gate's decision in this engine path. The recording uses cached real Codex advisory responses after the gate; it does not perform fresh inference on every replay.

Does approving a correction actually send it to the POS?

Approve Correction and Escalate only change their button labels and disable themselves in this demo. They do not release a HOLD, resubmit an order, log human action or write to a real point-of-sale system.

What does verifying an order receipt prove?

Local HMAC-SHA256 verification checks that a receipt body matches its signature under the same shared secret. Changing the decision without re-signing fails verification; the public demo key permits re-signing by anyone who knows it, so this is not independent custody or immutable storage.

Are these results measured in real restaurants?

The evaluation uses 43 fixed synthetic orders: 35 PASS, 7 HOLD and 1 BLOCK. All eight labelled review or block fixtures are intercepted with no false holds among the 35 normal fixtures; the simple baselines are local code comparisons, not vendor or restaurant measurements.

Define your restaurant's order-permission boundary

Discuss the rules and review path your operation needs.

We can help assess where vendor interpretation becomes transaction authority and design a validation approach for your menu and point-of-sale workflow.

Assess the decision boundary

  • ✓ Review structured order inputs
  • ✓ Map quantity and menu policies
  • ✓ Define confirmation cases
  • ✓ Plan representative evaluation

Design the implementation

  • ✓ Separate advice from permission
  • ✓ Specify operator confirmation
  • ✓ Plan POS submission controls
  • ✓ Define receipt trust boundaries

Technical Research

Explore related research for broader context on this demonstration.