A credit AI can recommend an outcome the record cannot support. It can also choose a policy-supported outcome without supplying the structured reasons the workflow needs. I want those failures to stay separate, because they call for different repairs.
That distinction shapes our Validation Firewall demo. It checks recommendations outside the model, using an authored credit policy and selected deterministic rules. The example here uses retained model outputs on twelve synthetic loan applications, replayed without fresh inference. The routes shown are simulated; no borrower receives a decision.
When an explanation points to the wrong outcome
One retained response recommends approval for an applicant whose FICO score is 559, below the demo policy's minimum of 620. The record also has a debt-to-income ratio of 55%, above the 43% ceiling. These values are enough to see the conflict with Credit Policy CP-1.
The rationale says the policy has not supplied permitted denial-reason keys, so a policy-compliant denial reason cannot be issued. That is an important integration problem. It supplies no basis for changing the credit criteria, however. Missing the information needed to express a denial cannot make this record eligible for approval.
The separate policy check marks the recommendation BLOCK and lists the failed criteria. It keeps the proposed outcome and the reasons for stopping it visible together.
The retained response recommends approval; the separate policy check identifies the record's conflicts with CP-1. This is a simulated gate, not a borrower decision.
I prefer this boundary because it prevents an explanation about the output format from becoming authority over the credit rule. A team can investigate the missing reason list without accepting the unsupported approval in the meantime.
A supported denial can still need repair
Another retained response recommends denial for a record that fails the configured credit criteria. Its policy-consistency check passes. Its structured principal-reason list is empty, so the reason-key check instead produces an ESCALATE result, indicating a need for review in the demo.
The prose names credit-policy problems, but the structured response does not supply the required principal-reason keys. A downstream workflow that expects those keys cannot treat explanatory prose as a substitute without an explicit design decision.
Policy consistency passes while the separate reason check needs proof. The displayed review route is simulated; this is not a completed human review.
Here, changing the proposed outcome is not the first repair suggested by the evidence. The team needs to inspect why the required reason fields are absent and whether the output contract gave the model enough information to populate them.
The source provides a concrete answer: the prompt supplies the policy criteria but omits the permitted reason list it asks the model to use. That omission matters to both examples. It also makes these responses a poor basis for ranking the model's general credit capability. The contract between the application and the model needs scrutiny alongside the response.
Repair the decision and the contract separately
For a technical decision maker, the useful review question is which failure the evidence actually establishes. Does the proposed outcome conflict with a configured criterion? Are required structured fields missing? Did the request supply the allowed values for those fields? The answers determine whether to reject a recommendation, repair the request, or send an incomplete result for review.
In a production design, I would retain those findings separately and assign the contract repair to the team that owns the model integration. Filling in a reason field would not erase a policy conflict. Blocking an unsupported approval would not fix the request that helped produce it. That separation makes a repeat failure easier to diagnose than one undifferentiated “AI validation failed” label.
The prototype's coverage remains narrow. It compares supplied evidence entries with the record; it does not verify every statement or catch every omission. Its rationale scan uses a finite phrase list. Passing these checks does not certify legal compliance or authorize a production loan decision, and the demo does not create a complete borrower notice.
The full demo breakdown explains the checks and their limits.
My design standard is that an incomplete explanation must remain an incomplete explanation. It should neither override a credit rule nor disappear behind a policy pass.