In a ForenChain synthetic test, a 63-year-old asks whether to move a $480k retirement balance into one crypto token. The policy gate withholds a specific recommendation, but an answer-only log would not explain why. When an AI response is challenged later, a team needs the release decision and its surviving record, not just the final words.
This is an architectural issue for AI liability controls. Improving the classifier may help identify a risky request, but classification is still advice unless a separate release gate enforces a policy. We built ForenChain to make that distinction visible in a synthetic interaction, not to claim that a local demonstration settles a legal question.
A decision has to exist before there can be a reviewable record
The scripted customer says a friend recommended the token. A model might phrase an answer cautiously and still provide a specific allocation recommendation. A log of that answer would tell a reviewer what appeared on screen. It would not establish whether a configured control authorized its release.
In ForenChain, an advisory classifier proposes FINANCIAL_ADVICE, a risk tier, and confidence. Plain Python then checks the raw request for hard signals and applies the loaded representative financial policy pack, FIN-SEC-FINRA-NO-SPECIFIC-REC. For this request, the gate returns TRANSFORM. Code withholds the specific recommendation and supplies general educational information with a disclaimer. The classifier does not have authority to override that action.
This is more than a preference for cautious wording. The route is an explicit decision between ALLOW, TRANSFORM, BLOCK, and HUMAN_REVIEW, with the applied control recorded alongside it. A team investigating the output has a specific policy decision to inspect rather than an after-the-fact interpretation of the model's prose.

In the captured interaction, the synthetic retirement question produces general information rather than a specific allocation recommendation. The trace shows the safeguard, transformed response, and committed decision record.
A record of the answer is incomplete if the release decision that produced it exists only as an inference.
The record has to be inspectable, and its limits have to be clear
ForenChain appends each local decision to a SQLite ledger. Its SHA-256 record hash incorporates the previous hash and canonical fields for the current decision. The chain verifier recomputes the sequence and identifies the first broken link. That gives an investigator a way to check whether the demonstrated local records still match their stored hashes.
The captured interaction above is a later run. The seeded instance of the same retirement request appears as record #3 with a TRANSFORM action. In a simulated alteration, the interface changes that action to ALLOW without recomputing the stored hash. The verifier points to #3 as the first broken record. The edit is visible because a linked record and a verifier exist; the model did not need to remember or explain what happened.

After the simulated alteration, the local verifier identifies record #3 as the first broken link. This detects the shown edit; it is not independent authentication of the ledger.
A local hash chain has a narrow, useful meaning. It does not prevent an administrator from replacing the entire database. It does not supply independent signing, a trusted timestamp, an external archive, or production retention controls. It cannot by itself establish legal chain of custody or admissibility. Those require their own design and assessment.
The standalone HTML evidence package is similarly bounded. It gathers the synthetic matter envelope, decision rows, applied controls, and hash-chain manifest into a technical-evidence scaffold, including a Reasonable Alternative Design argument scaffold for counsel to evaluate. It is not a legal opinion or a completed defense.

The export organizes technical facts for counsel's review. Its own text says the local verifier does not establish independent custody or prove admissibility.
A useful gate must also show where it stops
The architecture has another important boundary: the four loaded packs are representative, not universal coverage. In a scripted medication-dosage question, the classifier recognizes a medical-advice topic, but no medical pack is loaded. The gate returns HUMAN_REVIEW and records that route. It does not contact a clinician or provide medical guidance. A visible escalation is more informative than silently treating a recognized but unsupported request as safe.
Our fixed, synthetic labeled set contains 28 cases. The local regression run matched the expected action on 28/28, routed all 18/18 in-coverage high-risk cases to TRANSFORM or BLOCK, and sent both 2/2 out-of-coverage cases to HUMAN_REVIEW. Those figures test this rule set against those labeled inputs. They are not open-world accuracy or safety rates, and they do not show a customer deployment.
The meaningful boundary is not just where a policy acts, but where the system records that no loaded policy can decide.
The full breakdown shows the decision trace, ledger check, evidence export, and the scope of the synthetic test. If your team logs out-of-coverage escalations, which fields have proved essential when a later reviewer needs to reconstruct that handoff?