I watched CertaRoute's synthetic A-4471 case reach a 0.985-confidence denial while individual clinical factors supplied only 22.06% of the model's attribution. I realized the confidence number was answering a different question from the one a physician would have to answer. The case was a post-acute skilled nursing extension. A high model score could not tell me whether the person's own clinical factors had received enough weight.
The score I could not treat as a decision
I opened A-4471's case file expecting the confidence figure to be the most conspicuous number. The attribution bars were harder to ignore. The recovery-timeline gap contributed about 49% of absolute model attribution, and prior utilization about 19%. Individual clinical factors together contributed 22.06%. That did not mean the model had ignored the patient or that the denial was medically wrong. It meant the model's own explanation showed a large share coming from factors other than the individual's clinical evidence.

The synthetic case file shows the post-acute skilled nursing request, clinical indicators and a pending physician-review state. The bottom subtitle belongs to the captured walkthrough video.
I could have presented this as an explainability exercise: show the chart, attach a paragraph, move on. That would leave the routing decision where it began. A compliance or medical-management leader still needs to know what happens when the explanation reveals insufficient individual consideration. CMS's February 2024 coverage-criteria and utilization-management FAQ says Medicare Advantage coverage determinations must account for the individual patient's circumstances; an algorithm based on a larger dataset cannot substitute for that review. It does not prescribe our 35% threshold. That number is a configurable rule in this demonstration, not a regulatory safe harbor.
Why I would not let the score close the case
I followed the teal clinical bars to the review check. In this seeded case, the rule compares the exact Shapley share from individual clinical factors with a configured 35% floor for a salient denial. A-4471's 22.06% fell below it. The deterministic governance gate set NEEDS_PHYSICIAN_PROOF and the record was marked NEEDS_PROOF. The case is pending physician review, not a finalized denial with a polished explanation attached.

The fired check appears beneath the attribution bars. The interface rounds the individual-factor share to 22% and shows the 35% configured minimum.
I read the other green checks cautiously. The screen's Coverage requirement: PASS is an attestation based on routing, not a comparison with an actual Evidence of Coverage document. Its record-completeness check receives a field-presence flag that this pipeline always passes as True; it does not inspect the completeness of the underlying record.
I made that route consequential in the design. The optional model-generated text can help explain the case; the default no-key path produces a deterministic explanation. Neither text path gets to authorize the disposition. The code gate owns routing. That separation matters because a fluent rationale can make a denial sound considered even when the underlying factors tell a less reassuring story.
I also resisted calling the queue a physician decision. The current queue is a state label in a local demonstration, not a staffed workflow. A real reviewer would need the clinical record, the governing coverage criteria and authority to make an individualized determination. The demo shows where that handoff should occur, not that it has occurred.
The record I wanted to reopen
I went back to A-4471 after the worklist finished and reconstructed its local record. I wanted to see the input features, attribution, routing result and chain check together, not only the model's final score. In the fixed run, all 253 synthetic cases received a local hash-chained SQLite record and 253/253 links verified before tampering. A-4471's record retained the NEEDS_PROOF state, and the local verifier could check the chain. No completed physician assessment is represented.
I value the tamper test for its narrower lesson: when the demo control alters a stored record without recomputing its hash, the local verifier identifies a broken chain. I saw the modal name A-4471 as the altered record and display the failed verification. A hash check cannot establish independent custody, the clinical correctness of the assessment or legal admissibility. Those distinctions belong beside the attractive integrity indicator, not in fine print after it.

After the demo control changes A-4471 without recomputing its hash, the reconstruction modal reports a broken local chain. The visible 0/253 verification count is the tamper-test state, not the intact-run result.
What I would ask before a real deployment
I built CertaRoute to make a specific design choice visible: when a confident initial model assessment leans too little on individual clinical factors, the system should expose that fact and route the case for human judgment with a checkable technical record. The 35% floor and the synthetic A-4471 example are ways to interrogate that choice, not evidence that the threshold is clinically validated. The stand-in QNXT label is not a connector. The app's DEFENSIBLE label means its own demo-state logic found a reconstructible record; it is not a legal conclusion.
I would ask a plan to validate the policy against its clinical and coverage criteria, connect the actual data and reviewer workflow, test the record controls under its security model, and examine which cases the gate misses. In this synthetic run, the gate routed 92 of 253 cases to a physician-review label. That volume is a design consideration, not a deployment forecast. I would rather show that burden honestly than make the queue disappear in a presentation.
I put the full CertaRoute breakdown beside the walkthrough so the case, rule and limitations can be inspected together. The number I keep coming back to is not 0.985. It is the 22.06% share that forced the denial out of the straight-through path while a human decision was still missing.