Crucible · Model Vetting Firewall

A clean scan leaves the loading question open

A synthetic model passes the PickleScan baseline, then attempts to open a database during loading. Crucible records and blocks that configured effect, returning QUARANTINE with the evidence attached.

23/23 vs 19/23

Blocked-event detection vs PickleScan flags

Same 23 malicious-plus-evader synthetic fixtures

4/4 vs 0/4

Four constructed evaders

Behavioral detection vs PickleScan 1.0.4

2/2

Abstain fixtures routed to REVIEW

Further evidence needed; no signature issued

Cards report one fixed 33-artifact synthetic direct-pipeline reference run on 6 October 2026 with deterministic advisory. They do not estimate detection on unknown models. The video separately captures fresh configured checks in the local app: the main console uses cached Codex advisory, and the benchmark uses deterministic advisory. No fresh model inference is captured.

The decision needs evidence about loading

When a security team approves a serialized model, the relevant question extends beyond its declared identity: what operation does loading attempt, and what evidence remains unresolved?

Python warns that crafted pickle data can execute code during unpickling. Hugging Face's pickle-scanning documentation also describes limitations of import and opcode inspection. Python pickle documentation; Hugging Face pickle-scanning documentation.

Our synthetic SQLite fixture makes that distinction inspectable: a baseline with no infection flag sits beside a blocked database-opening attempt. The admission record preserves both findings instead of treating the clean scanner field as clearance.

How the configured decision is made

  1. Inspect supported pickle content. Static opcode disassembly records globals and approximates invoked callables. The installed PickleScan baseline supplies a separate comparison field; its flag does not directly decide the verdict.
  2. Observe an attempted load. A fresh Python subprocess uses a CPython audit hook to record selected events and raise before configured blocked effects, including SQLite connection and socket connection. This is Python instrumentation with process separation, without container or OS confinement.
  3. Apply the gate and retain uncertainty. A blocked behavioral event returns QUARANTINE. A runner crash or timeout, pickle-disassembly error, or unexercised dangerous static global returns REVIEW. Otherwise the base gate returns ALLOW; challenger doubt can change ALLOW to REVIEW, while QUARANTINE remains in force.
  4. Attach a bounded record. Every result receives a minimal CycloneDX-shaped model inventory and local hash-chain fields. Only ALLOW receives a development-key signature over model-name, artifact hash and inventory. Unknown upstream history stays UNKNOWN.

The advisory uses analyst and challenger roles in one combined request. The recorded main console uses cached advice; the benchmark uses deterministic advisory. These configured checks do not guarantee that every malformed file, unsupported format or analysis error routes to REVIEW.

Follow one artifact from clean scan to admission decision

All artifacts, model names and hf:// source labels below are synthetic local fixtures, not customer models or verified registry records. The first three screenshots capture fresh configured checks with cached Codex advisory; the separate benchmark capture uses deterministic advisory. No fresh model inference is shown.

Worked example: a clean baseline, a blocked database attempt

The generated trusted-looking/finetune-safe pickle attempts to open a SQLite database during deserialization. Its name is an authored fixture label, not evidence of trust. The useful question is whether the scanner finding and the observed loading behavior support the same admission decision.

Configured result

PickleScan: CLEAN. Observed operation: sqlite3.connect, attempted and blocked. Final verdict: QUARANTINE. Signature: none.

Crucible synthetic SQLite evader showing PickleScan CLEAN, blocked sqlite3.connect and QUARANTINE
Synthetic SQLite evader: PickleScan does not flag the artifact; the configured audit hook records and blocks its attempted database operation. QUARANTINE issues no signature. NO CODE SURFACE means no configured dangerous-global hit; _sqlite3.connect remains present. The UI's sandbox wording refers to subprocess audit instrumentation, without OS or container confinement. Open the image for full-size inspection.

1. Read the scanner and static findings separately

PickleScan 1.0.4 records _sqlite3.connect as suspicious without setting its infection flag. Crucible's static disassembly also retains that imported callable, but it is absent from the configured dangerous-global set. The visible NO CODE SURFACE badge therefore means no configured dangerous-global hit; it does not mean the file contains no executable callable.

Evidence for the synthetic SQLite artifact
CheckRecorded findingWhat it establishes
PickleScan baselineflagged: false; _sqlite3.connect [suspicious]This baseline does not flag the artifact. It does not establish harmless loading.
Static disassembly_sqlite3.connect in imports and approximate callables; no configured dangerous globalThe callable is visible even though the configured denylist has no hit.
Observed loadsqlite3.connect with blocked: true; loaded: falseThe audit hook raises before the configured database-opening effect.
Final gateQUARANTINE; signature: nullThe blocked attempt determines this verdict. No signature is issued.

2. Use the attempted effect to decide the route

The fresh Python worker reaches sqlite3.connect for /tmp/vp_demo_persist/.store.db. Its CPython audit hook records the operation and raises before the configured effect. The gate returns QUARANTINE because a blocked dangerous event was observed, regardless of the clean baseline flag. This evidence does not show a created database or successful persistence.

The UI calls this worker a sandbox. Its implemented boundary is a subprocess with selected Python audit hooks, without an operating-system sandbox, container confinement or network isolation. A production admission system needs a separately established containment boundary.

3. Keep the decision tied to the artifact and record

The downloadable JSON record associates the artifact SHA-256 with its static findings, baseline result, attempted calls, gate reason, minimal model inventory and local hash-chain fields. For this QUARANTINE result, the signature field is null. Reviewers can inspect the decision evidence without treating advisory recommendations as an approval or a completed registry action.

A hash identifies the inspected artifact bytes. The local chain supports consistency checks between records, but it has no independent custody or external anchor and is not an immutable archive. The signed ALLOW payload shown below covers a narrower set of fields than the full evidence record.

A quiet load leaves a conditional finding unresolved

The separate synthetic acme/experimental-rl fixture contains builtins.eval in static inspection. Its conditional branch is not exercised in this environment, and the observed load records no blocked dangerous event. The unresolved static finding sends it to REVIEW with no signature. This route preserves the need for more evidence; no completed human investigation is shown.

Crucible synthetic conditional fixture showing builtins.eval, no blocked runtime event and REVIEW
Synthetic conditional fixture: static inspection finds builtins.eval, while the observed load records no blocked dangerous event. REVIEW issues no signature and routes the artifact for further evidence rather than completed human review. Open the image for full-size inspection.

ALLOW signs the local inventory while provenance stays unknown

The generated acme/sentiment-mlp weights dictionary follows the clean path: no blocked dangerous event is recorded, the configured checks return ALLOW and an Ed25519 signature is issued. Its inventory names the artifact and hash, serialization format, inferred framework and declared source. Training-data provenance and fine-tuning history remain UNKNOWN.

Crucible synthetic clean weights inventory showing UNKNOWN provenance and local Ed25519 signature
Synthetic clean weights dictionary: ALLOW receives a local development-key signature over model-name, hash and inventory. Training provenance and fine-tuning history remain UNKNOWN. The visible framework-reference text is a configured label, without legal validation or a compliance finding. Open the image for full-size inspection.

The signature authenticates the canonical model-name, artifact-hash and minimal CycloneDX-shaped inventory payload relative to a local development key. It does not sign every verdict or the entire record, fill upstream history, establish training rights or prove an arbitrary model safe. The visible regulatory-reference text is configured fixture metadata, not validated compliance.

Read the fixed-set comparison with its denominator

One frozen direct-pipeline reference run on 6 October 2026 uses deterministic advisory, PickleScan 1.0.4, temporary development keys and a temporary ledger. Its 33 generated artifacts comprise 8 benign, 19 malicious, 4 constructed evaders and 2 abstain fixtures. The malicious-plus-evader comparison counts the same 23 artifacts in both columns.

Fixed 33-artifact synthetic direct-pipeline reference run
MeasurementObserved resultScope
Malicious-plus-evader detectionBehavioral 23/23; PickleScan 19/23Blocked-event detection compared with the baseline infection flag on the same 23 fixtures.
Constructed evadersBehavioral 4/4; PickleScan 0/4Four fixtures designed to illustrate the difference between these checks.
Benign decisions0/8 given a verdict other than ALLOWEight benign fixtures, not an estimate of false positives on unknown models.
ALLOW signatures8/8 issued and verifiedVerification uses the included function and local development key.
Abstain route2/2 REVIEWUnresolved fixtures remain unsigned; investigation is not completed.
Expected verdicts and local chain33/33 matched; hash links intactAuthored fixture expectations and a local consistency check, without an external anchor.

The screenshot below is a separate completed HTTP/SSE benchmark run in the local app, with deterministic advisory. It shows the same fixed-set comparison and 33/33 expected verdict matches. It is not the source of the frozen direct-pipeline reference measurement above; its displayed timings belong to that captured run.

Completed Crucible synthetic benchmark showing 33 of 33 expected verdicts, behavioral detection 23 of 23 and PickleScan flags 19 of 23
Completed local HTTP/SSE benchmark: all 33 synthetic fixtures match their expected verdicts. Behavioral blocked-event detection is 23/23 and PickleScan flags 19/23 on the same malicious-plus-evader set. Advisory is deterministic. Only ALLOW signs the inventory payload; REVIEW and QUARANTINE are unsigned. The local hash chain has no external anchor, and displayed timings are not production latency. Open the image for full-size inspection.

These constructed-fixture observations do not estimate detection on unseen models, production latency or breach reduction. ALLOW describes the outcome of the configured checks on the observed load; it does not establish exhaustive model security.

What each layer can establish

LayerEvidence in this demoBoundary to retain
Static inspection and PickleScanGlobals, approximate callables and baseline flagA clean flag alone does not resolve loading behavior
Behavioral observationSelected attempted effects on one observed loadA quiet load can leave conditional behavior unresolved
Signed inventoryLocal model-name, hash and inventory payload for ALLOWThe signature does not attest unknown upstream history
Local hash-chained ledgerHash links supporting local consistency checksNo independent custody or external anchor

What this demo does NOT do

Crucible is a local demonstration on synthetic artifacts. It has no public-registry connector, enterprise admission enforcement, OS sandbox, production key infrastructure or complete dependency reconstruction. It does not evaluate model quality, inference safety or training-data poisoning, and its framework-reference labels do not establish compliance.

Python cautions that audit hooks are unsuitable for implementing a sandbox. Python audit-hook documentation. Production work must establish containment, trust boundaries and controlled custody beyond this local demonstration.

Questions security and platform teams ask

What can a clean pickle scan establish?

A clean PickleScan result means that this baseline did not flag the inspected artifact. In Crucible's synthetic SQLite example, the configured audit hook records and blocks an attempted database operation during loading while the baseline remains clean. A scanner result alone does not establish side-effect-free loading.

What if suspicious static evidence is not exercised during loading?

Crucible routes a configured dangerous static global to REVIEW when the observed load does not exercise a blocked dangerous event. The synthetic conditional fixture contains builtins.eval and takes this route without a signature. REVIEW requests further evidence; it does not mean a human investigation is complete.

What does the signature cover?

Only ALLOW receives an Ed25519 signature over the canonical model-name, artifact-hash and inventory payload, using a local development key. It authenticates that payload relative to this key. It does not establish upstream custody, source authenticity or complete provenance.

What upstream provenance remains unknown?

The minimal CycloneDX-shaped model inventory records training-data provenance and fine-tuning history as UNKNOWN. It includes the artifact hash, serialization format, inferred framework and declared source. An ALLOW verdict and a valid local signature do not fill the missing history.

How is the worker isolated?

The worker is a fresh Python subprocess with a temporary working directory and a CPython audit hook that records selected events and blocks configured effects. It has no container or operating-system sandbox and is not a network-isolated environment. This demonstration does not establish production containment or exhaustive security.

Does this integrate with our registry and admission pipeline?

The demonstration reads generated artifacts from a local synthetic registry. Its hf:// source strings are fixture labels, and it has no public-registry connector or enterprise admission enforcement. Production integration would require registry trust boundaries, containment, key management and independently controlled audit storage.

Technical Research

Explore related research for broader context on this demonstration.

Define the evidence your admission gate needs

Discuss your model-intake workflow with our team.

We use these demonstrated distinctions to frame an assessment or implementation conversation around your registry, loading boundary and evidence requirements.

Admission design assessment

  • ✓ Artifact formats and intake paths
  • ✓ Checks and unresolved verdicts
  • ✓ Loading and isolation boundaries
  • ✓ Inventory and audit requirements

Production implementation planning

  • ✓ Registry and pipeline integration
  • ✓ Containment and deployment design
  • ✓ Signing keys and evidence custody
  • ✓ Representative evaluation plan