
- In 2022 a pharma drug-discovery model flipped one reward sign and generated 40,000 toxic molecules — including VX analogues — in under 6 hours. On a standard server. The gap between drug design and weapon design was a single reward sign in a Python config. 🧵
- The threat didn't stand still. In 2025 GeneBreaker jailbroke Evo 2-40B, a 40B-param open-weight DNA model, at up to 60% attack success across 6 viral categories. It never asks for a pathogen. It asks for a protein "homologous to" a benign one. (NeurIPS 2025)
- That's the whole problem with refusal training. Keyword filters and RLHF are built to catch "design me a nerve agent." Homology-guided beam search never says the word. It reads like comparative genomics right up until you analyze the function of what came out.
- "We unlearned the dangerous knowledge" isn't a defense either. Benign relearning on public medical articles jogs an RMU-unlearned model back toward its forgotten capability (CMU/ICLR 2025). Unlearning is deep obfuscation, not erasure.
- For any open-weight bio model on-prem: 10-50 fine-tuning examples and a few hundred dollars of GPU strip the safety alignment. On open Llama weights, near-frontier bio capability is back in days on a single H100 (arXiv 2508.03153). "Safety-aligned" means nothing here.
- The last line of defense fell too. The Paraphrase Project (Microsoft/Twist/IDT, Science, Oct 2025) used a protein diffusion model to generate thousands of ricin variants that slipped past the homology-based DNA screening every major synthesis provider uses.
- It took a 10-month coordinated patch effort to close that gap. So "our DNA vendor will catch it" was a false guarantee through most of 2024-2025. A pharma has to screen in-house, before the order ever leaves the building.
- Meanwhile the rules diverged. The US rescinded its AI biosecurity EO framework in 2025. The EU AI Act keeps tightening — full high-risk application Aug 2, 2026, penalties up to €35M or 7% of global turnover. EU operations means you comply to the EU standard.
- Chemistry42's 460+ medicinal-chemistry filters screen the output for known toxicophores. They don't touch the objective. Flip the reward sign — the MegaSyn move — and a model steering toward the CWA manifold emits novel structures that pass all 460.
- So no single technique holds. Latent-space governance at generation, unlearning at the weights, and a continuous relearning red-team — layered, with an audit trail — is the only honest posture. "We used RLHF refusal" isn't a control framework the plaintiff's bar accepts post-MFT.
- If your generative-chemistry pipeline runs an open-weight model, what is actually stopping a reward-sign flip today — and could you prove it to an ISO 42001 auditor? #BiosecurityAI #AIGovernance
- We wrote up the full threat model — the three attack vectors, the vendor landscape, and a defense-in-depth blueprint mapped to ISO 42001 and the EU AI Act: https://veriprajna.com/solutions/biosecurity-ai-safety