
40,000 toxic molecules — including VX analogues — generated in under 6 hours. The change? One flipped reward sign.
That was MegaSyn, 2022. A commercial drug-discovery model at Collaborations Pharmaceuticals, run on a standard server. Drug-design mode and weapons-design mode were separated by a single value in a Python config.
Most pharma generative chemistry pipelines still have that exact architecture today.
What changed is the attack surface. The defenses teams trust — refusal training, RLHF alignment, structural-alert filters — were built for prompts like "design me a nerve agent." The 2025 attacks don't look like that:
— GeneBreaker (NeurIPS 2025) hit a 60% attack success rate jailbreaking the open-weight Evo 2-40B DNA model — not by asking for a pathogen, but for a protein "homologous to" a benign one. Keyword filters are blind to it.
— The Paraphrase Project (Microsoft, Twist, IDT, in Science) showed thousands of AI-paraphrased ricin variants slipping past the homology-based DNA synthesis screening every major provider uses. The "last line of defense" failed for months.
— And "knowledge-gapped" unlearning isn't erasure: relearning on innocuous medical text can pull the forbidden capability back out.
Our read after working this: in 2026, "we added a refusal prompt" is not a defense, and it's not a liability shield. Under the EU AI Act — full application Aug 2, 2026, penalties up to €35M or 7% of global turnover — it reads as an audit finding.
Defense that actually holds is layered: latent-space governance at generation time, unlearning hardened against relearning, in-house pre-synthesis screening, and a documented red-team record. Not a product — a build that sits on the stack you already run.
If your biology model got exfiltrated tomorrow, what's in your control framework besides "we used RLHF refusal"? Save this for whoever owns your DURC or ISO 42001 file.
#AIBiosecurity #PharmaAI #AIGovernance #AISafety #DrugDiscovery