
It took one flipped sign in a config file to turn a drug-discovery AI into a weapons designer.
In 2022, a pharma research team inverted the reward function in their chemistry model, MegaSyn. In under six hours, on off-the-shelf tools, it produced 40,000 toxic molecules — including analogues of nerve agent VX. The team published it as a warning.
Four years on, most generative chemistry pipelines still share that architecture: the reward function is a setting, not a hardcoded limit. And 2025 made it worse:
→ GeneBreaker jailbroke an open biology model with up to a 60% success rate — never asking for a pathogen, only for proteins "homologous to" a benign reference.
→ 10–50 fine-tuning examples and a few hundred dollars of GPU time can strip safety alignment off an open-weight model, restoring near-frontier capability.
→ Even the last line of defense failed: AI-paraphrased toxin variants slipped past the DNA screening major vendors rely on (Science, 2025).
The uncomfortable truth from our research: refusal training is a behavioral constraint, not a capability one. It teaches a model to say no — not to forget. A config flip, a fine-tune, or a homology prompt walks around it.
Defense that survives an ISO 42001 or EU AI Act audit must live below the prompt layer — in the latent space, the weights, and a standing red team.
The question for your ML lead: does your safety live above the prompt layer — or below it, where all of these attacks operate?
#AIGovernance #Biosecurity