
The number of drug-like molecules that could theoretically exist is estimated at 10^60. The number of atoms in the observable universe is roughly 10^80. A best-in-class pharmaceutical screening campaign tests about one million compounds. That means even the most ambitious physical search covers 0.0000000000000000000000000000000000000000000000000001% of the available space.
We have been studying what happens when R&D organizations try to navigate this vastness using intuition and random screening. The short answer: they burn capital at an alarming rate. The average cost of developing a single new drug has climbed to $2.23 billion per asset. The internal rate of return on pharma R&D hit a 12-year low of 1.2% in 2022. This is not a funding problem. It is a methodology problem, and closed-loop AI discovery is the methodology that replaces it.
The Edison Method is akin to mapping the Pacific Ocean by dipping a teaspoon into the water at random intervals.
Why Does Trial and Error Hit a Wall?
Thomas Edison tested thousands of carbon filaments to find one that glowed. Even Nikola Tesla criticized the approach, noting that "a little theory and calculation" could have saved 90% of the labor. Yet most modern R&D labs still operate on a version of Edison's playbook: synthesize large libraries of compounds, screen them physically, and hope for a hit.
This worked when the targets were simple. It breaks down when the search space becomes astronomical. In materials science, battery development follows a pattern researchers call "punctuated evolution" -- long stretches of incremental optimization interrupted by rare, accidental jumps when someone stumbles on a new material class. The Edison approach handles the optimization phase reasonably well. But finding the next big jump by random sampling in a space of 10^60 or 10^100 candidates is statistically hopeless.
The financial consequences are severe. High-throughput screening (HTS) was supposed to industrialize trial and error. Instead, it often produces high false-positive rates and identifies compounds with poor real-world properties -- solubility issues, toxicity -- because the screening is random rather than rational. The capital required to maintain massive physical compound libraries is prohibitive for all but the largest organizations.
If you are testing new materials physically in a wet lab without simulating them first, you are lighting money on fire.
Why Do Most AI Tools Fail at Real Discovery?
Moving the search in silico -- onto computers -- is the obvious first step. But the market is flooded with AI tools that learn purely from data correlations. These "black box" models work fine within their training data. Ask them to predict properties of a genuinely new class of material, and they fail badly. Experimental data for novel materials is sparse, noisy, and expensive. A model that needs millions of training examples is useless when you have dozens.
Our approach centers on Physics-Informed Machine Learning (PIML) -- embedding the actual laws of physics (conservation of mass, thermodynamics, quantum mechanics) directly into the AI's architecture. Think of it this way: instead of asking the AI to figure out the rules of chess by watching a million games, we hand it the rulebook and let it focus on strategy.
This delivers three practical advantages. PIML models need far less training data because the physical constraints are already built in. They extrapolate reliably into unexplored chemical space because their predictions must obey physical laws. And they cannot "hallucinate" molecules that violate basic chemistry -- unlike standard generative AI, which can produce valid-looking molecular structures that are physically impossible.
We explored these architectural trade-offs in depth in our interactive analysis.
Why Graph Neural Networks Over Large Language Models?
There is a widespread assumption that Large Language Models (LLMs) like GPT-4 or Claude are the universal solution for every AI problem, including scientific discovery. They are not. Molecules are three-dimensional graphs defined by atoms, bonds, and geometric constraints. LLMs treat molecules as text strings, which strips away the 3D structure that determines how a molecule actually behaves.
Benchmarks consistently show that Graph Neural Networks (GNNs) -- models that explicitly represent molecular geometry -- outperform LLMs in property prediction tasks. GNNs are naturally suited to molecules because they process atoms and bonds as nodes and edges in a graph, preserving the spatial relationships that matter.
The architecture we advocate is a hybrid. LLMs serve as the reasoning layer -- parsing scientific literature, extracting synthesis recipes, designing experimental protocols. GNNs and PIML models do the heavy lifting of property prediction, inverse design, and stability analysis. The LLM orchestrates; the specialist models calculate.
How Does a Self-Driving Lab Actually Work?

The real transformation happens when prediction and experimentation become a continuous cycle. A closed-loop discovery lab -- sometimes called a self-driving lab -- connects the AI directly to robotic synthesis and characterization equipment. The AI predicts a candidate. Robots synthesize it. Sensors measure its properties. The results feed back into the AI, which updates its understanding and selects the next experiment. No human in the loop between cycles.
Berkeley's A-Lab synthesized 41 novel inorganic materials in 17 days. That would take human researchers months or years.
The A-Lab at Lawrence Berkeley National Laboratory demonstrated this at scale. Its AI agent autonomously corrected synthesis recipes when reactions failed, analyzing X-ray diffraction patterns, adjusting precursor ratios -- the relative amounts of each starting ingredient fed into a reaction -- and retrying, achieving a 71% success rate for novel materials, far exceeding human intuition-driven synthesis.
The mathematical engine behind this is Bayesian Optimization, which balances two competing goals: exploring unknown regions of chemical space and refining areas around known hits. Unlike random screening, it uses uncertainty -- tracking not just what the model predicts, but how confident it is -- to select the experiment that will teach the system the most. The ANI-1x potential, developed using this active learning approach, achieved the accuracy of expensive quantum mechanical calculations while using only 10% of the data required by random sampling.
Our team has also adopted Cost-Informed Bayesian Optimization, which factors in the actual price of each experiment. If two candidates offer similar scientific value but one requires a $5,000 reagent and the other a $50 reagent, the system chooses the cheaper path. Studies show this approach can reduce optimization costs by up to 90% while reaching the same outcomes.
The hardest part of building a self-driving lab is not the AI. It is getting the AI to talk to the hardware. Spectrometers, liquid handlers, and hotplates from different vendors speak different proprietary languages. Without a universal translation layer, the AI is a brain in a jar. The SiLA 2 standard (Standardization in Lab Automation) solves this by treating every instrument as a microservice. The AI sends a high-level command -- "Dispense 5ml" -- without needing to know the low-level serial port instructions of the specific machine. A 20-year-old HPLC (High-Performance Liquid Chromatography) system -- essentially a precise instrument for separating and analyzing chemical mixtures -- can be wrapped in a SiLA 2 driver and participate in a cutting-edge autonomous loop.
Before any physical robot moves, the experiment runs first in a digital twin -- a virtual replica of the lab. Thousands of virtual experiments validate the logic, timing, and collision paths of a protocol. This prevents costly crashes and lost samples, and provides real-time anomaly detection by comparing sensor data against the twin's predictions.
For the full technical methodology behind these optimization strategies, see our detailed research.
What About the Failed Experiments?
In traditional labs, negative results get buried in notebooks. In closed-loop discovery, failed experiments are some of the most valuable data the system collects.
To distinguish a viable drug from a non-viable one, the AI needs to know what failure looks like. Negative data sharpens the model's decision boundaries -- the lines the AI draws between "this will work" and "this won't" -- and prevents generative models from predicting reactions that are thermodynamically impossible, meaning they violate the fundamental energy rules of chemistry and could never occur in nature. That accumulated understanding of where the dead ends are -- a kind of permanent map of the failure landscape -- becomes durable intellectual property. The organization never wastes resources exploring those paths again.
Every failed experiment in a closed loop is a permanent lesson. In a traditional lab, it is a forgotten notebook entry.
The Economics Favor Simulation -- So Where Do You Start?

The return-on-investment case is straightforward. Traditional HTS burns operating expenditure on consumables, reagents, and personnel time. Active learning reduces the number of physical experiments needed to find a hit by 10x to 100x. Autonomous equipment runs 24/7 at near-100% utilization, compared to the 30-40% typical of human-staffed labs.
Speed matters as much as cost. AI-first biotech companies like Exscientia have moved AI-designed small molecules into Phase I trials in roughly 12 months, compared to the industry average of 4-5 years. Insilico Medicine took an AI-discovered fibrosis candidate from target discovery to preclinical candidate in under 18 months at a fraction of typical cost. Every day saved in R&D is an extra day of market exclusivity.
The hidden cost is the one most organizations never measure: every dollar spent physically testing a material that could have been ruled out by simulation is a dollar not spent on a viable candidate. With pharma R&D failure rates hovering around 90%, the ability to fail virtually instead of physically is the single largest lever for improving returns.
The shift from Edisonian to closed-loop does not require rebuilding your entire lab on day one. It starts with a simple discipline change: simulate before you synthesize. Run computational screens to filter candidates before committing reagents. Capture negative results as systematically as positive ones. Evaluate where Bayesian optimization could replace random experimental design.
The search space of 10^60 molecules is not an insurmountable abyss. It is a landscape to be navigated -- but only with the right map.
If your R&D organization is exploring how to make this transition, we would welcome the conversation. The methodology has moved well past theory and into production results.