Artificial intelligence has made enormous progress in language, vision and protein modeling. The next major challenge may be substantially harder: building models that can predict how living cells behave.
On October 7, Biohub announced an expansion of its Virtual Biology Initiative involving the nonprofit research organization, the U.S. Department of Energy, the National Institutes of Health, Google DeepMind, Isomorphic Labs and Meta.
The combined commitment is described as $1.8 billion in funding, data, computing resources and measurement technology. That distinction matters: the figure is not simply $1.8 billion in new cash going into a single fund.
What is a “virtual cell”?
The long-term ambition is to create predictive models capable of representing biological systems well enough to answer some questions digitally.
A sufficiently advanced system could potentially model questions such as:
- How does a cell respond to a drug?
- What happens when a gene is altered?
- How does a cancer cell change after an intervention?
- Which treatment combinations deserve laboratory testing?
- How do cells interact within tissues?
This is still a research goal. Biohub describes the creation of an accurate predictive model of biology as one of the major challenges of the next era of science.

Where does the $1.8 billion come from?
Biohub originally committed $500 million to the Virtual Biology Initiative in April.
The expanded effort adds:
- $300 million from Meta, Google DeepMind and Isomorphic Labs;
- more than $500 million from the DOE over five years;
- NIH coordination of datasets and resources associated with more than $500 million in previous federal investment;
- Biohub’s original $500 million commitment;
- additional data, computing and technology contributions.
The headline figure therefore represents a combination of cash, existing datasets, computing capacity and scientific infrastructure.
Data — not just GPUs — is the bottleneck
Building biological AI requires far more than increasing model size.
A cell contains multiple interacting layers of information:
- DNA and gene expression;
- proteins;
- molecular structures;
- spatial organization;
- cellular states;
- responses to interventions;
- interactions between cells;
- changes over time.
Biohub plans to invest in technologies including cryo-electron tomography, large-scale microscopy and biological engineering to generate much richer datasets.
Why Google, Meta and Isomorphic Labs matter
The partnership combines different strengths.
Google DeepMind brings expertise in AI for science and biological modeling.
Isomorphic Labs brings a strong focus on AI-driven drug discovery.
Meta brings major AI infrastructure, engineering expertise and capital.
Together, the three companies are investing $300 million in the initiative.
The DOE, meanwhile, plans to bring national-laboratory capabilities including exascale computing, X-ray and neutron scattering, electron microscopy and autonomous laboratories.
NIH brings existing biomedical data into the project
The NIH is not starting from scratch.
Its contribution involves coordinating biomedical datasets, repositories and research resources developed through previous federal investments. Biohub will work with NIH to standardize those resources for AI training.
That standardization could be critical.
AI models need datasets that can be compared and combined reliably. Otherwise, differences in laboratory protocols, formats and measurements can become as problematic as the biological signal itself.
Could AI reduce wasted experiments?
If the technology works, part of the drug-discovery workflow could eventually move toward:
hypothesis → AI simulation → experiment selection → laboratory validation.
That does not eliminate physical laboratories.
Instead, the model would help researchers prioritize which experiments are most worth performing.
Biohub argues that accurate predictive models could allow scientists to conduct more biological experimentation digitally and accelerate the path toward new treatments.
What could this mean for the pharmaceutical industry?
A reliable biological foundation model could eventually help pharmaceutical companies:
- identify new biological targets;
- prioritize drug candidates;
- model disease mechanisms;
- predict cellular responses;
- reduce unnecessary experimental cycles;
- integrate multiple biological data types.
But there is an important limitation:
An AI prediction is not proof that a drug works in humans.
Any promising prediction would still require laboratory studies, preclinical testing and clinical trials.
The access question
Biohub presents the project as a resource for the wider scientific community.
However, Reuters reports that commercial backers may receive an initial period of exclusive access to some of the newly generated data before public release.
That raises an important governance question:
How open should biological data be when public resources and private investment are both involved?
The answer could influence how valuable the initiative becomes for universities, startups and researchers outside the founding group.
What happens next?
Biohub is targeting a first dataset within roughly a year, while its broader ambition is to develop functional predictive biology models over a five-year horizon. These are project goals, not already-delivered capabilities.
The crucial test will be whether additional biological data and computing power actually produce models that generalize reliably to real-world biology.
That challenge is significantly harder than simply scaling a language model.
What does this mean for users?
There is no immediate consumer-facing change.
Longer term, if the approach works, the effects could reach users through:
- faster drug discovery;
- better disease understanding;
- more personalized treatments;
- improved therapeutic design;
- potentially lower costs and shorter research cycles.
But the distance between an AI model and an approved medical treatment remains substantial.
The Biohub initiative represents a shift in the AI race.
The objective is no longer simply to have AI read biological information.
It is to make AI capable of modeling and predicting biological behavior.
A universal virtual cell remains an ambitious research target, not a finished product. But the $1.8 billion commitment shows that some of the world’s largest technology, government and scientific organizations now see predictive biology as one of the major AI frontiers of the coming decade.