The US government and major technology companies Meta Platforms and Alphabet’s Google are joining nonprofit Biohub in a $1.8 billion effort to create large open datasets for training artificial intelligence models for biological research, as the technology industry and government seek to accelerate the use of AI in drug discovery and life sciences.
The initiative brings together $300 million from Meta, Google DeepMind and drug discovery company Isomorphic Labs, more than $500 million from the US Department of Energy and more than $500 million in earlier federal funding coordinated by the National Institutes of Health.
The commitments come on top of the $500 million Biohub committed to the project in April, bringing the total investment in the effort to $1.8 billion.
Register for the next Tekedia Mini-MBA.
Register for Tekedia AI in Business Masterclass.
Join Tekedia Capital Syndicate and co-invest in great global startups.
The funding will support Biohub’s Virtual Biology Initiative, an ambitious effort to generate biological data at a scale far beyond what scientists currently have available and use it to train predictive AI models.
The objective is to understand how cells respond to changes across a much larger range of biological conditions. Researchers hope that models trained on the resulting datasets will eventually be able to predict biological responses and help shorten drug development timelines that currently can stretch across years.
“Biology has been just sort of a clever discovery-based science until this point,” said Dr. Priscilla Chan, co-founder of Biohub, in an interview. “We have always held this as a community asset, not just for one group, so that it can build upon itself over time.”
The initiative reflects a growing realization across the AI and pharmaceutical industries that capable models require enormous quantities of high-quality domain-specific data. In biology, much of the challenge is not simply developing better algorithms but generating enough consistent experimental data for those models to learn from.
Biohub’s effort is aimed at closing a major gap between the amount of biological data currently available and what researchers believe will be necessary to build accurate predictive models.
Alex Rives, Biohub’s head of science, said existing cell datasets contain hundreds of millions of cells. Predictive models capable of accurately representing biological systems will eventually require billions and potentially trillions of cells’ worth of data.
“We need to capture the language of biology, we need to capture the language of the cell. And that doesn’t exist today,” Rives said.
The data will be generated using technologies including spatial transcriptomics, which maps molecular activity within intact tissue, as well as experiments that record how cells respond to changes in their surrounding environment.
A major part of the effort involves generating this information in a coordinated manner. Many of the relevant datasets either do not exist at the required scale or have been produced using different methods that make them difficult to combine for AI training.
The US Department of Energy plans to invest more than $500 million over five years in laboratory measurement, modelling and computing. The NIH will coordinate datasets and repositories created using more than $500 million in previous federal funding, with Biohub standardizing that material for AI training.
The government-funded work will be made available without the restrictions attached to privately funded datasets, according to Rives.
Private companies contributing funding, however, will receive an initial period of exclusive access to the datasets they help finance.
“With commercial funders we have embargo periods where there’s a period of time where the groups can work on the data, and then it becomes available as a public scientific resource,” Rives said.
That arrangement gives Biohub a mechanism for attracting private capital while maintaining its longer-term objective of creating a public scientific resource. Companies receive an early opportunity to work with the data, while the broader research community eventually gains access.
Biohub plans to approach pharmaceutical companies and philanthropic organizations for additional support.
AI Competition Moves Deeper Into Biology
The initiative comes as AI companies increasingly move beyond general-purpose models and into specialized scientific applications.
Rives said work that would ordinarily take decades could potentially be compressed into five years through the initiative. Biohub expects to have its first major dataset ready in about a year and aims to develop accurate predictive models within five years.
If successful, the project could give researchers a new way to study biological systems. Instead of relying primarily on individual experiments and retrospective analysis, scientists could use models trained on large-scale experimental datasets to predict how cells might behave under previously untested conditions.
But that is expected to have implications for drug discovery, where researchers currently face a long process of identifying potential compounds, testing their biological effects, and progressing promising candidates through increasingly expensive stages of development.
The project’s scale also underlines why AI-driven biology is becoming an area of competition among technology companies.
Anthropic has expanded its biology research through a wet lab, while the OpenAI Foundation has launched a grant programme of more than $125 million to support biological and medical datasets for AI research. The competing efforts point to a shift in the economics of AI research. Access to computing remains important, but high-quality proprietary and public datasets are becoming an equally important resource for building specialized models.
The problem is considered acute for biology because biological systems are highly complex and experimental conditions can produce dramatically different outcomes. A model trained on limited or inconsistent observations may struggle to generalize to new biological environments.
Biohub’s approach is therefore centered on generating data systematically rather than relying solely on datasets that already exist.
The $1.8 billion commitment also represents a significant public-private investment in AI-enabled scientific research. Meta, Google DeepMind and Isomorphic Labs are providing commercial capital and expertise, while US government agencies are funding measurement, computation and the organization of existing scientific data.
The eventual release of the datasets could also create a foundation that researchers outside the project’s original partners can use to develop new biological models and applications.
However, execution has been fingered as the immediate challenge. Generating billions or trillions of observations is substantially different from assembling existing datasets, and ensuring that the resulting data is accurate, standardized, and useful for predictive modelling will be critical.
Biohub is betting that a concentrated five-year effort can compress decades of biological data generation into a period short enough to match the rapid development cycle of modern AI. If the initiative succeeds, the resulting datasets could become a foundational resource for AI-driven biology, giving researchers and companies a much larger experimental base from which to build models of cells, disease and potential treatments.



