OpenAI Foundation Launches Public Data for Health Program – Unite.AI

0
1
OpenAI Foundation Launches Public Data for Health Program – Unite.AI



OpenAI Foundation Launches Public Data for Health Program – Unite.AI

The OpenAI Foundation introduced Public Data for Health on September 15, 2026, its second science program, with more than $125 million in initial grants to fund the creation and preservation of high-quality scientific datasets made broadly available to researchers.

The program was announced in a post authored by Abhishaike Mahajan and Jacob Trefethen. The foundation launched AI for Alzheimer’s, its first program in Life Sciences and Curing Diseases, in April 2026; where that effort targets one disease affecting many families, the new program is intended to enable progress across the life sciences on many diseases. Life Sciences and Curing Diseases is one of the foundation’s three initial priority programs, alongside AI Resilience and Civil Society and Philanthropy.

In the announcement, the foundation describes scientific data as observations about the world and the foundational input to research and discovery. It contrasts parts of mathematics, where it says AI systems have recently started contributing new knowledge without the collection of new data, with biology, where it says models can increasingly analyze biological information at scale and recover hidden structure even from incomplete evidence. The authors wrote that they expect many remaining breakthroughs in disease prevention and treatment to come from giving intelligent new models more observations of the world, and that some datasets of enormous public value may never be created or shared because no single institution has sufficient incentive or capacity to fund them.

Three Initial Grantee Projects

The initial tranche supports nonprofits and universities and spans molecular, epidemiological, and regulatory layers of data.

OpenADMET will build open datasets, benchmarks, and blinded competitions to test whether AI models can predict how small molecules are absorbed and distributed through the body, work the foundation said is aimed at making drug development more predictable and reducing the failure rate of new drugs. The announcement notes that 90% of drug candidates fail in clinical trials, often because it is difficult to predict how they will be absorbed and move through the body, and it cites AlphaFold2’s use of Protein Data Bank data in the CASP competition as precedent for pairing prediction challenges with high-quality data. “Drug discovery is filled with universal problems that no individual company or academic lab interested in curing a specific disease can solve alone,” said James Fraser, an OpenADMET Governing Board member and professor and chair of bioengineering and therapeutic sciences at UCSF.

CTD Commons will test whether Common Technical Documents from failed or shelved drug programs can be acquired and made openly available for research and analysis. According to the announcement, a CTD compiles an investigational drug’s full journey, spanning animal toxicology, manufacturing details, and correspondence with the FDA, and only a small fraction of that work appears in published papers. Josh Morrison, leading organizer of CTD Commons and president of 1Day Sooner, said the project will reduce duplication and cost across clinical research and maximize its impact.

The University of North Carolina will establish the Initiative for Generative Immunotherapy, creating public, multimodal data aimed at a future in which cancer patients receive rapid, personalized cancer vaccines at diagnosis. The announcement describes current neoantigen cancer vaccines as among the few medicines designed computationally for each patient: a tumor cell is sequenced to predict which tumor-specific targets appear on its surface, but that sequencing is a proxy because measuring the targets directly, and checking how strongly a patient’s immune system reacts to each one, is difficult and expensive. UNC will generate those missing links across hundreds of tumors and multiple cancer types, producing what the foundation described as one of the first training and evaluation datasets in the field, as de-identified public data that researchers worldwide can build on. “Personalized cancer vaccines are finally starting to show signs of clinical efficacy, but still have gaps which might take decades to fill under the traditional model of therapeutic development. High-quality data can help us close those gaps faster,” said Alex Rubinsteyn, an assistant professor of genetics at the UNC School of Medicine.

Connected, Scarce, and Direct Data

Alongside the grants, the foundation published starting hypotheses for the types of data where it believes its support can be most useful: connected data, which follows biology across multiple steps; scarce data, which saves what cannot be recreated; and direct data, which measures biological and clinical states closest to what matters. It said funded datasets will usually carry one or two of these properties, and in rare cases all three.

Under the connected-data rationale, the UCSF grant will let the OpenADMET team measure key molecular properties and interactions for tens of thousands of compounds, link that broad profiling to drug-transport measurements, including structures of transporter proteins bound to selected molecules and functional assays of those proteins, and test a subset of the compounds in human blood-brain barrier models, comparing the results with in vivo animal measurements.

Under the scarce-data rationale, the CTD Commons grant will examine whether records from failed drug development programs can be preserved before companies shut down and the records disappear. The foundation said such documents may help early-stage drug development teams understand what regulators have required of similar products in the short term and, over the longer term, could serve as resources from which machine learning systems discover patterns in why drugs fail or succeed.

Under the direct-data rationale, the UNC team at UNC Lineberger and UNC Health will directly measure tumor cells’ surface proteins and patients’ T cell responses rather than relying on tumor sequencing alone. The foundation said that, if the project succeeds, vaccine candidates that follow could enter human dosing backed by a stronger set of targets.

Data Accessibility and Next Steps

Across all of its grants, the foundation said data created with its support should be broadly available, with individual privacy and consent maintained wherever human data are involved. It said it encourages grantees to publish analyses as preprints, share data regularly rather than only at a project’s end, and work with the users of a dataset to assess its utility on biologically valuable problems.

The foundation said it is actively updating its views as the science and AI progress and invited ideas for the program at [email protected]. It is hiring for four open roles on the Life Sciences and Curing Diseases team.