The Chemist Who Did Not Live to See the Nobel Prize

The Chemist Who Did Not Live to See the Nobel Prize

- 21 mins

The economic case for public data infrastructure in the age of AI


The database that won the 2024 Nobel Prize in Chemistry started with seven entries and a man who would not live to see the eighth.

In October 1971, Walter Hamilton at Brookhaven National Laboratory established the Protein Data Bank [1]. It held seven crystal structures. Hamilton had championed the idea at crystallography meetings throughout the late 1960s, arguing that structural data was too valuable to sit in individual labs. He died in 1973, at 41. He never saw what his archive would become.

The accumulation was patient. Seven structures grew to 23 by 1976, a few hundred by the late 1980s, 10,000 by 2000, more than 170,000 by 2020 [1]. Each was painstakingly determined by X-ray crystallography or NMR, deposited voluntarily, made freely available. Fifty-three years of quiet, publicly funded work.

Then the payoff. Since 1994, a biennial competition called the Critical Assessment of protein Structure Prediction (CASP) had benchmarked the field, judging computational predictions against experimentally determined PDB structures [46]. In 2020, DeepMind’s AlphaFold 2 won CASP14 with a median accuracy score of 92.4 out of 100, solving the problem outright, trained on those 170,000+ structures [2]. In 2021, DeepMind partnered with EMBL-EBI to release predictions for over 200 million proteins under a CC-BY 4.0 licence [3]. In October 2024, Demis Hassabis and John Jumper received the Nobel Prize in Chemistry [4].

Seven structures to a Nobel Prize. The AI revolution in structural biology was built entirely on decades of publicly funded, openly shared data.

EMBL-EBI’s latest impact assessment shows a 73-to-1 return on investment from open biological data infrastructure [5]. That substrate is now under threat. Every AI company training models on PubMed, ClinVar, GenBank, the cancer atlases, and the imaging archives pays nothing for it. The institutions that produce this data face budget pressures that could erode the foundation on which the entire AI-in-health enterprise depends.


One Per Cent of the Budget, All of the Drugs

The numbers for public biomedical research are not speculative. They are the most thoroughly documented return on public investment in the history of American science.

The National Institutes of Health contributed to published research associated with 354 of 356 drugs approved by the FDA between 2010 and 2019 [6]. That is 99.4%. Not a rounding error. Virtually every drug your doctor prescribes traces back to publicly funded science. Every dollar invested in NIH research generated $2.56 in economic activity in FY2024, supporting 407,782 jobs and producing $94.58 billion in new economic output across all fifty states [7].

The $3.8 billion Human Genome Project generated an estimated $965 billion in economic activity, a return of roughly 65-to-1 [8]. That return did not occur due to a proprietary platform. It happened because of the Bermuda Principles, a one-page agreement adopted in February 1996 requiring all sequence data to be released within 24 hours [9]. No embargoes. No patents on raw sequence. No permission required. Open protocols, funder-enforced. The Bermuda Principles cost essentially nothing to implement and generated nearly a trillion dollars.

The pattern holds internationally. EMBL-EBI, Europe’s primary open biological data infrastructure, spends GBP 75 million a year to run its data services and generates an estimated GBP 5.5 billion in value, a return of roughly 73-to-1 at the midpoint of independent estimates. Half a million researchers across 190 countries depend on it. These are conservative figures: the assessors excluded secondary users entirely and noted the true user base may be double the headline count [5].

Consider three therapies that define modern medicine.

The COVID-19 mRNA vaccines were developed in record time, but they rested on decades of NIH-funded basic research stretching back to the late 1980s. Katalin Kariko and Drew Weissman’s work on modified nucleosides, funded in part by NIH grants, made mRNA vaccines viable. They received the 2023 Nobel Prize [10]. When the SARS-CoV-2 genome was sequenced in January 2020, researchers at NIH’s Vaccine Research Centre applied their prefusion-stabilising approach to the spike protein and partnered with Moderna within 24 hours [11]. The science was ready because the funding had been patient and focused on the long term.

GLP-1 receptor agonists, the drug class behind semaglutide (Ozempic, Wegovy) and tirzepatide (Mounjaro), trace to NIH-funded research on Gila monster venom in the early 1980s. Dr Jean-Pierre Raufman discovered unusual molecules in the venom that affected the pancreas. This led to the identification of exendin-4 [12], which became the foundation for what is now projected to be a $100 billion annual drug class by 2030 [13]. The gap between Raufman’s Gila monster research and Ozempic spans more than forty years. No venture capitalist would have funded that timeline.

CAR-T cell therapy, among the most transformative cancer treatments of the past decade, emerged from NCI-funded basic immunology research beginning in the 1970s. Steven Rosenberg’s patient work at the National Cancer Institute, from early T-cell activation experiments to the first successful CAR-T treatment for lymphoma in 2009, spanned roughly 40 years of sustained public investment [14]. That patient remains cancer-free.

The NIH budget of $48 billion represents approximately 1% of total US healthcare expenditure [15]. One per cent of the spend, underwriting 99.4% of the drug pipeline.


When Scientists Are Told Not to Say “mRNA”

This is not a uniquely American problem, but the United States illustrates it most starkly. The NIH FY2026 budget proposal would reduce funding by approximately 40%, to an inflation-adjusted level billions of dollars lower than at any point in the past twenty-five years [16]. Early FY2026 data showed grant award rates collapsed by 54% compared to the previous year [17]. A proposed indirect cost cap of 15%, far below most institutions’ negotiated rates of 40-60%, would have withdrawn roughly $4 billion annually from research infrastructure [18]. A court ruled the cap unlawful, but the disruption to institutional planning was immediate. In the UK, UKRI faces real-terms funding pressure after years of flat settlements. Across the EU, the next Horizon framework is being shaped by competing demands for AI investment and defence spending. Basic research funding is under strain everywhere.

Reports have emerged of researchers being advised to avoid certain terminology in grant applications, including references to mRNA vaccine technology [19], the same technology that produced the most significant public health intervention of the century. When the language of science is shaped by funding climate rather than evidence, something has gone wrong.

The proximate causes of these funding pressures vary by country: fiscal consolidation, shifting political priorities, competing demands from defence and industrial strategy. I am not arguing that AI hype caused the cuts. But I am arguing that the AI narrative functions as a permission structure. It makes the cuts seem less dangerous by implying that private-sector AI investment will pick up the slack. It lowers the perceived cost of losing public research capacity. When a politician can gesture at $725 billion in Big Tech AI capital expenditure and say “the private sector has this covered,” it becomes easier to cut $20 billion from basic research.

The effect is the same as the 2005 weather data bill [47], just achieved differently. Not a single piece of legislation to privatise public data, but a slow defunding of the institutions that produce it, accompanied by a vague assurance that AI will fill the gap.

It will not.


How Many AI Drugs Has the FDA Approved?

I should be honest about where AI works. I run data platforms for a living. I see the tooling every day.

AlphaFold is a genuine scientific revolution, proven at CASP14, the open benchmarking competition that has tested protein structure prediction since 1994 [46]. It has predicted structures for over 214 million proteins, been cited in more than 35,000 papers [20], and won the 2024 Nobel Prize in Chemistry [4]. AlphaFold researchers see a 40% increase in their submission of novel experimental protein structures [21]. It accelerates science. It does not replace it. And it was trained on the Protein Data Bank, which contains decades of experimentally determined structures funded by NIH and equivalent agencies worldwide. Without fifty years of publicly funded structural biology, and without the open competition framework that let researchers benchmark against shared data, AlphaFold would have had nothing to learn from. AlphaFold is not an outlier. Across the life sciences, 42% of researchers who use EMBL-EBI’s open data resources report training or evaluating AI and ML models on that data. As one developer of a machine-learning tool for cryo-EM put it: “Without EBI, there would be no ModelAngelo” [5]. Cut the public data infrastructure, and you cut the AI that depends on it.

AI weather forecasting has been transformative. Google DeepMind’s GenCast outperforms the European Centre’s ensemble forecast on 97.2% of verification targets [22]. In December 2025, NOAA deployed three AI-driven global weather models into operational use [23]. This is real progress. It was also trained on ERA5, whose backbone is physics-based numerical weather prediction that has been developed over decades with public funding.

In drug discovery, AI-discovered molecules achieve 80-90% success rates in Phase I trials, well above the historical average of 52% [24]. AI produces safer molecules. That matters.

Now the other side of the ledger.

Zero AI-discovered drugs have received FDA approval. Over 200 candidates are in clinical development [25]. None has completed the regulatory process. The most advanced, Insilico Medicine’s rentosertib, has completed Phase IIa [26]. It is years from approval, if it gets there at all.

Phase II success rates for AI-discovered drugs, approximately 40%, are indistinguishable from historical averages [24] [25]. Phase II tests efficacy: whether the drug actually treats the disease. This is where most programmes fail, and AI has not cracked it. AI produces molecules that are safe. It has not yet been demonstrated that it can produce molecules that work.

IBM spent over $4 billion on Watson Health, employed 7,000 staff at its peak, and sold the unit for approximately $1 billion [27]. Watson for Oncology recommended bevacizumab for a patient with severe bleeding, a potentially fatal error [28]. MD Anderson spent $62.1 million on a Watson-powered system over four years without treating a single patient [29].

The Epic Sepsis Model, deployed across hundreds of hospitals serving 54% of US patients, missed 67% of sepsis cases in external validation, with an AUC of 0.63, compared with the internally reported 0.76-0.83 [30]. A follow-up study found that its predictive accuracy dropped to 53% when restricted to data collected before a blood culture was ordered [31]. The model was largely cueing on clinician suspicion rather than providing early warning.

Google created and disbanded its Health division in three years [32]. Amazon shut down Haven, Amazon Care, and Halo within roughly two years each [33]. The Big Tech health AI graveyard is large and expensive.

Daron Acemoglu, the 2024 Nobel Laureate in Economics, estimates AI’s impact on GDP at 0.53-0.93% over a decade [34]. Goldman Sachs titled a 2024 report “Gen AI: Too Much Spend, Too Little Benefit?” [35]. The gap between Acemoglu’s evidence-based estimates and industry projections of 7-10% GDP growth is approximately one order of magnitude. Someone is wrong.


The IETF Runs the Internet on $14 Million a Year

The argument against cutting NIH is not merely scientific. It is economic, and the comparison is lopsided.

Microsoft, Alphabet, Meta, and Amazon committed over $320 billion in AI capital expenditure in 2025, with plans for up to $725 billion in 2026 [36]. NIH’s entire annual budget is $48 billion [15]. The AI sector spends 15 times the NIH budget each year on infrastructure that has yet to produce a single approved drug.

A single ChatGPT query consumes approximately 0.3-0.42 watt-hours of electricity [37]. At one billion daily queries, that is 310 GWh annually, enough to power 29,000 homes. Training GPT-4 consumed an estimated 50 GWh, enough to power San Francisco for three days [38]. The Stargate initiative aims to spend $500 billion on building data centres that may each require 5 GW of power, more than New Hampshire’s total demand [36].

Now consider the cost of open data infrastructure, and the returns it generates.

The cleanest precedent is GPS. On 1 September 1983, Soviet fighters shot down Korean Air Lines Flight 007 after it strayed into restricted airspace. 269 people died. Within weeks, Ronald Reagan announced that GPS, then a military system, would be made available for civilian use. On 1 May 2000, Bill Clinton ended Selective Availability, improving civilian accuracy tenfold [39]. The signal is free: no subscription, no API key, roughly $2 billion a year to maintain. In 2019, NIST and RTI International estimated that GPS had generated $1.4 trillion in cumulative economic benefits to the US private sector since the 1980s [40]. Ninety per cent of that value accrued after 2010. Loss of GPS would cost approximately $1 billion per day. Every ride-hailing app, precision agriculture system, and supply chain tracker depends on it.

GPS is what happens when a government builds infrastructure as a protocol rather than a platform: a public investment generating orders-of-magnitude private-sector returns, invisible to its beneficiaries, and perpetually at risk of restriction or defunding because it looks like a cost centre.

The same pattern holds across every open infrastructure that has endured. The IETF, the body that governs the protocols running the entire internet, operates on roughly $14 million a year [41]. TCP/IP has survived seven US presidents. The Bermuda Principles fit on a single page and generated $965 billion in economic activity [8] [9]. OHDSI, which runs federated analytics on 974 million patient records from 544 data sources in 88 countries, operates with minimal central funding and no central data warehouse [42]. GenBank has survived nine NIH directors over 44 years.

These are not platforms. They are protocols: open standards, funder-enforced, with decentralised implementation. They cost orders of magnitude less than platforms, and they last orders of magnitude longer. GPS, the IETF, GenBank: each was built with public money, each generates returns its funders never imagined, and each faces periodic threats from people who see only the cost. All of Us has already seen its funding swing from $541 million to $158 million, a 71% reduction [43]. N3C runs on a $60 million Palantir contract [44]. When the contract ends, the pipelines are stranded. But GenBank endures, and GPS endures, because they function as infrastructure, not as programme-specific portals.

Every dollar NIH has invested in N3C, All of Us and dozens of other platforms is at risk when budgets shift or vendors change. Protocols protect those investments by making them interoperable. The data, the governance rules, and the researcher’s credentials survive even if the platform does not.

I should know. As Chief Data Officer at Sage Bionetworks, I run one of those platforms. Synapse hosts over 4 petabytes of data for major NIH programmes. Through the DREAM Challenges, we have hosted more than 60 open benchmarking competitions across cancer, neurodegeneration, and drug discovery, the same open-competition model that produced AlphaFold applied to the diseases that matter most. I am proud of what we have built. And I know that without the protocol layer connecting our platform to others, every one of those data silos is a DC power station lighting up one city block, while the grid remains unbuilt.


The Most Dangerous Cut Is the One You Cannot See

Here is what is happening, stated plainly. Governments around the world are allowing pressure to build on the research infrastructure that produced 99.4% of approved drugs, that generated $2.56 for every dollar invested, that enabled mRNA vaccines, CAR-T therapy, and GLP-1 agonists, while directing investment toward an AI buildout that has produced zero approved drugs, that costs fifteen times as much, and whose most prominent health ventures have been expensive failures.

This is not a technology debate. It is an economics question with a clear answer.

The Protein Data Bank started with seven structures because one chemist believed data was too valuable to hoard. Fifty-three years later, a Nobel Prize proved him right. The parallel to today is exact: open data built the AI. Defund the data, and you defund the AI.

Biomedical research data does not yet have the public visibility it deserves. Most people do not know that their doctor’s prescriptions trace back to NIH-funded science. They do not know that Ozempic started with a Gila monster. They do not know that the mRNA technology saving lives during COVID was built on thirty years of patient, publicly funded work that no private investor would have underwritten.

The genomics bubble of 1998-2003 offers a precise precedent. Celera Genomics peaked at $247 per share in March 2000 and crashed to $14 by January 2003, a 94% decline [45]. The 74 companies in the genomics sector traded at an average of 25% of their highs. The publicly funded Human Genome Project outlasted them all. Celera’s proprietary database could not compete with GenBank, which was free. Open data won.

The current AI investment cycle will follow the same pattern. Not because AI is worthless, but because the hype exceeds the evidence by an order of magnitude, and the evidence sits on a foundation of public science that is being defunded even as the hype accelerates. When the cycle corrects, and it will, the companies that survive will be those that built on the public data commons. If that commons has been gutted, there will be nothing to build on.

The $3.8 billion Human Genome Project generated $965 billion because the Bermuda Principles made the data open. The $48 billion NIH budget underwrites 99.4% of approved drugs because the science is public. The IETF’s $14 million a year helps maintain the protocols used by billions because its standards are open.

Protocols, not platforms. Open infrastructure, not proprietary extraction. Patient public investment, not speculative private capture. These are not ideological preferences. They are what the evidence shows.

Somewhere right now, a graduate student funded by an NIH training grant is running an experiment that will not produce a clinical result for twenty years. That is not a waste. That is how mRNA vaccines happen. That is how CAR-T therapy happens. That is how a lizard in the Arizona desert becomes the most important drug class of the decade.

Cut her funding, and we will never know what we lost.


Susheel Varma is Chief Data Officer at Sage Bionetworks and the author of The Current Wars of Biomedical Data. He was previously CTO at Health Data Research UK, where he co-authored national TRE guidance, and Head of AI and Data Science at the UK Information Commissioner’s Office.

References

[1] Protein Data Bank. “History.” RCSB PDB. https://www.rcsb.org/pages/about-us/history

[2] Jumper, J. et al. (2021). “Highly accurate protein structure prediction with AlphaFold.” Nature, 596, 583-589.

[3] Varadi, M. et al. (2022). “AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models.” Nucleic Acids Research, 50(D1), D439-D444.

[4] NobelPrize.org. “The Nobel Prize in Chemistry 2024.” https://www.nobelprize.org/prizes/chemistry/2024/summary/

[5] Frontier Economics (2026). “Open Data: The Value and Impact of EMBL-EBI Data Resources.” Report commissioned by EMBL-EBI.

[6] Cleary, E.G. et al. (2018). “Contribution of NIH funding to new drug approvals 2010-2016.” PNAS, 115(10), 2329-2334. Extended analysis through 2019 by same group.

[7] United for Medical Research (2024). “NIH’s Role in Sustaining the U.S. Economy: FY2024.”

[8] Battelle Technology Partnership Practice (2011). “The Impact of Genomics on the U.S. Economy.”

[9] Maxson Jones, K., Ankeny, R.A. & Cook-Deegan, R. (2018). “The Bermuda Triangle: the pragmatics, policies, and principles for data sharing in the history of the Human Genome Project.” Journal of the History of Biology, 51(4), 693-805.

[10] NobelPrize.org. “The Nobel Prize in Physiology or Medicine 2023.” https://www.nobelprize.org/prizes/medicine/2023/summary/

[11] Corbett, K.S. et al. (2020). “SARS-CoV-2 mRNA vaccine design enabled by prototype pathogen preparedness.” Nature, 586, 567-571.

[12] Eng, J. et al. (1992). “Isolation and characterization of exendin-4, an exendin-3 analogue, from Heloderma suspectum venom.” Journal of Biological Chemistry, 267(11), 7402-7405.

[13] Goldman Sachs Research (2024). “The GLP-1/Obesity Market.”

[14] Rosenberg, S.A. & Restifo, N.P. (2015). “Adoptive cell transfer as personalized immunotherapy for human cancer.” Science, 348(6230), 62-68.

[15] Congressional Research Service (2025). “National Institutes of Health (NIH) Funding: FY1996-FY2026.”

[16] Nature (2025). “NIH budget cuts: what researchers need to know.”

[17] Lauer, M. (2025). NIH Office of Portfolio Analysis and Reporting (OPERA) data, as reported in Nature.

[18] Nature (2025). “US indirect-cost cap for research grants ruled unlawful.”

[19] Science (2025). “Researchers told to avoid certain terms in NIH applications.”

[20] Google DeepMind (2024). “AlphaFold impact.” https://deepmind.google/technologies/alphafold/

[21] Callaway, E. (2024). “AlphaFold is already changing the world.” Nature.

[22] Price, I. et al. (2025). “Probabilistic weather forecasting with machine learning.” Nature, 637, 84-90.

[23] NOAA (2025). “NOAA deploys AI-driven global weather models.”

[24] Boston Consulting Group (2024). “AI-Driven Drug Discovery: Early Clinical Signals.”

[25] Jayatunga, M.K.P. et al. (2024). “How successful are AI-discovered drugs in clinical trials?” Nature Biotechnology (commentary).

[26] Insilico Medicine press releases (2024-2025). Rentosertib (ISM001-055) Phase IIa clinical trial results.

[27] Strickland, E. (2019). “How IBM Watson Overpromised and Underdelivered on AI Health Care.” IEEE Spectrum.

[28] Ross, C. & Swetlitz, I. (2018). “IBM’s Watson supercomputer recommended ‘unsafe and incorrect’ cancer treatments, internal documents show.” STAT News.

[29] Herper, M. (2017). “MD Anderson Benches IBM Watson In Setback For Artificial Intelligence In Medicine.” Forbes.

[30] Wong, A. et al. (2021). “External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients.” JAMA Internal Medicine, 181(8), 1065-1070.

[31] Lyons, P.G. et al. (2022). “Factors associated with variability in the performance of a proprietary sepsis prediction model across hospitals.” Critical Care Medicine.

[32] Wired (2021). “Google Health Is Shutting Down.”

[33] CNBC (2023). “Amazon is shutting down Amazon Care.”

[34] Acemoglu, D. (2024). “The Simple Macroeconomics of AI.” NBER Working Paper 32487.

[35] Goldman Sachs (2024). “Gen AI: Too Much Spend, Too Little Benefit?”

[36] Bloomberg (2025). Company earnings reports and capital expenditure guidance for Microsoft, Alphabet, Meta, and Amazon.

[37] IEA (2024). “Electricity 2024: Analysis and forecast to 2026.”

[38] Epoch AI (2024). “Estimating training compute of notable ML models.”

[39] GPS.gov. “Selective Availability.”

[40] RTI International for NIST (2019). “Economic Benefits of the Global Positioning System to the U.S. Private Sector.”

[41] IETF Annual Report / Internet Society (ISOC) financial statements.

[42] OHDSI.org. “The OHDSI Network” (2025). https://www.ohdsi.org/

[43] Science (2025). “All of Us funding slashed.”

[44] NCATS reporting on the N3C Palantir contract.

[45] Mardis, E.R. (2011). “A decade’s perspective on DNA sequencing technology.” Nature, 470, 198-203.

[46] Moult, J. et al. (1995). “A large-scale experiment to assess protein structure prediction methods.” Proteins, 23(3), ii-v.

[47] S. 786, “National Weather Service Duties Act of 2005,” 109th Congress.

Originally published on LinkedIn.

Susheel Varma

Susheel Varma

Chief Data Officer at Sage Bionetworks

comments powered by Disqus