Why AI will not speed up science (yet)
You can see AI everywhere except in the productivity statistics
1. Enter the AI Scientist
A few weeks ago, Nature published two papers that are being described as landmarks in AI-driven biomedical science. Google DeepMind introduced Co-Scientist, a multi-agent AI system that generates research hypotheses, proposes experiments, and analyzes data. FutureHouse, a nonprofit AI research lab, introduced Robin, a similar system that identified a candidate treatment for dry age-related macular degeneration, a leading cause of blindness. Both systems accomplished in hours what would typically take a research team weeks or months of literature review, hypothesis refinement, and data interpretation.
The framing from the developers matches the ambition of the work. FutureHouse CEO Sam Rodriques said that “AI is going to massively accelerate the pace of science” and “increase productivity”, allowing scientists make more discoveries much faster. The company’s more recent system, Kosmos, is claimed to perform the equivalent of six months of research in a single day, on the basis of estimates from early users. These two papers are the latest entries in a genre that the leaders of the field have been developing for several years. Demis Hassabis, who shared the 2024 Nobel Prize in Chemistry for AlphaFold, describes the achievement as the equivalent of “a billion years of PhD time done in one year”. Dario Amodei, the CEO of Anthropic, predicts that AI-enabled biology will “compress the progress that human biologists would have achieved over the next 50 to 100 years into 5 to 10 years”. There are other examples, but the common thread we’ve been hearing from AI leaders is that (1) AI will dramatically accelerate productivity in science, and (2) it will do so soon.
I want to push back on this. I think we’re seeing a pattern we’ve seen before, and it did not play out the way the technology’s champions predicted.
2. The productivity paradox returns
A good framing of this pattern actually comes from economics, and is described in this recent piece in Fortune. Briefly, in 1987, the economist Robert Solow made an observation that became well-known in economics: “You can see the computer age everywhere but in the productivity statistics.” Although firms had been investing heavily in information technology, the measured productivity has not grown. This gap between obvious technological change and absent economic gain became known as the productivity paradox. Economists are now reviving this idea to describe the AI boom. In a study of business executives, about 90% of firms reported that AI had made no measurable difference to productivity over the past three years. As in the 80s, a transformative technology arrives, everyone agrees it will change everything, but the productivity doesn’t actually improve.

The original paradox did eventually resolve, and IT-driven gains arrived in the late 1990s. But the lag was over a decade long, and the benefits were unevenly distributed. Importantly, the gains ultimately required broad structural changes in how organizations worked.
I believe science and academia will follow the same trajectory, and possibly a worse one. The reasons are structural, and they have little to do with how good the AI is.
3. AI does not solve the real bottlenecks
What AI automates is useful: searching the literature, generating hypotheses from published findings, analyzing data, generating models to explain complex systems, editing text. These are all important tasks that consume real effort; however, these tasks are not, for most research programs, the rate-limiting step.
The actual bottlenecks in biomedical science are physical, institutional, and human. Getting a grant funded takes 12 to 18 months from submission to award, and that assumes it is funded at all; most proposals require years of revision and resubmission. Recruiting human participants for a clinical study takes years, and enrolling a cohort with rarer characteristics can take longer still. The physical constants of biology are indifferent to how quickly a model can read: cell lines must be cultured, samples collected and shipped, sequencing runs queued and completed, and mouse colonies bred to the correct genotype over many months, with some disease models yielding a first meaningful readout only after the animals have aged. Experiments fail and have to be repeated, protocols need optimization, reagents and specialized instruments have lead times, and shared core facilities have waiting lists. Layered on top of all this is an institutional apparatus that runs on its own clock: IRB approval for human subjects and IACUC approval for animal work, biosafety approvals for pathogens, data use agreements and material transfer agreements with legal back-and-forth between universities, and controlled-access approvals for protected genomic data all take many months. Another bottleneck is human: training a scientist to independence takes years, recruitment and hiring is slow, and visa and immigration delays for international trainees add a layer that can take months to years of bureaucracy. Peer review adds still more delay: revisions can take years, and even finding scientists willing to review a paper now takes months. All these processes, physical and institutional and human, set the real pace of research. Currently, AI improves almost none of them.
This is a well-known principle in engineering and operations research, sometimes called Amdahl’s law: the speedup achievable by optimizing one component of a system is limited by the fraction of total time that component represents. So if literature review and hypothesis generation account for 5% of a research project’s timeline, then even a 100X speedup in those tasks reduces the total project duration by less than 5%. The bottleneck just moves downstream, to the parts of the process that were already slow.
The most instructive illustration of this comes not from a skeptic but from a company built on the opposite premise. A couple of weeks ago, a San Francisco startup called Capable announced that it has built a laboratory capable of going from AI-led drug design to experimental data in 24 hours:
If you scroll through the X thread, what is striking is the description of how they achieved this. Their platform consists of six modules, only one of which is AI-related (drug design); the others are machine learning infrastructure, rapid synthesis, quality control, cell testing, and mouse testing. To compress the timeline they built an in-house in vivo facility, obtained pre-approved protocols, and put veterinarians and in vivo researchers on standby. They are now hiring automation engineers to integrate liquid handlers and robotic arms in order to, in their own words, unlock bottlenecks.
I do not raise this as a criticism of Capable, which appears to understand the problem clearly and is spending money and effort on the actual bottlenecks. I raise it because a company whose entire thesis is that AI will accelerate drug discovery has concluded that the way to accelerate drug discovery is to restructure the wet lab. The speedup did not come from the AI, but from a vivarium, pre-approved protocols, veterinarians on standby, and wet-lab pipelines re-engineered to run faster.
4. To accelerate science with AI, institutions must change
I am not arguing against using AI in research. The tools are powerful and I do believe will eventually transform science. My issue is with the expectations around AI.
When AI companies and prominent commentators claim that AI will dramatically accelerate science, university administrators hear that they can expect more output1 from fewer people. Funders hear that AI might offset the cuts to NIH budgets that are currently suffocating the research enterprise. Trainees hear that reading deeply and analyzing carefully matter less than learning to prompt a model. All of these conclusions are wrong.
To anyone inclined to answer all of this by noting that the original paradox did eventually resolve: it did, but not on its own, and not merely with time. It resolved when organizations restructured how work was done, retrained the people doing the work, and redesigned processes around what computers were actually good at. Organizations did not just wait for gains from technology to arrive; they changed the entire system to remove bottlenecks and make the gains possible.
For science, this points to a clear and somewhat unglamorous conclusion. If the real bottlenecks are the ones described above, then the way to capture gains from AI is to work on those bottlenecks directly. That is largely the job of institutions rather than individual scientists or better models. Funding agencies can shorten the grant cycle and reduce the administrative burden that consumes so much of a researcher’s time. Universities can speedup IRB and IACUC review, streamline the contracting and data use agreements, and invest in the shared infrastructure, core facilities, and technical staff. Journals and funders can rethink a peer review system that is already overloaded and is only coming under more and more strain. The companies most invested in AI for science understand this; it is time for academic institutions to reach the same conclusion.
This is the problem that no model will solve for us. AI that analyzes a dataset 10X faster moves a project forward only marginally, but a proposal funded within 4 weeks of submission cuts years of wasted time. The same is true for a manuscript reviewed and published within 2 weeks of submission. What would transform science are institutions that are willing to do the work and solve the real bottlenecks. The gains from AI tools will eventually be substantial, but they will only be unlocked by deliberately fixing the surrounding system.
Defining “output” or “productivity” in science is it’s own huge question that I will not get into here and may be worth its own post. I’ll just note that counting publications was never a good proxy for scientific productivity, and AI makes it even worse, as it lowers the cost of producing papers without lowering the cost of producing knowledge. The growing gap between new papers and new knowledge, and what it does to how we measure and reward science, is not something we are anywhere close to resolving.




One detail in the Robin paper supports your argument almost in passing. Robin proposed photoreceptor outer segments and primary RPE; the lab used beads and ARPE-19 instead, for availability and speed. The laboratory’s constraints changed the proposed experiment before testing began.
The ranking that shaped which candidates reached the screen was built on different grounds: scientific rationale, pharmacological profile, and the supporting literature. It was assessed against expert preference and its own consistency.
If candidate generation grows faster than experimental capacity, the tested fraction will shrink, and the ranking will decide more of what ever meets an assay.
That can be tested within the same screen. Carry two or three lower-ranked candidates alongside the top set and repeat the comparison across screens. The difference in hit rates would show whether the ranking adds experimental enrichment beyond its agreement with expert judgment.
Solid take