Why AI Agents won't cure diseases
On the Renaissance of the Polymath
In December 2023, ETH Zurich invited me to speak about how generative AI would change scientific research. At the time, scientists were mainly discussing LLMs for writing papers and searching the scientific literature. The term “AI agents” had not become the default of every AI pitch.
By then, I had spent years at Palantir working at the intersection of AI deployment and life sciences. Working with companies turning research into medicines, I had seen a different bottleneck. Scientific, clinical and regulatory work was fragmented across teams, software and disciplines, with valuable context lost at every handoff. Making individual experts more productive with AI would not solve that problem.
My argument was simple: an LLM alone would not get us there. Models would need to retrieve information, use software and execute complex workflows rather than merely generate text. The years since have borne that out. The industry’s progression from chatbots to retrieval, tools, memory and agents reflects the search for the architecture LLMs were missing.
But the emergence of agents did not resolve the deeper problem. Agents can retrieve data and execute tools. They do not produce coherence to scientific work that is fragmented across people, data, software and disciplines. A useful research system must form a connective tissue across the entire scientific workflow, while grounding every conclusion in verifiable evidence and exposing every inference to scrutiny.
At ETH, I could describe the kind of system biomedical research needed, but the technology to build it was not yet ready. To explain what I had in mind, I kept returning to a much older model: the polymath.
The polymath problem
Leonardo da Vinci dissected cadavers and designed flying machines in the same notebooks. His anatomy informed his engineering; his engineering sharpened the questions he asked of anatomy. The polymath’s advantage was not simply breadth. It was the ability to make connections between fields that specialists in isolation could not see.
That way of working became a casualty of scientific progress: Da Vinci lived in a world whose recorded knowledge was still navigable. Today, no scientist can absorb more than a fraction of the knowledge relevant to their work, let alone develop deep expertise across every discipline involved in bringing a drug to patients (one estimate already placed the scholarly literature above 100 million documents a decade ago, and 12 years later PubMed alone now indexes tens of millions of publications).
So science scaled through specialisation. A medicinal chemist understands molecular binding. A clinical pharmacologist understands dose response. A bioinformatician interprets sequencing data. Each is an expert for one part of the system, but none can hold the whole drug program in their head. The result is a loss of coherence when crossing the boundaries of departments, disciplines and data systems. Every handoff sheds context and loses depth.
A model is not a scientific system
Throughout history, tools have extended human cognition. Writing externalised memory, libraries made knowledge retrievable, and databases and ontologies organised millions of discoveries into machine-readable form. Each expanded our access to knowledge, but connecting evidence across disciplines, judge its relevance, and carry those into the next question remained manual.
Then, large language models promised something different. For the first time, it seemed possible to interact with a system that was knowledgeable across domains. A universal scholar that had read almost everything. But while these models have become dramatically more capable, and with retrieval and tool use, they can answer many specialist questions accurately, this is not the same as doing complex science.
Scientific usefulness demands more than a plausible answer. Researchers need to know which evidence supports a claim, how sources were weighed (in particular conflicting evidence), what remains uncertain (data gaps) and whether the reasoning can be reproduced. A model’s internal knowledge, compressed into its weights, cannot provide that accountability on its own. That is the difference between systems that automate tasks and systems that augment scientific reasoning.
Follow the drug
The idea of AI systems that actively coordinate complex scientific work across data, software and experiments is becoming a key focus area of tech and pharma companies. Virtual scientist agents are already moving beyond question answering, helping researchers explore hypotheses, execute code, and process experimental data.
With agents becoming more capable and specialised, ****the bottleneck will no longer be finding information, or executing tools. It is connecting those research loops to the wider chain of scientific judgement that shapes a drug programme, from target biology and indication selection to biomarkers, clinical development and regulatory strategy.
We believe that modern biomedical research will be a distributed cognitive system: a hive mind across experts (each supported by their agent tools). Every meeting, experiment, analysis (human- or AI-operated) produces knowledge. Too much of it is lost as work moves between teams, software and disciplines.
Once viewed this way, the design problem changes. Instead of building agents that answer questions or completes isolated tasks, building a persistent scientific reasoning system that follows the drug rather than the department, carrying evidence, assumptions and context from one decision to the next.
That is the principle behind Perceptic. We organise AI workers around the lifecycle of a drug programme rather than around individual scientific functions. Each worker uses the tools, databases and methods appropriate to its domain, but contributes to a shared, continuously evolving understanding of the programme. Instead of resetting at every organisational handoff, knowledge compounds.
Importantly, the purpose is not to replace scientific judgement, but to extend it. Biology is inherently noisy and our understanding incomplete. Scientists should be able to reason across the entire programme without sacrificing the depth of their deep expertise in their disciplines.
Nearly three years after that talk at ETH, models, tool use and persistent context have matured enough to make that architecture practical rather than aspirational. The polymath disappeared because knowledge outgrew the individual mind. Its return will not come as one omniscient model, nor as an army of specialised agents. It will come through systems that bring together specialised human and AI experts to allow researchers to reason with both breadth and depth. That is the problem we set out to solve with Perceptic.
--
Tilman Flock is co-founder and CEO of Perceptic. Before founding the company, he held senior enterprise leadership roles at Palantir Technologies, working at the intersection of AI deployment and life sciences. Reach out at enquiries@perceptic.ai or visit https://perceptic.com




