Beyond one billion tokens per day: a small experiment in AI for drug discovery
Published on September 10, 2026
Contact: Gianni De Fabritiis (g.defabritiis@acellera.com)
Repository:
Acellera AI lab notes
Report comparison:
Comparative evaluation of seven IL23R-p19 reports
Report example:
IL-23R–p19 Inhibition for Psoriasis
Acellera currently uses about one billion tokens per day through agentic operations. I expect that volume to grow tenfold over the next year. At that scale, using frontier models for every step of increasingly complex scientific work would be unaffordable for us. I suspect many other teams will face the same constraint.
That is why I am interested in how far we can get with models whose weights we can run ourselves, including models that fit on a single GPU. The useful question is how much scientific work they can do, and how much correction their output needs.
I ran a small experiment comparing Claude and ChatGPT with PlayMolecule® AI, using a demanding drug-discovery task.
The task was to retrieve and understand scientific literature and structural data, then produce a target product profile, or TPP, and a small-molecule discovery strategy for the IL23R-p19 interaction in psoriasis. A TPP describes the properties a proposed medicine should have. To be useful before clinical development, it also needs measurable criteria that help a team decide which compounds to advance.
IL-23 is one of the targets Acellera uses to validate its technology, so it was a natural choice. It is also an interesting reasoning challenge. Published work describes an induced pocket on the p19 subunit, inhibitory peptides and a fragment-screening approach. Turning that evidence into a credible small-molecule program requires care: binding near a receptor interface does not, by itself, establish that a small molecule will block the interaction. Lecomte et al., 2025
I gave every system the same prompt:
"Create an informed, validated report with text, tables, and figures based on publicly available information about this target, useful for drug discovery. Include structural information and visualizations, possible discovery strategies with small molecules, and a target product profile (TPP) of a small-molecule inhibitor therapeutic for the IL23R-p19 interaction for psoriasis at the preclinical level, and possible structural strategies to target it. "
I collected seven reports, anonymized them, and asked ChatGPT Astra-Extra-High to compare them. Only afterward did I reintroduce the system names. The names in the figures follow the experiment's model mapping.
The evaluation focused on whether the reports could help someone start a discovery campaign. It assigned up to 20 points for the TPP and 20 for the small-molecule strategy, using ten equally weighted criteria. These covered product definition, potency, exposure and dose, translation, developability, structural evidence, starting chemistry, assay controls and experimental decisions. The review included targeted checks against primary sources, structural coordinates and numerical calculations. The full comparative evaluation contains the rubric, evidence and report-specific criticisms.
The first surprise was how much the local model could produce. I ran Qwen3.8-28B on a single RTX 5090, in the office, through PlayMolecule AI. Seeing a model on one local GPU produce a substantial report spanning structural biology, pharmacology and discovery planning was encouraging.
Its limitations were equally clear. It scored 11/40, the lowest result in this set. The evaluation identified consequential errors in pharmacological interpretation, structural contacts and the product profile. Its output was useful as background material to examine, but it needed substantial correction before it could guide a campaign. That distinction matters: generating a substantial scientific report is an achievement; making it reliable enough to direct experiments is a harder one.
I still find this promising. A local model can already contribute to a complex task, even if its useful role today requires close review and narrower responsibilities. This experiment does not establish that Qwen is the only small model capable of doing so; I did not run that comparison. It does make me optimistic about what we will be able to do locally.
The second surprise was that the ranking broadly matched my expectations of the models' reasoning capabilities. My interpretation is that model capability remains an important constraint on this kind of scientific work. Retrieving a paper is only part of the job. A system also has to reconcile conflicting evidence, recognize unsupported assumptions and connect its proposed experiments to a decision.
The detailed scores help explain the differences:
For example, the PlayMolecule AI on Astra-high report did a particularly good job checking the published chemical starting points, including discrepancies between the main paper and supplementary data. That affects which compounds a team would buy, reproduce and trust. The ChatGPT 6 Astra-high report was especially useful in connecting potency, free exposure and dose through a worked calculation. It showed why an attractive potency target can still imply an impractical dose under particular assumptions.
Those are the kinds of distinctions I wanted the experiment to reveal. A long report can contain many reasonable individual statements while still failing to provide a coherent experimental plan. The lower-scoring reports often had useful ideas, but left important gaps between a binding hypothesis, functional inhibition and a feasible product.
I would be cautious about turning this ranking into a general measure of “model intelligence.” This was one task and one report per configuration. The systems ran in their respective environments; the evaluation does not establish that tool access, token budgets or execution time were matched. It also did not measure intelligence independently. The results are consistent with my impression that stronger reasoning matters, but they do not isolate its contribution from retrieval and the harness. By “harness,” I mean the software around the model: the tools it can use, how it accesses information, and how it carries a task through to a finished result.
The third surprise was the comparison between PlayMolecule AI and ChatGPT using Astra. Their reports scored 37/40 and 36/40, respectively. Both scored 18/20 for the TPP; the difference was a single point in the strategy assessment.
My takeaway is that PlayMolecule AI was competitive with ChatGPT on this task using the same base model, with different strengths in the final reports. I would not claim a meaningful win from one point, or general equivalence from one pair of runs. But the result is encouraging for a scientific harness: in this example, it supported a report at the same operational level as the standard ChatGPT environment.
There are limits to using an LLM to judge other LLMs. Anonymizing the reports reduced explicit brand cues, but did not remove possible preferences in the judge or clues in the writing. The evaluation remains an assessment of these documents, rather than an independent expert consensus or a demonstration that any proposed compound works. It also does not include comparable cost and latency measurements, so this is not yet a price-performance benchmark.
For our next experiments, I want repeated runs, more targets and expert review of the scientific errors. I also want to measure the cost of getting to an acceptable result, including the correction work. At a billion tokens per day, that is the quantity that matters to us. A smaller model becomes valuable when the work it saves exceeds the work needed to check and repair its output.
I came away encouraged by what a single GPU could produce, and reminded of how demanding reliable scientific reasoning remains. There is substantial room to improve both the models and the way we use them. Having this much capability running locally in the office makes the future look good.
If you work on IL-23, small-molecule discovery or scientific AI, I would welcome comments.
Gianni De Fabritiis, CEO
g.defabritiis@acellera.com