Today, Raycaster introduces Biopharma Bench V0.1. We placed frontier AI agents into employee seats across 12 simulated biopharma companies built from private industry records and genuine product histories: the submissions, quality systems, clinical files, and working drafts someone in that seat would actually encounter on the job.
Frontier models write fluent, persuasive technical prose in minutes. But completing an entire professional assignment without introducing a single unforced error is another matter entirely. Across 71 active assignments and 752 frozen evaluation criteria, GPT-6 Astra led on average rubric coverage with a 65.9% macro task mean (0 of 71 full passes). Claude Opus 5 reached 63.8% (0 full passes), and Grok 4.6 scored 63.0% while taking the top spot for complete task delivery with 2 full passes (2.8% Pass@1). DeepSeek V4.1 Flash reached 57.0%, Kimi K3 scored 48.5% (1 full pass), GLM-5.3 Flash reached 47.6%, and Gemini 3.8 Flash tied GPT-5.6 Sol at 46.6%.
Pass@1 asks whether every frozen substantive criterion passed on the first attempt. Mean task score asks how much of the rubric the model satisfied on average. Neither is regulatory sign-off, and neither means an AI is ready to operate without an accountable human in the loop.
The 752 evaluation criteria aren't grading for superficial formatting or polite prose. They capture the substantive judgments a director or VP with 15+ years in the work demands before putting their name on a file: did you identify the controlling protocol, did you reconcile contradictory batch records, and will your conclusions survive an FDA audit?
Why this benchmark
In drug and medical device development, evidence rarely comes in a neat, pre-packaged folder. It is scattered across regulatory submissions, quality deviation logs, clinical trial protocols, agency correspondence, and informal working drafts.
Those records frequently contradict one another, become superseded over time, or carry completely different legal and regulatory weight. An informal team memo cannot override an approved Common Technical Document (CTD) specification, and an exploratory Phase 1 protocol does not govern commercial release testing. Figuring out which document actually carries authority is the real job.
When a benchmark tests an agent on a clean prompt or a hand-picked evidence packet, it bypasses the hardest part of professional work. In the real world, typing the prose is easy. The challenge is investigating conflicting records and deciding what the company can defensibly do next.
Benchmark contents
The benchmark roster contains 71 assignments across 12 distinct company environments, built around authentic product histories and private biopharma records. The roster includes genuine development milestones—such as the complex dual-IND history of Allena Pharmaceuticals. Alongside these stands Havenor Therapeutics, a completely synthetic showcase company created so researchers can inspect full trajectories, coworker interactions, and deliverables in public without exposing confidential partner files.
These environments span major health authorities (US FDA CDER, CBER, CDRH; the European Medicines Agency; Health Canada; Korea MFDS; and China NMPA / CDE), diverse modalities (mRNA vaccines, antibody-drug conjugates, targeted small molecules, peptide injectables, oral solid fixed-dose combinations, and IVD companion diagnostics), and essential company disciplines across regulatory affairs, quality compliance, MSAT, and clinical site operations.
Rather than answering an abstract question, the agent sits down at an employee's desk at a specific moment in the company's history. It receives the actual files, tools, and background noise that a person in that seat would have had. Its job is to investigate the company state and produce the required work product: an updated regulatory assessment, an out-of-specification investigation, or a sign-off package. In Havenor Therapeutics, agents can also query simulated systems of record, pull lab data, and message coworkers while the work is underway.
You can explore curated run trajectories and actual deliverables from Havenor on the live benchmark page.
Results
Figure 1 ranks all eight systems by macro task mean—the primary metric, where every assignment carries equal weight. We also show Pass@1 on every row: the percentage of assignments where an agent satisfied every single frozen criterion in its single scored trial. You can sort the leaderboard by macro mean, complete passes, or criteria passed.
Figure 1 · Leaderboard
Partial performance versus full-task success
8 systems · 71 tasks · 12 companies · 752 criteria
Sorted by Macro mean, highest first.
01GPT-6 Astra
71/71 tasks · Codex · medium thinking- Pass@1
- 0.0%
- 0/71 full passes
- Macro mean
- 65.9%
- Equal weight per task
- Criteria
- 499/752
- 66.4% of items
02Claude Opus 5
71/71 tasks · Claude Code · medium thinking- Pass@1
- 0.0%
- 0/71 full passes
- Macro mean
- 63.8%
- Equal weight per task
- Criteria
- 485/752
- 64.5% of items
03Grok 4.6
71/71 tasks · Cursor CLI · high thinking- Pass@1
- 2.8%
- 2/71 full passes
- Macro mean
- 63.0%
- Equal weight per task
- Criteria
- 480/752
- 63.8% of items
04DeepSeek V4.1 Flash
71/71 tasks · Pi · high thinking- Pass@1
- 0.0%
- 0/71 full passes
- Macro mean
- 57.0%
- Equal weight per task
- Criteria
- 433/752
- 57.6% of items
05Kimi K3
71/71 tasks · Cursor CLI / Pi · default thinking- Pass@1
- 1.4%
- 1/71 full passes
- Macro mean
- 48.5%
- Equal weight per task
- Criteria
- 374/752
- 49.7% of items
06GLM-5.3 Flash
71/71 tasks · Pi · high thinking- Pass@1
- 0.0%
- 0/71 full passes
- Macro mean
- 47.6%
- Equal weight per task
- Criteria
- 365/752
- 48.5% of items
07GPT-5.6 Sol
71/71 tasks · Codex · medium thinking- Pass@1
- 0.0%
- 0/71 full passes
- Macro mean
- 46.6%
- Equal weight per task
- Criteria
- 364/752
- 48.4% of items
08Gemini 3.8 Flash
71/71 tasks · Cursor CLI · high thinking- Pass@1
- 0.0%
- 0/71 full passes
- Macro mean
- 46.6%
- Equal weight per task
- Criteria
- 356/752
- 47.3% of items
Across 71 scored trials per system, GPT-6 Astra tops the leaderboard with a 65.9% macro task mean (passing 499 of 752 criteria), followed closely by Claude Opus 5 at 63.8% (485 criteria) and Grok 4.6 at 63.0% (480 criteria). Grok 4.6 achieved the highest rate of complete work, recording 2 full passes (2.8% Pass@1, successfully clearing both an NMPA methotrexate deficiency response and a Korea ADC manufacturer review). DeepSeek V4.1 Flash reached 57.0% across all 71 tasks, Kimi K3 scored 48.5% with 1 full pass, GLM-5.3 Flash achieved 47.6%, and Gemini 3.8 Flash tied GPT-5.6 Sol at 46.6%(356 and 364 criteria passed, respectively).
However, looking only at top-line averages conceals how agents actually perform on the job. Because each environment represents an authentic professional seat, models frequently trade functional strengths:
Figure 2 · Domain Comparisons
Complementary capabilities across specialized biopharma seats
Cablivi (mRNA/biologics)
Astra vs. Opus leadOutside Cablivi, Astra and Opus pass the exact same number of items (445 of 685). Astra's 14-item overall lead comes entirely from Cablivi (54/67 vs. 40/67).
FoundationOne & Valsartan
Opus vs. Grok comparisonOpus leads FoundationOne (44/67 vs. 31/67) and the semaglutide freeze (24/32 vs. 18/32), while Grok leads the Valsartan GVP/CAPA packet (52/60 vs. 41/60). Overall, Opus leads Grok by just 5 items.
Korea MFDS (ADC)
Three-way parityKorea MFDS is a virtual tie across three architectures: Opus (54/78), Grok (53/78), and DeepSeek (52/78), with Astra scoring 44/78.
Site Notes & Metformin
Gemini vs. Sol tie at 46.6%Gemini and Sol tie on the macro mean at 46.6%, but solve opposite seats: Gemini leads on clinical site notes (92/127 vs. 72/127) and the metformin reply (30/48 vs. 16/48), while Sol leads on Korea, Pemazyre, Cablivi, and FoundationOne.
Take Astra and Opus: outside the Cablivi biologics environment, the two models pass the exact same number of criteria (445 of 685). Astra's entire 14-item lead comes from Cablivi. Opus leads FoundationOne CDx (44/67 vs. Grok's 31/67), while Grok pulls ahead on the Valsartan GVP/CAPA crisis investigation (52/60 vs. Opus's 41/60). On Korea MFDS ADC review, Opus (54), Grok (53), and DeepSeek (52) finish in a virtual dead heat. And while Gemini 3.8 Flash and GPT-5.6 Sol tie on the overall macro mean at 46.6%, they solve completely opposite seats: Gemini leads on clinical site notes (92 vs. 72) and NMPA metformin replies (30 vs. 16), whereas Sol leads on Korea, Pemazyre, Cablivi, and FoundationOne.
Scores only tell half the story—cost matters just as much when considering deployment. Figure 3 plots task performance against execution cost on a common set of 63 assignments where consistent telemetry was tracked.
Figure 3 · Cost-Performance Frontier
Cost-performance Pareto frontier across the 71-task benchmark
All three frontier systems land on different points of the cost-performance frontier:
- GPT-5.6 Sol (Codex): The most economical option at $1.32 per task, achieving a 45.5% mean score.
- Claude Opus 5 (Claude Code): Mid-tier cost at $2.86 per task, reaching a 63.6% mean score.
- GPT-6 Astra (Codex): Highest performance at 68.2%, at an average cost of $3.50 per task.
Across these runs, higher performance required more inference budget: no evaluated system was both cheaper and higher-scoring than another. These figures track candidate model inference only and exclude grading and container overhead.
What current agents get wrong
When an agent falls short on these tasks, it is almost never because the writing looks unpolished. The prose is usually fluent, confident, and neatly formatted. The failure almost always comes down to authority: which document actually governs?
You can see this clearly in Havenor Therapeutics Task 03, where a principal engineer is asked to review sterile-filter validation studies and draft a technical brief ahead of a health authority audit.
Figure 4 · Havenor Therapeutics Task 03
Three records. One claim that does not survive the protocol.
- 01 · ReportAVT-VAL-107
Bacterial retention report
Within the protocol-defined 30-minute window; organism viability control conformed.
Observation VAL-FLT-BRT-002-OBS-01. The second replicate began 14 minutes after the nominal inoculation window. The report records it as an accepted observation.
- 02 · Governing protocolAVT-VAL-107-P
Approved protocol v01
The approved protocol contains no 30-minute inoculation window.
Its acceptance criteria cover organism identity, challenge density, filtrate recovery, recovery controls, and a bubble-point limit. None of those criteria define a 30-minute inoculation window.
- 03 · Candidate briefDefensibility review
Inspector-facing brief
The brief repeats the 30-minute window as protocol-defined.
In the published Havenor Task 03 trial, the candidate opened the retention report and the scanned protocol, then carried the report’s claim into the final review. The brief does not record that AVT-VAL-107-P contains no such window.
The model demonstrated real capability: it navigated complex folder hierarchies, opened validation reports, and extracted data from executed protocol scans that simple text extractors couldn't parse.
The breakdown was subtler—and far more dangerous. In a secondary retention report, an observation noted that an inoculation began 14 minutes late, claiming this was “within the protocol-defined 30-minute window.” Yet when you open the controlling protocol (AVT-VAL-107-P), no such 30-minute window exists anywhere in the document.
An experienced human engineer knows that a secondary report cannot invent a protocol requirement out of thin air. If it isn't in the protocol, you flag the discrepancy. Instead, the model took the secondary report at its word, repeated the phantom 30-minute window as an established fact, and carried it straight into the brief prepared for the inspector.
In a real audit, handing an inspector a brief that misstates or invents a protocol requirement creates immediate compliance exposure. This illustrates the central finding of the benchmark: an agent can search thoroughly, reason through complex data, and write beautifully—and still fail at the essential act of professional verification.
How we evaluated
Every benchmark attempt takes place inside an isolated container. The agent is assigned an employee seat with access to the local company files, relevant tools, and a standard set of instructions: explore the company records, make reasonable inferences when encountering ambiguity, and continue working unless genuinely blocked.
We evaluate complete model-and-harness pairings, not bare models in a vacuum. On the 71-task roster, Astra and Sol ran on Codex at medium thinking, Opus on Claude Code at medium thinking, Grok and Gemini on Cursor CLI at high thinking, DeepSeek and GLM on Pi at high thinking, and Kimi on a blend of Cursor CLI and Pi at the provider default. DeepSeek has no medium setting, and the Grok and Gemini seats are the high model slugs. The harness governs how an agent inspects files, runs tools, and manages context, so results reflect the combined system.
Figure 5 · Evaluation Flow
How an assignment becomes a benchmark score
Employee seat & records
Role-appropriate seat, realistic company desk, files, and authorized tools.
Model + harness
Autonomous investigation under standard non-clarification instructions.
Deliverable
Document, spreadsheet, or authorized system action.
Mean task score
Share of frozen substantive criteria passed for that task, macro-averaged equally across all assignments in the benchmark suite.
Pass@1
The percentage of assignments where the candidate satisfied every frozen requirement in its single scored trial without human intervention.
Once an agent finishes, its deliverable is scored against frozen criteria that were quarantined away during the run. When a check is strictly quantitative—such as a specific calculated yield or a key spreadsheet value—we evaluate it with deterministic code.
Most professional work, however, cannot be captured by a regex. Two senior regulatory writers might structure a justification memo differently while both making sound, defensible arguments. The scores on this page are a hand grade of each deliverable against those frozen criteria. A check passes only when the submitted work itself supports it.
A brief word on what these scores mean: strong performance shows that an agent can navigate complex internal records, reconcile conflicting evidence, and draft defensible technical work. It does not certify regulatory compliance, and it does not mean an agent is ready to operate without accountable human oversight.
What comes next
This release represents an initial snapshot across eight leading systems. As frontier models and agent harnesses continue to evolve, Raycaster will publish updated results and expand the benchmark roster.
We believe evaluations are only as credible as their transparency. You don't have to take our word for any score: we invite you to inspect the full transcripts, coworker interactions, and final deliverables for yourself on the live benchmark page.
Inspect the work, not just the score
Explore the Havenor trials, step-by-step agent trajectories, and complete deliverables for Biopharma Bench V0.1 on Raycaster Eval.
