Sponsored benchmark research

We pay researchers to build the benchmark their field is missing.

If you know a kind of professional work well enough to say when it has been done right, you have the part that’s hard to buy. The rest of a benchmark is software, and we build that.

Write us two pages. If it fits, we fund two to three months of your time, cover the AI credits, and put engineering behind it. It’s mostly PhD students, postdocs, independent researchers, and people who do the job.

Time and money

A full stipend for two to three months, longer if the work keeps going somewhere.

Credit

Your name on the benchmark, plus an invite to our EvalScience day and retreat.

Support

You write the domain content. We build the software, cover the AI credits, and run the models.

Publishing

Open, access-controlled, or private. We agree on which one before you start.

What we’re funding right now.

Healthcare delivery and payer operations

Coverage and care decisions pieced together from systems that don’t agree.

Prior authorization and utilization review · care quality and facility risk · clinical documentation · coding and revenue integrity · safety reporting

What we’ve published here ↗

Life sciences and biopharma

Where a governing document sets the right answer, and general reasoning still gets it wrong.

CMC and change control · regulatory submissions · GMP deviations and CAPA · clinical operations · pharmacovigilance · promotional review

What we’ve published here ↗

Energy and industrial infrastructure

Where one missed line in a vendor package becomes rework, downtime, or a safety finding.

Document control against NACE, API, and ASME · grid and transmission planning · project controls and turnover · project finance

What we’ve published here ↗

Financial services and markets

Where a break has to decompose exactly, and plugging it to make the numbers tie is the mistake.

Trading and treasury operations · post-trade reconciliation · credit and underwriting · valuation and deal models · regulatory reporting

What we’ve published here ↗

Legal and regulatory compliance

Where citing the wrong clause is worse than saying you don’t know.

Contract and obligation review · transaction diligence · litigation evidence · regulatory interpretation and filings

What we’ve published here ↗

Something we haven’t listed

Same terms. What matters is that the work has consequences and can be graded.

Insurance claims · tax and the accounting close · public procurement · trade compliance · laboratory and scientific operations

Tell us about it →

What it looks like once it’s out.

A care operator’s manual review, a grid curtailment analysis, and a merger model. You get the task, the files, what the model produced, and how the rubric scored it. Not all of them passed, which is the point.

Published runs · live product surface
APEX-Agents · Law4 / 10 · FailPublic run

Find the 2025 operations-manual changes that create regulatory exposure

Please review Grove’s 2023 and 2025 operations manuals. Let me know if there are any changes in the 2025 manual that may present issues for Grove’s regulatory compliance. For any regulatory issues identified, tell me what changed in the 2025 manual and why it presents an issue.

Failure mode · Named the changed sections but not the rules they break

GPT-5.5 · xhigh · Raycaster harness · public tier

Live public run · transcript · files · rubric

Inspect full run ↗

How it actually goes.

01

You send two pages

The workflow, the tasks you’d build, and why you think models fail them.

02

We scope and fund

A short statement of work: deliverables, timing, how it gets published, and the stipend.

03

You write one task first

We review it before the rest, so a format problem surfaces in week one instead of at delivery.

04

We build it and run models

Your files become a working environment. You see exactly where models fail.

You write the cases, the ground truth, and the rubric criteria. Nobody is asked to write environment code, verifiers, or schema tooling.

Some of this can be open. Some of it can’t.

Open

Synthetic or public-source data, rights-cleared. Tasks, rubrics, and graders get published and open-sourced, minus a holdout slice we keep back.

Access-controlled

Realistic and cleared, but not for posting publicly. Scores and method are public; source files sit behind access review.

Private

Can’t be de-identified, or publishing it would contaminate what you’re measuring. Stays a private holdout. Your credit doesn’t change.

Two pages is enough.

No template, no deadline, no formal application.

evalscience.org ↗

Worth including

The specific workflow, not the industry

The tasks you’d build, and what each one tests

Why you think models fail them

What data you’d use, and what would have to stay private