Evaluation and training data

Data that contains the work.

Not prompt-response pairs stripped of context. Raycaster captures the task, environment, trajectory, artifact, rubric, and judgment as one inspectable record—including failures labs can learn from.

One evaluation record

The useful unit is bigger than an answer.

01

Task

A bounded request with real deliverables and acceptance conditions.

02

Source packet

The documents, workbooks, references, and initial workspace state.

03

Trajectory

Every visible action, tool result, intermediate file, and recovery step.

04

Artifact

The finished workbook, document, presentation, analysis, or decision.

05

Rubric

Criteria written around what a domain expert would actually inspect.

06

Judgment

Deterministic checks, grader evidence, expert review, and adjudication.

Published record · open now

Inspect the task, run, artifact, and judgment together.

Choose a real task below. The record exposes the prompt, workspace scale, trajectory metadata, finished output, rubric result, cost, and a direct path into the full public run.

Published runs · live product surface
APEX-Agents · Law5 / 5 · Pass6m 12s

Review warranty claims and update refund amounts

Review the attached warranty claims, then edit the existing product purchases spreadsheet to show the maximum refund amount a customer could receive for each product purchased.

Failure mode · Missed governing agreement constraints

GPT-5.5 · xhigh · Raycaster harness · public tier

Live public run · transcript · files · rubric

Inspect full run ↗

Built with experts

Domain judgment is part of the dataset.

01CaptureExperts provide representative work, source material, and edge cases.
02SpecifyWe translate tacit standards into rubrics and observable acceptance conditions.
03RunModels and harnesses operate in the same instrumented environment.
04AdjudicateEvidence, graders, and expert review resolve what actually passed.