Research · methods · field notes
Evaluation methods, benchmark design, reference harnesses, and field notes from agents working across documents, workbooks, presentations, and critical domains.
MAY 23, 2026
Inside seven weeks this spring, Trimble bought Document Crunch and Trunk Tools shipped its Autodesk Forma integration. The commercial-construction version of AI submittal review is closing fast. The EPC technical review surface - vendor docs against NACE, API, ASME, and a project spec - is a different shape, and it is still up for grabs.

MAY 22, 2026
The FT and SynMax reported in April that ~40% of US data centers planned for 2026 will miss their deadlines by 3+ months. The cited causes - power, transformers, labor - sit outside the owner's fenceline. Inside the fence, with modular construction compressing build time from two years to thirty weeks, the bottleneck shifts to review. At $14M a day in delay cost, that matters.

MAY 21, 2026
Oracle Primavera codifies them: Approve, Approve as Noted, Reject, Revision Needed. Every major EPC contractor inherits some variant - A, B, C, D - and they decide whether procurement releases, fabrication proceeds, and the project moves. Most AI tools treat them as a sentence at the end of an answer. They are not.

MAY 20, 2026
Aconex's own Smart Manual documentation says it directly: handover manual quality is determined by the asset register, not the folder. The register is what facilities inherits and what eventually flows into Maximo. AI review's job at closeout is to keep the register honest.

MAY 18, 2026
2025 brought us the new idiom for building AI: give it tools to operate a computer. Capability is no longer the hard part - the translation layer between the agent's workspace and the user's files of record is. Here's what we learned building that layer.

APR 9, 2026
OfficeQA Pro is a serious grounded-reasoning stress test: search a massive Treasury Bulletin archive, recover the right rows after restatements, run real analysis, and pass a deterministic grader. In the paper's human study, annotators average ~35% when they search the full corpus and ~51% even when the exact PDF pages are handed to them - still harsh for work this finicky. We built the document agent we want in production and hit SOTA on their Table 4 setup with Gemini 3 Flash alone. Bigger models weren't the missing ingredient; the harness was.

MAR 24, 2026
Every week a new model tops the leaderboards, yet they still fail at complex knowledge work. Here is why the illusion of Pass@1 is holding agents back, and why reliability requires a version-controlled architecture of trust.

SEP 28, 2025
How document-native AI agents can connect CMC source material, tech-transfer workflows, and regulatory submissions while keeping experts in control.
