Evaluating AI Agents Live at the Grounded Reasoning Cup
Databricks, Tuesday, August 18th, 2026
Eleven academic teams tested agents live on OfficeQA Pro V2, built from about 120,000 pages of US Treasury documents.
Databricks reports on the Grounded Reasoning Cup, which challenged 11 academic teams to apply agents developed on OfficeQA Pro to OfficeQA Pro V2.
The new benchmark was built from approximately 120,000 pages of US Treasury documents and released for the event.
Results showed that generalization cannot be assumed: approaches developed on a familiar benchmark did not always transfer reliably to the new one.
The live format made the gap between benchmark tuning and general capability visible in real time. Databricks draws lessons for how enterprises should evaluate agents before deployment.