The AI Front Page

Reading signals from this article are folded back into your front page ranking on this device.

Research/arXiv AI/ML/July 29, 2026 at 5:56 PM

arXiv paper: APEX-Accounting

A new arXiv AI paper by Julien Benchek, Austin Bennett, and Jasmin Kern, and 8 more studies APEX-Accounting.

ResearchAI
Research / arXiv AI/ML
Source

Follow arXiv AI/ML to make it a durable For You signal.

arXiv ID: 2607.27189v1 Title: APEX-Accounting Authors: Julien Benchek, Austin Bennett, Jasmin Kern, Ryan Stevens, Rene Sultan, Charis Ching, Hayley Popiel, Vaibhav Mittal, Felix Mercier, Brendan Foody, Bertie Vidgen Primary category: cs.CL Categories: cs.CL, cs.AI, cs.HC Comment: Public dev set: https://huggingface.co/datasets/mercor/apex-accounting Published: 2026-07-29T17:56:49Z Updated: 2026-07-29T17:56:49Z Abstract: We introduce APEX-Accounting, a benchmark built by Mercor in partnership with Ramp, to assess whether frontier models can do the real work of accountants. Tasks include reconciling accounts, accruing expenses, posting transactions, and producing reports. The private eval set comprises 160 tasks, split across 10 worlds. Each world contains an accounting system, as well as spreadsheets, PDFs, and other files. Every task was authored and solved by experts in accounting and bookkeeping, who also wrote grading rubrics. Across nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, ahead of Muse-Spark-1.1 (xHigh) at 52.6%. No model scores more than 2.6% Pass^8 (GPT-5.6-Sol (Max+Pro)) and the highest Pass@8 is 21.5% (Muse-Spark-1.1 (xHigh)). We experiment with increasing the token budget from $1 to $50 and observe an instance of Simpson's paradox: scores increase as the token budget increases but within a given budget-constrained harness, scores are lower on tasks where the model spends more tokens. As APEX-Accounting is a closed benchmark, leaderboard evals can be run for any frontier model on request. PDF: https://arxiv.org/pdf/2607.27189v1