ramp-75def206·2 events·first seen Aliases: Ramp
Mercor and Ramp introduce APEX-Accounting, a closed benchmark of 160 expert-authored accounting tasks across 10 simulated worlds, covering reconciliation, accruals, transaction posting, and reporting. Across nine frontier models, Claude-Fable-5 (Max) leads with 56.4% Mean Criteria@3, while no model exceeds 2.6% on the stricter Pass^8 metric, indicating substantial headroom. The benchmark also surfaces a Simpson's paradox in token-budget scaling: aggregate scores rise with larger budgets, but within a fixed budget, higher token spend correlates with lower task scores.
Ramp's engineering team has deployed OpenAI's Codex with GPT-5.5 to automate and accelerate code review workflows, reducing feedback time from hours to minutes. The case study highlights an enterprise deployment pattern where agentic coding tools are integrated into production software development pipelines. This represents a concrete example of GPT-5.5 and Codex being used in real-world enterprise settings.