AI & Models
Mercor benchmark shows AI models struggle with professional tasks
Mercor’s new APEX-Agents benchmark reveals that most AI models struggle to perform complex, multi-domain white-collar tasks, with the highest-performing model achieving 24% accuracy.
Nearly two years after Microsoft CEO Satya Nadella predicted AI would replace knowledge work, a new benchmark called APEX-Agents, developed by the training-data company Mercor, reveals that most AI models are struggling to perform white-collar work—defined as professional services like law, banking, and accounting. According to the research, every AI lab is getting a failing grade on the benchmark, as models struggle to complete complex tasks that mirror actual professional environments.
The benchmark is designed to reflect real-world professional services by testing systems on their ability to handle multi-domain environments. Mercor CEO Brendan Foody noted that models struggle with tracking down information across multiple domains, which is a core requirement of human professional work. Foody explained that human jobs do not involve a single individual providing all context in one place, but rather require operating across various tools like Slack and Google Drive.
The benchmark’s scenarios are drawn from actual professionals. For example, one legal task asks whether a company can treat log exports containing personal data to a U.S. analytics vendor during the first 48 minutes of an EU production outage as consistent with Article 49 of EU privacy laws. Answering such questions requires deep assessment of company policies and regional regulations.
When tested on these tasks, the models showed low success rates. Gemini 3 Flash achieved the highest score with 24% one-shot accuracy—defined as a model’s performance on a single attempt. GPT-5.2 followed closely with 23% accuracy, while Opus 4.5, Gemini 3 Pro, and GPT-5 all scored roughly 18%. This evaluation contrasts with OpenAI’s GDPval benchmark. While GDPval measures general knowledge across a wide range of professions, the APEX-Agents benchmark focuses on sustained tasks within a narrow set of high-value professional services.
Despite the low initial scores, Foody expects rapid improvement. “Right now it’s fair to say it’s like an intern that gets it right a quarter of the time, but last year it was the intern that gets it right five or 10% of the time. That kind of improvement year after year can have an impact so quickly,” Foody said.
Why it matters
The APEX-Agents benchmark provides a new way to measure how AI models perform on complex, multi-domain tasks typical of professional services, revealing that current models struggle to meet the standards of human professionals.