Abstract

Comprehensively evaluating LLM-based AI agents across interactive environments is difficult because tasks, scaffolds, verifiers, and scoring rules are fragmented across sources. Existing efforts to unify these evaluations remain limited in scale and domain, making costly reruns necessary and leaving available results difficult to compare.

We introduce Messier, a corpus that combines public evaluation results with new runs on underrepresented professional and scientific benchmarks. It brings these heterogeneous sources into a common data model while retaining the context needed to interpret each result in greater depth.

The current release spans 31 benchmarks, including SWE-bench, BFCL, OSWorld, τ²-bench, HarveyAI-Lab, and Toolathlon. Its 725 agents include recent models such as GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro, and DeepSeek V4 Pro, together with scaffolds such as OpenHands, Claude Code, and Terminus. Across 11,999 tasks, Messier describes 72,000 verifiers by type, including LLM judges, scripts, exact-match checks, and human labels. We also record whether each verifier checks the final output or environment state or examines intermediate interactions. The corpus includes 118,089 trajectories from 15 benchmarks, including MathArena, TerminalBench, τ²-bench, Toolathlon, SWE-bench Pro, and the six benchmarks evaluated through Harbor.

Using this corpus, we find that frontier progress is uneven across benchmark groups and show that scoring rules can change measured performance and agent rankings. We further derive open-data capability scores that correspond closely to Epoch's Evaluation Capability Index (ECI). Researchers can readily extend these scores to estimate performance across custom subsets, such as specific domains, occupations, action spaces, or verifiers.

One of several models and environments is shown. A model and scaffold form an agent. The agent completes repeated trials of a task, each evaluated by one or more verifiers whose outputs form a trial result.
Fig. 1. Benchmark evaluation flow. An environment supports one or more tasks, each associated with one or more verifiers. A model and scaffold form an agent, which may execute the same task over repeated trials. The verifiers evaluate each trial, and the scoring rule combines their outputs into one trial result. Benchmarks usually report only the final result. Meaningful comparison also requires the individual verifier results and the rule used to combine them, so researchers can identify what contributes to success or failure more precisely.

Corpus composition

Benchmarks by group

Tasks by group

Models, scaffolds, and agents

Scoring rules

Verifier types

Frontier by model release quarter

First and latest observed frontier

Fig. 2. Corpus composition and frontier analysis. MESSIER records how agents are formed, what actions and state each environment exposes, and how verifiers and scoring rules produce results. Because a task may provide more than one action type or verifier, we can examine these choices across the corpus. The five benchmark groups are qualitative and may overlap. For example, an enterprise task may also involve function calling. GUI evaluations remain notably underrepresented. The lower panels use model release dates to trace the observed frontier over time. For each quarter, the frontier averages the best result observed on each task among agents whose models had been released by then. Vertical bars show one standard error across the benchmarks in each group.

Item Response Theory

Item Response Theory (IRT) is a family of mathematical models used to design, analyze, and score tests, surveys, and other psychological or educational assessments. It models the probability of an observed response as a function of latent characteristics of the respondent and the item. In our setting, respondents are models, items are tasks, and each response records whether a model completed a task successfully. The 1PL model places model ability θ and task difficulty β on a shared scale, with success becoming more likely as ability exceeds difficulty. We estimate both from the observed results across overlapping tasks.

By placing all tasks on one shared scale, the analysis allows us to compare difficulty across otherwise different domains, such as mathematics and enterprise workflows. This comparison depends on the strong assumption that the same scale meaningfully represents difficulty across those domains. Separate, multidimensional, or hierarchical scales may better represent differences between occupations, industries, model families, and scaffolds. Additionally, the strength of these comparisons depends on coverage. The data include 250,048 of 1,674,845 possible model-task observations, or 14.9%. Because IRT connects models and tasks through observed results, comparisons are better supported when models are evaluated on overlapping sets of tasks. The matrix below shows the number of models represented in each benchmark group and shared between groups. We therefore welcome new benchmarks, models, and runs that expand the corpus and connect evaluations that currently have little overlap.

Model coverage and overlap by group

Fig. 3a. Model coverage and overlap.

Distributions of model ability and task difficulty

Fig. 3b. Ability and difficulty distributions.

Comparison with Epoch ECI

The panels below compare Messier ability estimates with Epoch's published general, coding, and mathematics ECI scores. To do so, we match models across the two sources and compute Spearman rank correlation, using the full-corpus fit for general ability and re-estimating model ability on tasks classified under SOC 15-1200 and SOC 15-2000 for coding and mathematics, respectively, while holding task difficulty fixed at its original value.

General

Coding

Mathematics

Fig. 4. Comparison with Epoch ECI. Each panel compares Messier model ability with the corresponding Epoch ECI score for models represented in both sources.

Ability by industry

We repeat this methodology using NAICS classifications for the three industries with the broadest coverage across benchmark groups. For each industry, we re-estimate model ability with the IRT task-difficulty parameters β fixed at their original fit values. The panels compare these estimates with overall ability and include models evaluated on at least 10 tasks from at least two benchmark groups in that industry.

Professional, Scientific, and Technical Services

Health Care and Social Assistance

Finance and Insurance

Fig. 5. Ability by industry. Each panel compares full-corpus model ability with ability re-estimated from tasks in the indicated NAICS industry.

Citation

@article{krsteski2026messier,
  title={Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation},
  author={Krsteski, Stefan and Meyer, Charlotte and Allegre, Guillaume and O'Halloran, Tony and Sallinen, Alexandre},
  journal={arXiv preprint arXiv:2607.25891},
  year={2026}
}