Skip to main content

Autoresearch
Bench

A benchmark for agents that iteratively test
and improve solutions through experimentation

Autoresearch Bench

A benchmark for agents that iteratively test and improve solutions through experimentation.

Emulated

Overview

Autoresearch Bench tests whether AI agents can make research progress on their own. Each task gives an agent a well-defined objective, a metric it can evaluate its own performance along, and a fixed amount of time. The agent runs an experiment, measures the result, and chooses what to try next. It repeats this loop until time expires or the agent stops.

This benchmark covers optimization, model training, inference, and scientific machine learning. Among the five models tested, Claude Opus 5 ranks first overall. On several tasks, it also improves on the human baseline, sometimes with solutions we had not previously found.

Leaderboard

Models ranked by mean normalized score across the tasks with graded runs.

Results

We measure the performance of Claude Opus 5, GPT-5.6 Sol, Gemini 3.7 Flash, Kimi K3, and Grok 4.6 on Autoresearch Bench and report their scores over the course of 4-hour autoresearch loops. The models are generally not given access to the internet and submit their solutions to an external grader outside of the agent's sandbox (see Evaluation and scoring for more details). We optionally show 12-hour rollouts in the graphs below, but they are not averaged into the leaderboard score due to low sample size.

How models progress over time

To evaluate model performance over time, we take the mean of the step-wise autoresearch curves within and across tasks for each model (after normalization) and plot them below.

Figure 1. Score over time for all tasks. Mean normalized score per model over hours since each run started. In the all runs view solid lines are drawn for successful runs, whereas dotted lines are shown for runs with failures (e.g. infra failures, rate limits, etc.). These failures are not averaged into the mean curves or the leaderboard, and are only shown for context.

The sharp separation between models is notable.

Opus 5 is better at optimization tasks

When comparing model performance across task domains, Opus 5 performs relatively much better on optimization tasks than the other four models, scoring roughly 2x higher.

Figure 2. Score over time for optimization tasks. The same curves restricted to the optimization tasks. As in Figure 1, dotted lines in the All runs view are rollouts that failed QA and do not enter the means or leaderboard.

Opus 5 solutions generalize best to held-out data

Most tasks hold out a private split the agent never sees and is evaluated against. As an agent you want to score high on the private set, but psychologically might get trapped into optimizing for the public score that you can see. Looking at how well each model is aligned, we find three tasks with generally larger public/private disagreement:

Figure 3 plots the public versus the private score on every scored submission for these tasks. A submission with good public/private alignment falls along the diagonal, whereas a submission that overfits to the public score falls below it.

Figure 3. Tasks where models overfit to the public score and underperformed on the held-out set.

In two of these three tasks, Opus 5 solutions hold up best when applied to the held-out set. On these tasks Opus's public scores are better aligned with the held-out private scores when compared to the other models.

Figure 4. Public-to-private alignment. Each model's private score as a fraction of its public score on the three tasks above. A ratio near 1 means the private split confirms what the public board showed.

GPT-5.6 Sol submits the most

One interesting behavioral trait that we find is that Sol submits to the external grader far more often than the other four models.

Figure 5. Submissions over time for all tasks. Each model's share of a task's graded submissions, cumulative over the run clock and averaged across tasks with equal weight. Within a task, a model's share is its mean submissions per run over the sum of every model's mean, so within a task the shares sum to 100% at the end and an average model holds 20%. Sharing within each task first keeps a task where grading is cheap and runs bank hundreds of submissions from dominating the mean.

Although, we find that submitting more often does not necessarily transfer to the quality of each change. Sol and Grok submit far more, but each of their submissions moves the score less. Kimi and Opus submit less and gain more per submission.

Figure 6. Many small steps or few large ones. One faint logo per graded run: the graded submissions it spent against the average gain a submission banked when it raised the run's best score, in normalized score. The solid logo is each model's median run.

Gemini 3.7 Flash gives up the earliest

Agents are given four hours to complete the task, but can optionally choose to end early. We find that while Sol, Kimi, Grok, and Opus share similar behavior and place their last graded submission inside the final half hour of the budget on a typical run, Gemini chooses to stop early. In nearly a third of its runs it writes a closing summary and ends the session on its own, usually with hours of budget left.

Figure 7 shows this as the share of each model's runs still submitting at each hour.

Figure 7. When models stop submitting. The proportion of each model's graded runs that have not yet made their last graded submission.

Evaluation and scoring

Each agent works in an isolated workspace with no general internet access. It can experiment and score itself locally as often as it likes, and in most tasks additionally receives access to an external grader in a separate sandbox that it can submit its solutions to.

This grader returns a public score that the agent can see and use as feedback, and in many cases (but not all) holds out a private score that the agent cannot see. A run's final score is taken leniently as the best private score it reached.

Comparing across tasks

Within a task, every model is scored on the same metric, so the comparison is direct. Across tasks it is not. Many tasks are not scored on a [0, 1] scale, and many have no known ceiling. In cell segmentation, for example, a perfect mAP of 1.0 is possible in principle, but no model trained on the training set generalizes perfectly to the held-out set, so the true ceiling is unknown and well below 1.

We therefore normalize with a quantity we can measure. Each score is divided by a per-task frozen divisor, the average over models of each model's mean final score on that task:

s~=sd,  d=1M∑m=1Msˉm\tilde{s} = \frac{s}{d}, \; \qquad d = \frac{1}{M}\sum_{m=1}^{M} \bar{s}_m

This gives an interpretable, continuous score:

  • Score = 1: An average score
  • Score > 1: An above average score
  • Score < 1: A below average score

A model that reaches twice the average final score on a task scores 2 there. The divisors are frozen so the published numbers do not drift as new models are added or more runs are included. The frozen scale is reported here. A model's benchmark score is the mean over its runs within a task, then the mean across tasks with equal weight, so a task we happened to run more often does not count for more.

Acknowledgements

We thank the task authors and contributors who helped us build this benchmark.

Autoresearch Bench is built at Emulated by:

Ryan Peters
Sid Patllollu
Joseph Wang
Amit Prakash
Lalithadithya Nataraja
Scott Sauers

Citation

If you use Autoresearch Bench, cite this post as:

bibtex
@misc{emulated2026autoresearchbench,
  title        = {Autoresearch Bench: Evaluating Agents on Iterative Research Tasks},
  author       = {Peters, Ryan and Patllollu, Sid and Wang, Joseph and Prakash, Amit and Nataraja, Lalithadithya},
  year         = {2026},
  month        = sep,
  howpublished = {Emulated},
  url          = {https://www.autoresearch-bench.com/}
}
@misc{emulated2026autoresearchbench,
  title        = {Autoresearch Bench: Evaluating Agents on Iterative Research Tasks},
  author       = {Peters, Ryan and Patllollu, Sid and Wang, Joseph and Prakash, Amit and Nataraja, Lalithadithya},
  year         = {2026},
  month        = sep,
  howpublished = {Emulated},
  url          = {https://www.autoresearch-bench.com/}
}

Join us

If you're interested in learning more about the technical challenges we're solving at Emulated, reach out to us at hiring@emulated.so.