Overview
Autoresearch Bench tests whether AI agents can make research progress on their own. Each task gives an agent a well-defined objective, a metric it can evaluate its own performance along, and a fixed amount of time. The agent runs an experiment, measures the result, and chooses what to try next. It repeats this loop until time expires or the agent stops.
This benchmark covers optimization, model training, inference, and scientific machine learning. Among the five models tested, Claude Opus 5 ranks first overall. On several tasks, it also improves on the human baseline, sometimes with solutions we had not previously found.
Leaderboard
Results
We measure the performance of Claude Opus 5, GPT-5.6 Sol, Gemini 3.7 Flash, Kimi K3, and Grok 4.6 on Autoresearch Bench and report their scores over the course of 4-hour autoresearch loops. The models are generally not given access to the internet and submit their solutions to an external grader outside of the agent's sandbox (see Evaluation and scoring for more details). We optionally show 12-hour rollouts in the graphs below, but they are not averaged into the leaderboard score due to low sample size.
How models progress over time
To evaluate model performance over time, we take the mean of the step-wise autoresearch curves within and across tasks for each model (after normalization) and plot them below.
The sharp separation between models is notable.
Opus 5 is better at optimization tasks
When comparing model performance across task domains, Opus 5 performs relatively much better on optimization tasks than the other four models, scoring roughly 2x higher.
Opus 5 solutions generalize best to held-out data
Most tasks hold out a private split the agent never sees and is evaluated against. As an agent you want to score high on the private set, but psychologically might get trapped into optimizing for the public score that you can see. Looking at how well each model is aligned, we find three tasks with generally larger public/private disagreement:
Figure 3 plots the public versus the private score on every scored submission for these tasks. A submission with good public/private alignment falls along the diagonal, whereas a submission that overfits to the public score falls below it.
In two of these three tasks, Opus 5 solutions hold up best when applied to the held-out set. On these tasks Opus's public scores are better aligned with the held-out private scores when compared to the other models.
GPT-5.6 Sol submits the most
One interesting behavioral trait that we find is that Sol submits to the external grader far more often than the other four models.
Although, we find that submitting more often does not necessarily transfer to the quality of each change. Sol and Grok submit far more, but each of their submissions moves the score less. Kimi and Opus submit less and gain more per submission.
Gemini 3.7 Flash gives up the earliest
Agents are given four hours to complete the task, but can optionally choose to end early. We find that while Sol, Kimi, Grok, and Opus share similar behavior and place their last graded submission inside the final half hour of the budget on a typical run, Gemini chooses to stop early. In nearly a third of its runs it writes a closing summary and ends the session on its own, usually with hours of budget left.
Figure 7 shows this as the share of each model's runs still submitting at each hour.
Evaluation and scoring
Each agent works in an isolated workspace with no general internet access. It can experiment and score itself locally as often as it likes, and in most tasks additionally receives access to an external grader in a separate sandbox that it can submit its solutions to.
This grader returns a public score that the agent can see and use as feedback, and in many cases (but not all) holds out a private score that the agent cannot see. A run's final score is taken leniently as the best private score it reached.
Comparing across tasks
Within a task, every model is scored on the same metric, so the comparison is direct. Across tasks it is not. Many tasks are not scored on a [0, 1] scale, and many have no known ceiling. In cell segmentation, for example, a perfect mAP of 1.0 is possible in principle, but no model trained on the training set generalizes perfectly to the held-out set, so the true ceiling is unknown and well below 1.
We therefore normalize with a quantity we can measure. Each score is divided by a per-task frozen divisor, the average over models of each model's mean final score on that task:
This gives an interpretable, continuous score:
- Score = 1: An average score
- Score > 1: An above average score
- Score < 1: A below average score
A model that reaches twice the average final score on a task scores 2 there. The divisors are frozen so the published numbers do not drift as new models are added or more runs are included. The frozen scale is reported here. A model's benchmark score is the mean over its runs within a task, then the mean across tasks with equal weight, so a task we happened to run more often does not count for more.
Acknowledgements
We thank the task authors and contributors who helped us build this benchmark.
Autoresearch Bench is built at Emulated by:
- Ryan Peters
- Sid Patllollu
- Joseph Wang
- Amit Prakash
- Lalithadithya Nataraja
- Scott Sauers
Citation
If you use Autoresearch Bench, cite this post as:
@misc{emulated2026autoresearchbench,
title = {Autoresearch Bench: Evaluating Agents on Iterative Research Tasks},
author = {Peters, Ryan and Patllollu, Sid and Wang, Joseph and Prakash, Amit and Nataraja, Lalithadithya},
year = {2026},
month = sep,
howpublished = {Emulated},
url = {https://www.autoresearch-bench.com/}
}@misc{emulated2026autoresearchbench,
title = {Autoresearch Bench: Evaluating Agents on Iterative Research Tasks},
author = {Peters, Ryan and Patllollu, Sid and Wang, Joseph and Prakash, Amit and Nataraja, Lalithadithya},
year = {2026},
month = sep,
howpublished = {Emulated},
url = {https://www.autoresearch-bench.com/}
}Join us
If you're interested in learning more about the technical challenges we're solving at Emulated, reach out to us at hiring@emulated.so.