If an AI research agent improves its own harness using one model, does the improvement still hold on other models?
The arXiv paper on AIDE², our RSI system, tests this and compares it with more AI research agents.
We're releasing the arXiv paper on AIDE².
The RSI system where AI research agents improve their own research efficiency.
It includes new results on transfer across models and comparisons with more AI research agents: [1/4]
Can autoresearch agents find better training data, not just better training code?
Here’s our new EMNLP paper:
AutoData: Agentic Search for Pre-training Data Selection [1/6]
Thousands of search trees.
99 rewrites of the agent's own code.
9 in 10 rejected by a held-out gate.
Eight unattended days.
The AIDE² pilot report is out. AIDE_85, the discovered agent and more detailed PDF tech report, will be released when the remaining analysis lands
The first experimental evidence of recursive self-improvement (RSI).
Autoresearching the autoresearch agent for eight days.
The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)
Everyone assumes RSI would lead to an intelligence explosion. Is that really the case?
We tried to clean up the logic after 8 months of building RSI at the harness layer.
Not all RSI is equal. We split it into 4 levels: delegation, net positive, ignition, and inflection. (1/9)
One of our autoresearch runs sat flat for 60 steps. One message got it moving again.
Everyone is pushing research agents toward full autonomy, but what helps most is being able to step in when a run goes wrong, without breaking the loop. (1/6)