ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research Paper • 2606.07591 • Published May 28 • 106
SearchSwarm: Towards Delegation Intelligence in Agentic LLMs for Long-Horizon Deep Research Paper • 2606.09730 • Published Jun 8 • 56
DeNovoSWE: Scaling Long-Horizon Environments for Generating Entire Repositories from Scratch Paper • 2606.10728 • Published Jun 9 • 37
Towards Diverse Scientific Hypothesis Search with Large Language Models Paper • 2606.10587 • Published Jun 9 • 2
TreeSeeker: Tree-Structured Trial, Error, and Return in Deep Search Paper • 2606.11662 • Published Jun 10 • 11
EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery Paper • 2606.13662 • Published Jun 11 • 33
Speaking the Language of Science: Toward a General-Purpose Generative Foundation Model for the Natural Sciences Paper • 2606.16905 • Published Jun 15 • 7
Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark Paper • 2606.18648 • Published Jun 17 • 16
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? Paper • 2606.24530 • Published Jun 23 • 67
Towards Automating Scientific Review with Google's Paper Assistant Tool Paper • 2606.28277 • Published Jun 26 • 10
AI translation of literary texts is "fine", but readers still prefer human translations Paper • 2606.26040 • Published Jun 24 • 5
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill Paper • 2608.11924 • Published Aug 12 • 109
Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of Intelligence Paper • 2608.12036 • Published Aug 12 • 89
DarwinX: Evolving Agent Harnesses Through Natural Selection Paper • 2608.07545 • Published Jul 31 • 116
PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback Paper • 2608.30241 • Published Aug 31 • 11
StudentSim: Training LLM-based Student Simulators Paper • 2609.01591 • Published about 1 month ago • 493