| Literature Retrieval |
| 1 |
CiteME |
arXiv'24 |
Retrieval |
Citation fidelity benchmark. |
|
| 2 |
LitLLM |
arXiv'24 |
Retrieval |
LLM + academic database integration. |
|
| 3 |
LitSearch |
arXiv'24 |
Retrieval |
Retrieval precision benchmark. |
|
| 4 |
PaperQA2 |
arXiv'24 |
Retrieval |
Matches/exceeds expert on 3 tasks; 70% contradiction validation. |
|
| 5 |
OpenResearcher |
EMNLP'24 |
Retrieval |
RAG + graph traversal for literature exploration. |
|
| 6 |
PaSa |
arXiv'25 |
Retrieval |
Agentic multi-step iterative retrieval. |
|
| 7 |
Search, Inspect, Fetch |
arXiv'26 |
Retrieval |
Structure-aware Boolean retrieval for deep-research agents |
|
| 8 |
Rubric Reranker |
arXiv'26 |
Retrieval |
Trains document rerankers against search rubrics |
|
| 9 |
Personalized DR Refinement |
arXiv'26 |
Retrieval |
Graph-scaffolded evidence grounding for query refinement |
|
| Survey & Related Work Generation |
| 10 |
ChatPaper |
GitHub'23 |
Generation |
19K+ GitHub stars; arXiv summarization tool. |
|
| 11 |
PaperQA |
arXiv'23 |
Generation |
8K+ GitHub stars; RAG for scientific Q&A. |
|
| 12 |
AutoSurvey |
arXiv'24 |
Generation |
First end-to-end LLM survey drafting system. |
|
| 13 |
GPT Researcher |
GitHub'24 |
Generation |
26K+ GitHub stars; comprehensive report generation. |
|
| 14 |
LLMs for Lit. Review |
arXiv'24 |
Generation |
Hallucination analysis; models still generate errors. |
|
| 15 |
STORM |
arXiv'24 |
Generation |
Multi-perspective question-asking for outlines. |
|
| 16 |
Agentic AutoSurvey |
arXiv'25 |
Generation |
Multi-agent role decomposition. |
|
| 17 |
Citegeist |
arXiv'25 |
Generation |
Dynamic RAG pipeline on arXiv corpus. |
|
| 18 |
IterSurvey |
arXiv'25 |
Generation |
Iterative outline planning with stability checks. |
|
| 19 |
LiRA |
arXiv'25 |
Generation |
Multi-agent retrieval + verification + narrative. |
|
| 20 |
SurveyForge |
arXiv'25 |
Generation |
Outperforms AutoSurvey on outline quality. |
|
| 21 |
SurveyG |
arXiv'25 |
Generation |
Three-layer citation graph (Foundation/Dev/Frontier). |
|
| 22 |
SurveyX |
arXiv'25 |
Generation |
+0.259 content quality improvement; near expert level. |
|
| 23 |
InteractiveSurvey |
arXiv'25 |
Generation |
User-customizable reference categorization + outlines. |
|
| 24 |
CiteLLM |
arXiv'26 |
Generation |
Hallucination-free via trusted repository routing. |
|
| 25 |
STRUCTSURVEY |
arXiv'26 |
Generation |
Structured agentic retrieval for survey generation. |
|
| Deep Research Agents |
| 26 |
ASReview |
Nature MI'21 |
Deep Research |
Active learning; up to 95% effort reduction. |
|
| 27 |
CHIME |
arXiv'24 |
Deep Research |
Hierarchical organization of scientific studies. |
|
| 28 |
DeepResearch-Agent |
GitHub'25 |
Deep Research |
Hierarchical multi-agent; planner + sub-agents. |
|
| 29 |
DeerFlow |
GitHub'25 |
Deep Research |
Sub-agents with shared memory; sandboxed execution. |
|
| 30 |
OpenScholar |
Nature'26 |
Deep Research |
45M papers; +6.1% over GPT-4o, +5.5% over PaperQA2. |
|
| 31 |
AutoAgent |
arXiv'25 |
Deep Research |
Universal LLM compatibility; GAIA benchmark. |
|
| 32 |
Tongyi DeepResearch |
GitHub'25 |
Deep Research |
30.5B params (3.3B activated); SOTA on Deep Research. |
|
| 33 |
O-Researcher |
arXiv'26 |
Deep Research |
Multi-agent distillation + agentic RL. |
|
| 34 |
OpenResearcher (2026) |
arXiv'26 |
Deep Research |
54.8% BrowseComp-Plus; 97K+ trajectories. |
|
| 35 |
AREX |
arXiv'26 |
Deep Research |
Recursively self-improving deep research agent |
|
| 36 |
Predictive Navigation |
arXiv'26 |
Deep Research |
Pretraining objective tailored to research navigation |
|
| 37 |
On-Device DR (4B) |
arXiv'26 |
Deep Research |
Exposure bounds faithfulness; retrieval bounds coverage |
|
| 38 |
Carnot |
VLDB'26 |
Deep Research |
Interpretable, interactive, optimized deep-research execution |
|
| 39 |
Marginal Value Est. |
arXiv'26 |
Deep Research |
Stops search when added tokens no longer change the answer |
|
| 40 |
Retrieval-Aware Control |
arXiv'26 |
Deep Research |
Repairs stagnation in deep research reasoning |
|
| 41 |
Analogical Deep Research |
arXiv'26 |
Deep Research |
Retrieves historical analogies for foresight analysis |
|
| 42 |
Plato-Bio |
arXiv'26 |
Deep Research |
Verification-first biological novelty screening |
|
| 43 |
Albilich |
arXiv'26 |
Deep Research |
Proof-state orchestration with CAS integration |
|
| Retrieval and Synthesis Quality Assessment |
| 44 |
DeepScholar-Bench |
arXiv'25 |
Evaluation |
Coverage, coherence, factual accuracy benchmark. |
|
| 45 |
ReportBench |
arXiv'25 |
Evaluation |
100-prompt benchmark from 678 filtered survey papers. |
|
| 46 |
IDRBench |
arXiv'26 |
Evaluation |
100 tasks; interactive Deep Research evaluation. |
|
| 47 |
ScholarGym |
arXiv'26 |
Evaluation |
2,536 queries; query planning + tool invocation. |
|
| 48 |
SciNetBench |
arXiv'26 |
Evaluation |
18M papers; relation-aware retrieval <20%. |
|
| 49 |
Self-Evolving Retrieval |
arXiv'26 |
Retrieval |
Agent that adapts its own literature-search policy over time |
|
| 50 |
MasterSet |
arXiv'26 |
Retrieval |
Large-scale must-cite citation recommendation benchmark |
|
| 51 |
DeepSurvey |
arXiv'26 |
Generation |
Evidence-constrained citation assignment; analytical depth |
|
| 52 |
AutoResearchBench |
arXiv'26 |
Evaluation |
Benchmark for complex multi-hop scientific literature discovery |
|
| 53 |
PaperMind |
arXiv'26 |
Evaluation |
Multimodal reasoning + critique over papers; rebuttal concerns |
|
| 54 |
DRNOISE |
arXiv'26 |
Evaluation |
Benchmarks agents in misleading evidence environments |
|
| 55 |
HiEviDR-Bench |
arXiv'26 |
Evaluation |
Hierarchical evidence aggregation in deep research |
|
| 56 |
SciExplore |
arXiv'26 |
Evaluation |
Scientific navigation through information integration |
|
| 57 |
WANDR |
arXiv'26 |
Evaluation |
Benchmark spanning wide as well as deep research |
|
| 58 |
QA-to-DR Bench |
arXiv'26 |
Evaluation |
Verifiable tasks grown by iterative task evolution |
|
Paper2Social
Posts crafted from the survey across X, LinkedIn, Reddit, and Mastodon — each tuned to its platform's tone, length, and audience.
A fully automated AI system can now generate a research paper for as little as $15. But under pressure, every frontier LLM still fabricates results. The capability-vs-integrity tension is real. Our survey of 200+ papers maps the boundary. #AIforScience #LLM
Three findings every research lab should know in 2026: 1. AI handles the mechanical — but research-level code plateaus near 37% success. 2. 95.8% of rejected papers are misclassified as acceptable by LLM reviewers. 3. The most successful auto-research systems converge on a 3-layer architecture: exploration, execution, verification. Read the full survey →
[D] We surveyed 200+ AI auto-research papers. Here's what works, what doesn't, and what to do about it. Tools are now generating papers in 2.3 hours at $15. But the failure modes are getting harder to see — not easier. Ideation looks novel until execution; LLM reviewers are systematically lenient; cost decouples from quality past a modest budget. Full breakdown of the capability boundary, by stage. AMA in comments.
2/8 The ideation-execution gap is real: LLM ideas score 5.38 on novelty → drop to 3.41 after a human implements them. Brilliant on paper, brittle in practice. We see the same shape across stages.
Up to 17.5% of CS papers already carry detectable AI modification. The community needs to shift from detection (a losing race) to declaration. We propose stage-by-stage disclosure norms in the playbook. #AIethics #scicomm
Excited to release the Practitioner's Playbook — a stage-by-stage guide on what to delegate to AI and what to keep under human ownership. For each of the 8 research stages: ✅ delegate, ⚠️ retain, ❌ key risk. Built from controlled experiments and 250+ papers. Free + open. @worldbench