Submitted 05 Jul 2026

ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes

Abstract

Large language models have made research ideation increasingly accessible, yet effective idea development requires more than generating candidate directions. Researchers must ground a problem in current literature, identify meaningful bottlenecks, differentiate from existing solutions, and evaluate risks before committing to implementation. We present ResearchStudio-Idea as a reusable skill suite for this first mile of research ideation. The suite includes Paper-Search, a standalone multi-source literature search skill; Scoop-Check, a standalone prior-art collision checker for novelty claims; and IdeaSpark, the end-to-end skill that composes evidence grounding, pattern-guided generation, collision retrieval, audit, and idea-card rendering into one workflow. IdeaSpark is constructed from a corpus of 1,947 machine learning conference papers collected from ICLR, ICML, and NeurIPS between 2021 and 2025, including Oral papers, a separately tracked high-citation subset, and rejected submissions. Analysis of these outcomes reveals 31 recurring ideation sub-patterns, consolidated into 15 reusable ideation patterns. Each pattern is operationalized as a structured card containing research contexts, bottleneck types, differentiation strategies, supporting precedents, and common failure modes. Given a research problem and an evidence bundle, IdeaSpark evaluates evidence readiness, reconstructs the surrounding research context, identifies unresolved bottlenecks, selects relevant patterns, instantiates one candidate direction, retrieves potentially conflicting prior work, and performs outcome-informed auditing. This workflow transforms reusable ideation patterns into traceable research proposals. Blind automated-judge evaluations show that IdeaSpark consistently produces stronger research proposals than no-skill and generic-skill baselines while maintaining competitive novelty.

View Paper
PDF

AI Overview

Our new overview generator adds more detail and page citations

Grounding AI Research Ideation in Empirical Evidence

The initial stage of research, often called the "first mile," involves identifying a viable problem and proposing a defensible solution. While large language models (LLMs) have become proficient at summarizing literature and suggesting hypotheses, they often struggle with the "novel-but-empty" problem. This refers to generated ideas that appear creative due to their vagueness but lack the depth, feasibility, or strategic grounding required for high-impact scientific work.

ResearchStudio-Idea is a suite of skills designed to bridge this gap by anchoring the ideation process in the actual outcomes of major machine learning (ML) conferences like ICLR, ICML, and NeurIPS. Unlike generic ideation tools that rely solely on the internal knowledge of an LLM, this system utilizes a corpus of successful, highly cited, and—importantly—rejected papers to derive recurring innovation patterns and common failure modes.

Overview of the ResearchStudio-Idea data construction and IdeaSpark pipeline Figure 1: The system's workflow begins with an empirical analysis of conference papers to extract ideation patterns, which are then used by the IdeaSpark pipeline to ground, generate, and audit new research proposals.

Extracting Strategic Signatures from Conference Outcomes

To understand what makes research effective, the authors analyzed 1,947 unique papers from the 2021–2025 conference cycles. These papers were categorized into three groups:

  1. Oral Papers: Those selected for top-tier presentation by conference program committees.
  2. High-Cited (HC) Papers: The most influential papers in terms of community adoption.
  3. Rejected Papers: Submissions that did not meet the bar for acceptance, providing a critical "negative signal" for ideation.

A key challenge in analyzing such a diverse set of papers is separating the topic (e.g., computer vision or reinforcement learning) from the underlying research strategy. To solve this, the researchers employed a two-stage abstraction process. First, they extracted structured fields such as the "innovation approach" and "why it is non-obvious." Second, they rewrote these fields into domain-agnostic language, replacing specific technical terms with generic placeholders.

For example, a paper about "optimizing transformer attention" might be rewritten as "refining a standard scoring mechanism to reduce computational bottlenecks." This abstraction allows a clustering algorithm to group papers based on their strategic moves rather than their subject matter.

The 15 Patterns of Machine Learning Innovation

Through unsupervised clustering, the study identified 15 high-level "ideation patterns" that define the strategic landscape of modern ML research. These patterns are not just descriptive categories; they are operational "recipes" for innovation.

One significant finding is that research is rarely driven by a single pattern. On average, a paper utilizes approximately 2.3 patterns, with a modal composition size of k=2k = 2. This suggests that high-quality ideation is a compositional process, where researchers might "Audit and Pivot an Assumption" while simultaneously "Decomposing for Differentiated Treatment."

The researchers discovered that rejected papers actually use the same vocabulary of 15 patterns as accepted ones. The difference between success and failure often lies in execution rather than the choice of strategy. This led to the creation of "Operational Cards" for each pattern, which document not only success conditions but also specific failure modes derived from the reviews of rejected submissions.

Clustering of research strategies across Oral, High-Cited, and Rejected papers Figure 2: Strategic clusters showing the distribution of acceptance classes. Most clusters contain a mix of Oral, HC, and Reject papers, highlighting that the strategy itself is less important than its execution.

Differentiation in Success: Committees vs. Communities

The study reveals a distinct divergence between what conference program committees (PCs) value and what the broader scientific community adopts.

  • Committee Preference (ΔOR\Delta_{OR}): Peer reviewers often favor structural insights and theoretical pivots. A pattern like "Audit and Pivot an Assumption" is strongly associated with Oral acceptance.
  • Community Adoption (ΔOH\Delta_{OH}): The community at large often gravitates toward usable infrastructure and unifying frameworks. "Unify Heterogeneous Inputs into One Space" is a pattern that shows high citation impact even if it is not always favored by Oral selection committees.

"Reframe as a Solvable Object" emerged as a consistently high-performing pattern for both acceptance and long-term impact, indicating its value as a core research strategy.

The IdeaSpark Pipeline: Grounding and Auditing

The ResearchStudio-Idea suite culminates in a system called "IdeaSpark," which operationalizes these findings into a five-phase workflow.

Phase 1: Literature Grounding and Bottlenecks

Instead of brainstorming in a vacuum, the system starts with a multi-source search across arXiv, OpenReview, and Semantic Scholar. It builds a method-lineage tree to identify "bottlenecks"—structural gaps in existing work that prevent progress.

Phase 2: Pattern-Guided Generation

Based on the identified bottleneck, the system selects 1–3 appropriate ideation patterns. It then instantiates a candidate research direction, specifying the core mechanism, the rationale for why it should work, and how it differs from adjacent literature.

Phase 3: The Quality Gauntlet

This is a critical departure from standard LLM ideation. The system performs a targeted search to check for "prior-art collision" (ensuring the idea hasn't already been published). It then subjects the idea to an audit based on the "Reject Lessons" extracted from the conference corpus. If an idea matches a known failure mode or is too similar to existing work, the system issues a "revise" or "abandon" verdict.

Phase 4: Artifact Rendering

Finally, the system expands the idea into a structured research proposal, including LaTeX-formatted equations and implementability audits. This produces a defensible "idea card" that a human researcher can then evaluate and execute.

Evidence of Improved Idea Quality

The performance of IdeaSpark was tested against several baselines, including "bare" state-of-the-art LLMs (GPT-5.5 and Opus-4.8) and a generic automated ideation skill. The evaluation used "automated judges" to score 100 held-out problems across diverse ML domains.

Comparison of IdeaSpark against bare LLM baselines in terms of novelty and quality Figure 3: IdeaSpark achieves significantly higher quality scores while maintaining competitive novelty, avoiding the "novel-but-empty" failure mode of unguided LLMs.

The results indicated that bare LLMs often achieve high "novelty" scores simply because their ideas are so vague they do not collide with anything in the literature. However, their "quality" scores were low. IdeaSpark, by contrast, achieved the highest quality ratings and ranked first in 88% of cases. Its ideas were found to be more grounded, specific, and structurally sound.

To illustrate the technical depth the system can achieve, consider an example of a generated research proposal for improving Reinforcement Learning (RL) in long-context scenarios. The system might propose a "Phase-Stratified Advantage" calculation to concentrate credit on pivotal decision steps:

Ag=Rg−RˉqσqA_g = \frac{R_g - \bar{R}_q}{\sigma_q}

The proposal might further derive a self-score based on internal model confidence:

rg,sself=z-norms(1∣s∣∑t∈slog⁡πθ(xt∣x<t))r_{g,s}^{\text{self}} = \text{z-norm}_s \left( \frac{1}{|s|} \sum_{t \in s} \log \pi_\theta(x_t \mid x_{<t}) \right)

And finally blend these signals into a per-phase credit density dg,sd_{g,s}, where:

dg,s=softmaxs(β1contrastg,s+β2rg,sself)d_{g,s} = \text{softmax}_s (\beta_1 \text{contrast}_{g,s} + \beta_2 r_{g,s}^{\text{self}})

Subject to the constraint that:

∑sdg,s=1\sum_s d_{g,s} = 1

This leads to a stratified advantage assigned to each token in a specific phase ss:

Ag,s=Ag⋅(Sgdg,s)A_{g,s} = A_g \cdot (S_g d_{g,s})

Ensuring that the total credit is conserved:

Es[Ag,s]=Ag\mathbb{E}_s[A_{g,s}] = A_g

This level of mathematical and conceptual specificity is a result of the system's focus on evidence-grounded recipes rather than unconstrained generation.

Significance for the Research Community

ResearchStudio-Idea represents a shift toward more responsible and structured AI assistance in science. By treating ideation as a skill that can be grounded in empirical conference data, it provides several advantages:

  • Operationalizing Tacit Knowledge: It makes the "unwritten rules" of high-impact research explicit and accessible, especially for early-career researchers.
  • Learning from Failure: By incorporating data from rejected papers, it helps researchers identify and avoid common pitfalls before investing compute resources.
  • Modular and Auditable: The system's multi-phase workflow allows for human intervention and verification at every step, ensuring that the AI acts as a scaffold for human creativity rather than a replacement for it.

The work acknowledges that while it focuses on the "idea stage," the ultimate test of research remains implementation and peer review. However, by improving the quality of the "first mile," such systems have the potential to make the overall scientific discovery process more efficient and robust.

Can LLMs generate novel research ideas? a large-scale human study with 100+ NLP researchers
This paper provides the core motivation for ResearchStudio-Idea by empirically demonstrating that LLM-generated ideas are often judged as more novel but less feasible than expert ideas. This establishes the 'novel-but-empty' problem that the main paper's evidence-grounded and failure-aware approach aims to solve.
Chenglei Si, Diyi Yang, and Tatsunori Hashimoto. Can LLMs generate novel research ideas? a large-scale human study with 100+ NLP researchers, 2024. arXiv:2409.04109; checked May 10, 2026.
Sci-reasoning: A dataset decoding AI innovation patterns
This work is a direct methodological precedent, as it also induces innovation patterns from scientific papers. The main paper builds upon and differentiates from this by using a contrastive dataset of accepted, high-citation, and rejected papers to create 'failure-aware' patterns, rather than analyzing only successful papers.
Jiachen Liu, Maestro Harmon, and Zechen Zhang. Sci-reasoning: A dataset decoding AI innovation patterns, 2026. arXiv:2601.04577; checked May 10, 2026.
Towards end-to-end automation of AI research
This paper represents the broad 'AI scientist' systems that automate the entire research lifecycle. ResearchStudio-Idea positions itself as a more focused alternative, addressing the critical 'first-mile problem' of ideation rather than the full pipeline, making this a key point of comparison for the paper's scope and contribution.
Chris Lu, Cong Lu, Robert Tjarko Lange, Yutaro Yamada, Shengran Hu, Jakob Foerster, David Ha, and Jeff Clune. Towards end-to-end automation of AI research. Nature, 2026. Nature 2026; originally arXiv:2408.06292; checked May 10, 2026.
Novbench: Evaluating large language models on academic paper novelty assessment
This citation is highly relevant to the evaluation portion of the main paper, which introduces its own novelty-checking skill (Scoop-Check) and evaluates ideas on a novelty axis. Novbench is a large-scale benchmark for this exact task, providing essential context for the paper's approach to automated novelty and quality assessment.
Wenqing Wu, Yi Zhao, Yuzhuo Wang, Siyou Li, Juexi Shao, Yunfei Long, and Chengzhi Zhang. Novbench: Evaluating large language models on academic paper novelty assessment, 2026. arXiv:2604.11543; checked May 10, 2026.

Audio

0:00--:--
Transcript

Similar papers

Discussion