← Back to research
•·16 min read·industry

Autoresearch Tools

Comparison of 19 autoresearch tools for program experiments, deep research, and scientific discovery: architecture, evaluation, access, costs, and limitations.

Key takeaways

  • Program experiments, information research, and scientific discovery need different feedback and acceptance tests.
  • A strong score, cited report, or highly ranked hypothesis is evidence to inspect, not proof of correctness.
  • Google access remains restricted Preview, and Kosmos has changed its platform and subscription arrangements.
  • Evaluate total experiment cost and preserve histories, failures, held-out results, and human acceptance.

FAQ

What is autoresearch?

It is a family of agent-directed research loops that propose work, gather evidence or run experiments, and choose what to try next. This comparison separates program optimization, knowledge synthesis, and scientific discovery.

Do all autoresearch tools edit one file for five minutes?

No. That is a specific reference setup. Other tools maintain candidate populations, use variable evaluation budgets, research sources, or generate hypotheses that need external experiments.

Is Google Co-Scientist generally available?

The current Google Cloud documentation labels Co-Scientist and AlphaEvolve as restricted Preview offerings. Contact the account team for access; public documentation does not imply unrestricted availability.

Where does Tembo fit in an autoresearch workflow?

Tembo is adjacent engineering delivery infrastructure: agents can work with repositories and existing work systems to produce reviewable changes. This report does not claim a built-in integration with every research framework or count Tembo as a scientific experiment engine.

Executive Summary

Autoresearch is a family of workflows, not one product category with a single leaderboard. This comparison covers 19 tools across measurable program experiments, information research, and scientific discovery. Their outputs differ: a faster implementation, a sourced report, and a plausible scientific hypothesis require different acceptance tests.

The most useful current distinction is what closes the feedback loop. Karpathy's autoresearch evaluates a small training run; ShinkaEvolve maintains a population of candidate programs; deep-research agents retrieve evidence; scientific systems combine literature, code, and expert interaction.[1][2][3][4] None makes a convincing result automatically correct.

Commercial access also needs precise language. Google's current Co-Scientist and AlphaEvolve documentation describes restricted Preview access. Edison has reworked Kosmos around persistent R&D collaboration, so its older launch price and fixed-run framing should not be treated as the current universal offering.[4][5]

The recommendations below focus on task fit, architecture, evaluation, and operational cost. A quiet repository is not enough evidence to declare a project dead, and a funding announcement does not establish scientific reliability.

Market Definition

Members must document an agent-directed research or experiment workflow with inspectable output: measured candidates, sourced reports, hypotheses, or research artifacts. Public research frameworks and restricted commercial previews are both included, with availability labeled. This is a curated comparison, not a census of every coding agent or web-search assistant.

The previous 16 members remain covered. AlphaEvolve and ShinkaEvolve add substantial program-evolution approaches. EvoMap AutoResearch is a newly reviewed research workflow linked directly to its repository under the small/new-project exception; this review did not establish enough independent usage evidence for a full profile.

Pure benchmarks are contextual evidence rather than members. General agent frameworks and autonomous engineering tools overlap with the implementation work but are not automatically specialist research systems. Tembo's role is discussed separately in engineering delivery.

Sources were reviewed September 15–16, 2026. This is a documentation and published-experience review, not a shared benchmark or a claim of hands-on use of every tool.

Experiment Loops and Program Evolution

Market map

ToolWhat it changesFeedback and executionMain boundary
Karpathy autoresearchA small model-training programFixed training budget and validation bits per byteNVIDIA-oriented reference setup; hardware changes invalidate direct score comparisons[1]
pi-autoresearchCode against a configured metricMeasure script, durable journal, optional acceptance checksRuns with the agent's permissions; a branch is not a sandbox[6]
AutoAgentAgent harness, prompts, tools, routingHarbor-compatible evaluationsTasks are not bundled; provide the benchmark and execution setup[7]
AutoKernelGPU kernelsProfiling, correctness gates, end-to-end measurementLocal kernel speed and model-level benefit are separate[8]
autoresearch-at-homeTraining experiments shared across workersShared claims, successful and failed results, candidate codePublic coordination service availability was not verified[9]
Hyperspace AGIDistributed experiments across several domainsPeer coordination and shared experiment recordsProject labels and peer agreement do not establish AGI or scientific validity[10]
AlphaEvolvePrograms with automated objectivesGemini proposals, candidate archive, supplied evaluatorsRestricted Google preview; service and evaluation compute are separate[11][4]
ShinkaEvolveA population of programsModel mutations, novelty filtering, executable scoresOperate the framework and validate the evaluator yourself[2]

Karpathy autoresearch

The reference project deliberately keeps the experiment small: preparation and evaluation live separately from the editable training code, while the human supplies research direction in program.md. Its five-minute training budget excludes startup and compilation. The reported validation measure is useful within a controlled setup; the README explicitly warns against comparing results across different hardware.[1]

This is an approachable starting point for understanding the loop. It is not evidence that every task can run twelve meaningful experiments per hour, or that a five-minute improvement transfers to a full training job. Keep the original narrow contract clear when adapting the pattern.

pi-autoresearch

The Pi extension generalizes the loop through configurable instructions, measurement scripts, a persistent JSONL journal, and optional checks. Its confidence indicator compares improvement with variability after enough segment measurements; it is advisory, not a statistical guarantee or an automatic reason to discard a candidate.[6]

Choose it for a metric-driven workflow inside an agent environment you already understand. Establish repeatability, a bounded budget, and a recovery path first. The current extension package remains pi-autoresearch; the underlying Pi runtime's organizational move is not evidence that the extension itself changed to the same npm scope.[6]

AutoAgent and AutoKernel

AutoAgent moves the editable object up a level: it optimizes the agent harness itself. The repository expects user-supplied Harbor tasks and an execution environment, rather than shipping a universal benchmark suite. Its planned higher-level product should not be described as already available.[7] The main evaluation question is whether a harness gain survives tasks outside the search set.

AutoKernel starts with profiling, extracts a bottleneck, proposes Triton or CUDA changes, and checks correctness before measuring the wider workload. Its Amdahl-oriented scheduling recognizes that optimizing a small fraction of runtime may produce little overall benefit. The README's cycle-throughput figures are estimates, not a guaranteed rate on every GPU or kernel.[8] This is a useful distinction from loops that optimize an isolated microbenchmark and stop there.

Distributed experiments

autoresearch-at-home uses a shared coordination workspace to claim work, reduce duplication, exchange hypotheses, and publish failed as well as successful trials with candidate code. It can fall back to a solo workflow when coordination is unavailable.[9] That architecture is documented; this review did not verify a continuously operating public network.

Hyperspace AGI now spans a broader peer-to-peer runtime, private Pods, inference, and experimental workloads. Its repository describes propagation and shared state, but also cautions that experiment snapshots do not establish statistical significance.[10] Treat “breakthrough” labels and agreement among agent peers as project signals to inspect, not expert review. Its expansion does not justify reducing the whole project to either proven AGI or an abandoned training loop.

AlphaEvolve and ShinkaEvolve

AlphaEvolve is relevant when a correct program can be improved against an automated evaluator. Google's current skills workflow supports Python and documents a tested single-location editing scope; this is narrower than the general program-evolution idea.[12] A local evaluator does not mean the full managed service runs offline.

Klarna's first-hand collaboration report illustrates the value of hard constraints: a faster candidate was initially unusable because it violated deterministic execution, and changing the evaluator changed the search. The reported roughly twofold training gain belongs to that experiment and its scale limitations, not every production workload.[13]

ShinkaEvolve exposes a framework for population-based search, combining archive sampling, model-generated mutations, and novelty checks.[2] Its CLI independently controls proposal and evaluation concurrency, which helps match a search to scarce evaluation hardware.[14] Local model and embedding routes are documented, but all configured roles need review; missing price metadata recorded as zero is not free inference.[15]

Deep Research and Knowledge Synthesis

Market map

ToolApproachUseful distinctionEvaluation focus
GPT ResearcherPlanner and parallel retrieval over web/local sourcesReport generation with configurable models and sourcesClaim support, missing evidence, and retrieval cost[3]
dzhng/deep-researchRecursive breadth/depth searchSmall TypeScript implementation using search and LLM servicesSearch coverage, follow-up quality, and service limits[16]
Tongyi DeepResearchSpecialized research model plus tool harnessReleased 30.5B-total/3.3B-active modelWeights, tools, and deployment dependencies are separate[17]
LangChain Open Deep ResearchConfigurable LangGraph research workflowSeparate model roles and search/MCP backendsEnd-to-end quality and actual per-task cost[18]
Skywork DeepResearchAgentSelf-evolving agent protocol/runtimeVersioned resources and propose/assess/commit/rollback lifecycleGeneral framework requiring a configured research workflow[19]

Retrieval workflows

GPT Researcher separates planning, source gathering, and report synthesis, with web, local-document, and MCP paths. Deep research adds iterative depth and breadth. The README's cost examples are tied to particular configurations; an old inexpensive run is not a standing price for every report.[3] Evaluate a few questions with known answers and missing-information traps before relying on generated citations.

dzhng/deep-research provides a compact recursive implementation in TypeScript. It collects learnings and follow-up directions through search, then produces a report or answer. Firecrawl and model configuration introduce their own limits and costs.[16] Its small codebase can be useful for adaptation; a quiet release history alone does not establish that the design stopped working.

Open Deep Research exposes distinct summarization, research, compression, and final-writing model roles. Its repository includes evaluation with DeepResearchBench and cost tables covering a 100-task benchmark.[18] Those totals must not be relabeled as the cost of one research query. Review all stages when changing a model, because a stronger final writer cannot recover sources the search stage never found.

Specialized models and evolving runtimes

Tongyi DeepResearch pairs released model weights with a research agent harness. The published September 2025 checkpoint has 30.5 billion total and 3.3 billion active parameters with a 128K context window. Documented execution modes and dependencies include search, parsing, summarization, and code-execution services.[17] Downloading weights does not recreate the full evaluated system. Historical benchmark leadership should remain attached to its date and setup.

Skywork DeepResearchAgent now presents a broader self-evolution runtime. Its resource protocol versions prompts, tools, agents, environments, and memory, while a separate lifecycle proposes, assesses, commits, or rolls back changes.[19] It remains in this comparison as an evolution framework with research lineage, not as an interchangeable turnkey web-report product. Test the configured workflow rather than assuming the repository name establishes its current behavior.

Scientific Discovery and Research Artifacts

Market map

ToolCurrent scopeAccess/cost boundaryWhat needs validation
KosmosPersistent biopharma R&D collaboratorApply for new platform access; legacy terms are not current universal pricingTraceable claims, organization data, and experimental follow-through[5]
Google Co-ScientistLiterature-grounded hypotheses and proposalsRestricted PreviewRanking and review by agents do not replace experiments[4]
AutoscienceMira improves customer ML workloadsContact access; hosted or on-premises offeringCustomer-defined objectives and independently accepted results[20]
AutoResearchClawStaged research-to-paper workflowRepository plus configured models/computeHuman gates, experiment validity, and citation relevance[21]
AI ScientistComputational experiments and paper generationRunnable code under a custom source licenseMethodological rigor, generated code, and mandatory disclosure[22][23]
EvoMap AutoResearchIdeas through experiments and reviewable evidenceNew repository; bring compatible model services and execution environmentPreserved logs and cross-model review are aids, not correctness guarantees[24]

Kosmos

Edison's May announcement replaced the simple fixed-run story with a persistent, interactive R&D collaborator. It describes integration with organizational data and ongoing feedback, with access expanding according to capacity. Initial Incyte deployment focuses on target discovery, target validation, and translational biology; broader R&D deployment is an ambition, not an already-completed rollout.[5]

The same announcement closed founding subscriptions to new customers and changed academic credit arrangements. This review did not verify a current public price that could replace the old blanket $200/run claim.[5] A June partnership with Population Health Partners adds a stated company-creation and regulatory-authoring direction, with prospective outcomes clearly separated from delivered clinical results.[25]

Google Co-Scientist

The documented system assigns generation, reflection, ranking, evolution, proximity, and meta-review roles under a supervisor. Its output is a refined hypothesis or research proposal rather than automatic experimental truth.[4] Google's May Gemini for Science announcement described gradually available experiments and enterprise private previews; it did not establish general availability for every scientist.[26]

Choose this kind of system when the bottleneck is developing and prioritizing research directions. Plan separately for access, domain review, and the experiments that could reject its most convincing proposal.

Autoscience

Mira is positioned around improving ML quality, speed, cost, or footprint against customer-defined evaluations, with hosted and on-premises deployment offered.[20] The company's introduction reports a Kaggle silver-medal result with initial human setup.[27] That is a specific vendor-reported result, not evidence that every customer's production model will improve or that no expert work is needed.

The useful buying questions concern baseline reproducibility, objective ownership, allowed changes, and acceptance testing. Demonstration counters on a product page should not be mistaken for independently measured deployment scale. No public self-service price was established in the pages reviewed.

AutoResearchClaw and AI Scientist

AutoResearchClaw documents a 23-stage pipeline with modes for different levels of human involvement, three review gates, and iteration paths for refining or changing direction. Citation checks and a verified-result registry make evidence inspection part of the workflow; they do not guarantee relevance or eliminate hallucinations.[21] Its paper evaluates experiment-stage research tasks and reports benefits from targeted human help.[28] The amount and location of that help should remain visible when comparing results.

AI Scientist v2 uses agentic tree search for computational experiments and produces figures and a manuscript. Its README warns about executing generated code, and v2's broader freedom does not imply it always beats a strong v1 task template.[22] The custom license requires prominent disclosure in generated scientific manuscripts and includes use restrictions; it should not be collapsed into an ordinary permissive-license row.[23]

Sakana's March Nature announcement is a methodology milestone. The same account acknowledges weak ideas, methodological problems, inaccurate citations, and figure mistakes. Workshop acceptance was followed by a planned withdrawal.[29] Peer review of a method is not certification of every future output.

EvoMap AutoResearch

This newer workflow preserves plans, code, metrics, failures, and review artifacts from idea generation through execution. It uses several distinct models for cross-review and emphasizes lower-cost pilots and negative results.[24] The direct GitHub link is intentional: the project is new to this review, and independent operational evidence was insufficient for a full product profile. Cross-model agreement is still a signal to inspect, not a substitute for correct experiments.

Architecture: What Must Survive the Loop

The shared pattern is broader than “one file, five minutes, keep or revert.” Some systems maintain candidate populations, some pursue search branches, and others generate hypotheses that require external validation. The following contract is useful across those designs:

LayerRequired artifactFailure to watch for
ObjectiveExplicit goal and hard constraintsImproving a proxy while breaking the real requirement
BaselineReproducible starting resultComparing against a noisy or misconfigured reference
ExecutionVersioned inputs and bounded environmentHidden resource changes or unintended permissions
EvidenceCandidate source, measurements, sources, and failuresA polished conclusion without reproducible support
AcceptanceIndependent tests and accountable reviewPromoting the search winner without checking transfer

This is an evaluation framework, not a claim that every member implements all five layers. The artifact matters more than whether the interface calls the process autonomous.

A concrete engineering example

PostHog's June account describes Pi and pi-autoresearch investigating slow ClickHouse queries in a disposable cluster with anonymized production-shaped data. The team organized work into campaigns, lanes, hypotheses, and experiments. It periodically retested candidates on the original larger query range rather than trusting improvements on a narrowed case.[30]

The reported fix reduced scanned granules by 62% on the benchmark query; the article explains why latency improvement varied by time range. Its proposed automated pipeline—discover queries, run investigations, deduplicate ideas, create reviewed changes—was described as work in progress, not an already-proven unattended service.[30] This is a first-hand case study with a useful harness design, not a general performance guarantee.

Measuring generalization

Emulated's September Autoresearch Bench uses time-bounded agent runs and separate grading, with private evaluation splits on many tasks. It reports cases where public and private scores diverge and excludes infrastructure-failed runs from leaderboard averages.[31] The lesson for a team evaluation is to record failures and held-out performance, not just the best visible score. A benchmark with filtered runs answers a different reliability question from an unattended production service.

Pricing, Safety, and Selection

Compare total experiment cost: model calls, retrieval, embeddings, candidate execution, failed attempts, storage, and expert review. A framework license does not pay for any of those. A per-run commercial quote is incomplete without the run's limits, available data, and acceptance criteria.

For code experiments, separate the editable candidate from the evaluator and limit credentials and network access. For information research, inspect whether the cited passage actually supports the claim. For science, check data provenance, experimental validity, and disclosure requirements. These are different review tasks even when the interface for all three is a chat box.

Starting shortlist: use a narrow reference loop to learn the pattern; evaluate pi-autoresearch for configurable coding experiments; compare AlphaEvolve and ShinkaEvolve for program search; evaluate GPT Researcher or Open Deep Research for configurable report generation; and assess scientific platforms against a specific domain workflow and access arrangement. Preserve unsuccessful trials so that the next experiment begins with evidence rather than repeated guesses.

Where Tembo Fits

The research loop and the engineering delivery workflow are connected but distinct. Tembo connects coding agents to repositories and work systems, supports automated engineering tasks, and brings their output into a reviewable change workflow.[32] That makes it relevant when a measured improvement needs tests, documentation, integration, and human review.

For example, a team could require an accepted optimization to include its benchmark artifact and correctness tests before assigning the surrounding integration work. This is a proposed workflow, not a claim of a built-in integration with every research framework. Tembo is discussed as an adjacent delivery platform rather than counted as a scientific experiment engine. Disclosure: I am Tembo's co-founder and CEO.

Outlook

The strongest direction is better evidence flow: inspectable experiments, durable histories, affordable evaluation, and clear human acceptance. The available sources do not justify predicting that closed products will eliminate open frameworks or that publication milestones settle the market.

Select the tool whose feedback loop matches the decision. A report needs supported claims; an optimization needs repeatable gains under constraints; a hypothesis needs an experiment capable of disproving it. Useful autonomy makes those checks easier to perform and preserves their results.


Research by Ry Walker Research • methodology

Sources