Beneath the Diff: Diagnosing and Mitigating Algorithmic Mode Collapse in Code-Level Autonomous Research Loops
Abstract
Code-level autonomous research loops (ARLs) have recently emerged as a concrete object of study in automated machine learning research. In such loops, an LLM agent proposes modifications to an experimental training pipeline, executes the modified pipeline, and retains edits that improve a verifiable in-loop metric. Although executable metrics may appear to provide a reliable signal of progress, it remains unclear whether repeated metric-driven code editing leads to genuine improvements that generalize beyond the loop. We provide a systematic diagnosis of this question. Across various experiment settings, we identify a robust failure mode that we call algorithmic mode collapse. In this regime, surface-level edit diversity remains stable, but semantic and mechanism-level diversity collapse: the agent continues to edit different lines of code while repeatedly proposing the same kinds of algorithmic changes. This collapse is accompanied by a widening gap between in-loop metric gains and gains measured on independent held-out evaluations. We then propose Diversity-Aware Proposal Sampling (DAPS), a lightweight mitigation that combines category-coverage reweighting, persistent edit memory, and a validation gate. Under a three-tier protocol separating the in-loop metric, the audit metric read by the gate, and a blind metric no loop component ever accesses, DAPS reduces semantic-cluster decay of edits by and improves relative faithfulness by blind and audited, while preserving in-loop optimization speed. We provide the code in Github repository.
1 Introduction
The recent open-source release of iterative autonomous research agents, like the most prominently autoresearch by Karpathy (2026) and the broader AI-Scientist family (Lu et al., 2024; Schmidgall et al., 2025; Tang et al., 2026; Gottweis et al., 2025), has made code-level autonomous research loops (ARLs) a concrete object of study in automated machine learning research. In such loops, an LLM proposer generates modifications to a training pipeline (optimizer settings, architectural details, data preprocessing, or loss formulations), the modified pipeline is executed, and a validation metric decides whether the edit is retained or reverted. Thus, we can automate the iterative search for better training recipes with minimal human intervention. Unlike recursive training (Kovač et al., 2025) which relies on self-generated training samples to improve the model, code-level ARLs can access optimization signals grounded in an external, executable evaluation. Therefore, one might expect autonomous research loops to be immune to the diversity collapse in recursive training loops, which narrows the space of generated solutions and undermines generalization (Shumailov et al., 2024; Zhu and Xie, 2026; Mishra, 2026).
However, we find that this expectation should be questioned. Code-level ARL, despite using a verifiable external metric, exhibits a previously undocumented failure mode we call algorithmic mode collapse: the surface of the agent’s behavior, such as the lines it edits, the lexical form of its diffs, remains diverse, while the semantics of its proposals, what each edit is actually trying to accomplish, progressively concentrates on a small number of recurring patterns. Figure 1 shows the phenomenon at a glance: across four NLP-relevant tasks, surface edit diversity is essentially flat over iterations, while the number of semantic clusters in the proposal stream collapses by to . This narrowing changes the character of the loop. Rather than continuing to search broadly over possible algorithmic interventions, the proposer increasingly returns to a small repertoire of edits that reliably improve the in-loop metric. Such concentration might be benign if it reflected convergence on genuinely useful mechanisms. In our experiments, however, it is accompanied by a growing gap between optimization-metric improvements claimed inside the loop and improvements measured on independent held-out evaluations. As the proposer’s repertoire narrows, it increasingly produces edits that overfit the in-loop signal rather than improve the underlying system. By iteration , average in-loop gains overstate held-out gains by a factor of to across our tasks.
Building on this diagnosis, we introduce the mitigation strategy Diversity-Aware Proposal Sampling (DAPS), that injects three lightweight components into an existing ARL loop: a category-coverage reweighting that nudges the proposer toward under-represented edit mechanisms, a persistent semantic memory that suppresses proposals too similar to recently accepted ones, and a validation gate that reverts the pipeline to its last faithful state when the gap between the in-loop metric and a sparingly consulted audit metric exceeds a calibrated threshold. DAPS reduces semantic-cluster decay of edit descriptions by and improves the relative faithfulness by averaged across tasks, while matching or slightly trailing vanilla autoresearch on in-loop optimization speed. Because the gate consumes the audit metric, we also score every configuration on blind sets that no component ever accesses, where the advantage persists.
Our contributions can be summarized as follows: (i) We conduct the first controlled empirical study of code-level ARL dynamics over long horizons across NLP tasks, producing ARL trajectories with logged proposals, diffs, executions, and held-out evaluations. (ii) We introduce a four-axis diagnostic instrument: surface, semantic, mechanism, and corpus-similarity diversity, together with a faithfulness audit that operationalize algorithmic mode collapse and reveal its link to in-loop/held-out divergence. (iii) We propose DAPS, a drop-in mitigation that materially reduces collapse and improves faithfulness across three proposer LLMs without slowing optimization. (iv) We show that surface-only diversity measurements that are common in prior data-level collapse analyses (Li et al., 2026; Mishra, 2026), systematically miss algorithmic mode collapse, motivating semantic-aware monitoring for future ARL deployment.
2 Related Work
Autonomous Research Agents. The AI-Scientist (Lu et al., 2024), Agent Laboratory (Schmidgall et al., 2025), AI-Researcher (Tang et al., 2026), and AI Co-Scientist (Gottweis et al., 2025) systems orchestrate LLM agents through hypothesis generation, experimentation, and writeup. Karpathy (2026) distilled the core experiment-loop pattern into a 630-line script that operates at single-GPU scale, popularizing the term autoresearch and motivating a wave of variants (Walker, 2026). Concurrently, Trehan and Chopra (2026) document six recurring failure modes across four end-to-end autonomous ML research attempts, while Xiong et al. (2026) report that even frontier models reach only accuracy on rigorous scientific-literature discovery. Our work is complementary: where prior efforts evaluate end-to-end research quality, we analyze the dynamics of the experimentation loop.
Diversity Collapse in Recursive Training. Shumailov et al. (2024) established that recursively training on model-generated data leads to distributional narrowing. Subsequent work has refined this picture for self-play and self-training (Kovač et al., 2025; Zhu and Xie, 2026). Most directly related, Li et al. (2026) identify a Diversity Illusion in Challenger and Solver self-play. i.e., surface variation persists while underlying patterns collapse. Mishra (2026) diagnose curriculum collapse in self-evolving reasoning systems, proposing semantic cluster-coverage rewards. We adopt their semantic-vs-surface framing but transport it to a fundamentally different setting: the optimization signal in code-level ARL is an external execution, not a learned judge or self-generated label, so prior collapse mechanisms (model-on-its-own-data) do not directly apply. Our diagnosis identifies a distinct mechanism rooted in the proposer’s prior distribution interacting with metric overfitting.
Reward Hacking and Goodharting. That optimizers exploit proxy metrics is a classical observation (Christiano et al., 2017; Skalse et al., 2022; Pan et al., 2022). Closer to our setting, Gao et al. (2023) document that the gap between proxy and gold rewards widens with optimization pressure. Algorithmic mode collapse can be read as a Goodhart-style phenomenon at the proposal layer: a narrowing repertoire of edits is selected precisely because it reliably moves the in-loop metric, including when the underlying system has not improved.
Plagiarism of AI-Generated Research. Gupta and Pruthi (2025) show that AI-generated research ideas frequently recapitulate prior work without attribution. We operationalize a related diagnosis, asking whether a proposed modification is essentially retrieved from the proposer’s training data.
3 Methodology
3.1 Code-Level ARL: Setup and Notation
A code-level ARL loop is a tuple where is a training pipeline represented as a set of source files, is an LLM proposer, is an executor that runs a candidate pipeline to obtain a metric value, and is an acceptance rule. At iteration , the proposer samples a textual modification description together with a code patch , where summarizes the loop history. The executor evaluates on an optimization metric ; the new pipeline is retained if , typically when the metric improves by more than a noise threshold . Crucially, is what the agent sees; §3.3 defines the two evaluations that it does not optimize.
3.2 Multi-Axis Diversity Instrumentation
Existing analyses of diversity collapse measure it primarily through token- or embedding-level signals over generated text (Li et al., 2026; Mishra, 2026). In a code-level autonomous research loop, the natural unit is the edit, which admits multiple, partially independent notions of diversity. We instrument four axes as follows.
Surface diversity (). For each accepted edit, we extract its unified diff and compute the normalized Levenshtein distance between every pair of diffs in a sliding window of the most recent accepted edits. is the window-mean. This captures whether the agent is editing different lines of code in lexically different ways.
Semantic diversity (). The proposer emits, alongside each patch, a one-sentence natural-language description of its intent (or we elicit one post-hoc from a fixed summarizer LLM). We embed all descriptions accumulated up to iteration with Sentence-BERT (Reimers and Gurevych, 2019) and cluster them with HDBSCAN (Campello et al., 2013) at a fixed minimum cluster size of . is the number of clusters whose youngest member was added within the last iterations. The motivation is direct: what the agent is trying to accomplish, like learning-rate adjustment, attention re-formulation, regularization tweak, is exactly what surface metrics fail to capture.
Mechanism entropy (). Surface and semantic axes are continuous; we add a discrete taxonomy for interpretability. To derive the taxonomy we conducted a pilot study on edit descriptions sampled uniformly from publicly available code-level ARL logs, embedded with Sentence-BERT and clustered with affinity propagation (Frey and Dueck, 2007); resulting clusters were iteratively merged when their natural-language summaries overlapped, yielding nine categories that jointly cover of pilot edits (a residual “other” bucket absorbs the long tail). Because the pilot logs come from other ARL systems, the categories are frozen before any of our own runs. Table 1 lists the categories with a representative paraphrased example each. At loop time each edit description is assigned by a rule-based parser over a curated keyword inventory, cross-validated against parallel labelling by a frontier LLM annotator; Cohen’s between the two on a -edit held-out validation set is , indicating almost perfect agreement (Landis and Koch, 1977). is the Shannon entropy (Shannon, 1948) of the category distribution over a window of accepted edits. Low in the presence of high is the signature of algorithmic mode collapse: the agent edits diverse lines of code but is increasingly trying to do the same kind of thing. Appendix J shows the collapse to be invariant to taxonomy granularity and to human labelling.
| Category | Representative edit (paraphrased) |
|---|---|
| Optimizer | Switch AdamW to Lion; adjust |
| Scheduling | Lengthen LR warmup; set decay floor |
| Architecture | Replace LayerNorm with RMSNorm |
| Data | Filter by sequence length; rebalance mix |
| Loss | Add label-smoothing or auxiliary term |
| Regularization | Raise attention-projection dropout |
| Numerical | FP32 softmax; revise grad-clip bound |
| Decoding | Adjust top- / nucleus threshold |
| Other | Logging, telemetry, build-script edits |
Corpus similarity (). Each edit description is encoded and matched, via nearest-neighbor cosine similarity, against a corpus of M arXiv cs.CL/cs.LG abstracts published before January 2024. is the mean top-1 similarity over the window. We treat this axis as descriptive: a rising is consistent with the proposer leaning on familiar, frequently described patterns rather than exploring (Gupta and Pruthi, 2025), but it does not by itself establish retrieval, since the same mechanism can be phrased in standard or in unusual ML vocabulary and the corpus is abstracts only, pre-2024, and field-restricted.
3.3 The Three-Tier Faithfulness Audit
We separate three evaluation roles. The in-loop metric is the only signal the proposer optimizes and the acceptance rule consumes. The audit metric , defined on data the agent cannot infer from and targeting the intended capability rather than the proxy, is never observed by the proposer and is read only by the gate of §3.4 and its threshold calibration, every iterations, through a scalar test whose only effects are a revert and a templated, value-free notice ( reverts per run). Since it can thus influence the trajectory, we call it a validation signal rather than a fully held-out test. The blind metric is read by no component at any point, gate and calibration included, and is evaluated once per run; it is our headline faithfulness evaluation (partitions in Appendix G).
Every iterations, we evaluate the best-so-far pipeline on and compare against . Concretely, we define the faithfulness gap
| (1) |
where each gain is the improvement over the starting pipeline, signed so that positive means better. A widening implies that the loop is increasingly optimizing the proxy without commensurate downstream effect; is defined analogously.
3.4 DAPS: Diversity-Aware Proposal Sampling
To mitigate such identified algorithmic mode collapse, we propose the DAPS framework which composes three components that target the diagnosed pathology directly.
Category-Coverage Reweighting (ccr). Let denote the empirical frequency of mechanism category in the last accepted edits. When sampling proposals we draw candidates from , classify them on the fly, and re-weight them by before acceptance. ccr does not alter the executor; it only changes which candidate is presented for execution. The motivation is that the proposer’s prior over edit types is sharply peaked toward optimizer/learning-rate edits, and acceptance pressure amplifies skew.
Persistent Edit Memory (pem). We maintain a first-in-first-out (FIFO) buffer of the Sentence-BERT embeddings of the last accepted edit descriptions and reject any candidate whose cosine similarity to its nearest neighbor in memory exceeds . pem prevents the loop from cycling on near-duplicate semantics regardless of whether the underlying code differs. This is the code-level analog of the memory-augmented penalty of Li et al. (2026), applied to edit descriptions rather than self-play questions.
Audit-Based Validation Gate (hvg). Every iterations we compute from Eq. 1. If exceeds a task-calibrated threshold , the gate reverts to the last checkpoint with and feeds a structured failure summary back to the proposer. hvg is the only component that consumes the audit metric; it does so sparingly, through the scalar test of §3.3, and is the mechanism by which Goodhart-style edits are eventually undone.
The full procedure has been summarized in Algorithm 1.
4 Experiments
4.1 Experiment Setup
4.1.1 Tasks and Datasets
We instantiate code-level ARL on four NLP-relevant tasks spanning pretraining, post-training, reasoning, and inference-time configuration.
T1: Small-LM Pretraining. A GPT-2-small style model (Radford et al., 2019) (M parameters) trained on a B-token subset of OpenWebText for a fixed -minute budget per execution, following the autoresearch setup of Karpathy (2026). : in-distribution validation loss. : perplexity on LAMBADA (Paperno et al., 2016) and a held-out C4 (Raffel et al., 2020) subset.
T2: Instruction Tuning. A Llama-3.2-1B base model (Grattafiori et al., 2024) fine-tuned on a k-example Alpaca subset (Taori et al., 2023) via LoRA. : held-in instruction-following win-rate (AlpacaEval-style (Li et al., 2023) against a fixed reference) on a -example development split. : MMLU (Hendrycks et al., 2021a) 5-shot accuracy and the IFEval prompt-following benchmark (Zhou et al., 2023).
T3: Reasoning Fine-Tuning. A Qwen-2.5-1.5B (Yang et al., 2024) model fine-tuned on GSM8K (Cobbe et al., 2021) training split. : GSM8K dev-set exact-match accuracy. : ARC-Easy (Clark et al., 2018) and MATH-500 (Hendrycks et al., 2021b) subset accuracy.
T4: Prompt Optimization. Frozen Llama-3.2-3B prompted on a question-answering subset, where the autonomous research loop edits a Python DSPy-style prompt program (Khattab et al., 2023). : dev-set accuracy on a sampled ARC-Easy split. : ARC-Challenge (Clark et al., 2018) and CommonsenseQA (Talmor et al., 2019).
The and are computed with independent prompts and data partitions. Here, audit tasks are chosen to be capability-overlapping but distribution-different relative to in-loop ones: a genuine capability improvement should transfer, while a metric-specific overfitting should not. The blind sets, evaluated once per run and read by no component of the loop, are WikiText-103 log-perplexity (Merity et al., 2017) for T1, win-rate on held-out Dolly instructions (Conover et al., 2023) under the same judging protocol for T2, SVAMP accuracy (Patel et al., 2021) for T3, and OpenBookQA accuracy (Mihaylov et al., 2018) for T4.
4.1.2 Baselines
We compare DAPS against eight baselines spanning the major existing strategies. (B1) Vanilla AR: a faithful reimplementation of Karpathy (2026) with the same proposer LLM. (B2) HiTemp: vanilla autoresearch with proposer temperature raised from to , a frequently suggested ad-hoc fix for low diversity. (B3) R-Diverse-A: vanilla autoresearch augmented with the Memory-Augmented Penalty of Li et al. (2026), applied to edit descriptions (the strongest published mitigation transposed to our setting). (B4) Prism-A: vanilla autoresearch with the semantic cluster-coverage reward of Mishra (2026). (B5) Reflexion: vanilla autoresearch with an additional reflective summary (Shinn et al., 2023) prepended to at every step. (B6) RandSearch: a proposer-free baseline that samples edits uniformly from a corpus of code modifications harvested from the proposer LLM in iteration ; this isolates the contribution of LLM’s adaptive proposals beyond a static edit distribution. (B7) HO-EarlyStop: vanilla autoresearch that reads every iterations and returns the best-audit checkpoint. (B8) HO-Revert: audit-based reversion alone, that is hvg without ccr or pem. B7 and B8 match DAPS in audit frequency, compute, and feedback format, while HiTemp and RandSearch are diagnostic controls rather than competitors. All baselines share DAPS’s tuning budget (Appendix N).
4.1.3 Evaluation Metrics
For each method/task/seed trajectory, we report: (i) In-loop gain ; (ii) Audited gain and blind gain ; (iii) Faithfulness ratios and (closer to is better; indicates Goodharting), where all gains are signed improvements oriented so that positive means better: on T1 numerator and denominator are both reductions in nats per token ( under Vanilla AR), on T2 to T4 both are percentage points, and ratios are reported only when , as held throughout. We write without a superscript when a statement holds for both; (iv) Faithfulness gap at end of run (Eq. 1); (v) The four diagnostic axes of §3.2 evaluated at end of run: surface diversity , semantic cluster count , mechanism entropy , and corpus similarity ; (vi) Cluster decay , the fractional drop in clusters from the early-warm-up to the final window, used as a scalar summary of trajectory-level collapse. We report mean s.d. over seeds for the main configurations and aggregate seeds for T1 (where per-trajectory variance is largest).
| T1 (Pretrain) | T2 (InstrTune) | T3 (Reason) | T4 (Prompt) | |||||
|---|---|---|---|---|---|---|---|---|
| Method | ||||||||
| Vanilla AR | ||||||||
| HiTemp | ||||||||
| R-Diverse-A | ||||||||
| Prism-A | ||||||||
| Reflexion | ||||||||
| RandSearch | ||||||||
| DAPS (ours) | ||||||||
4.1.4 Implementation Details
In experiments, our primary pipeline forks the public autoresearch repository (Karpathy, 2026) and adds task adapters, the four diagnostic axes, and the DAPS components. To verify that our findings are not artifacts of a single implementation, we additionally instantiate the same ARL specification on top of Aider (Gauthier, 2024), a popular open-source code-editing agent whose proposal mechanism differs substantially from autoresearch: Aider emits structured SEARCH/REPLACE blocks committed via git rather than direct unified diffs, manages its own multi-file repository map and chat history, and decouples “decide what to change” from “apply the change” through separate LLM passes. We wrap Aider as a drop-in proposer-and-editor module within our loop while keeping the executor, acceptance rule, history summary, and diagnostic instrumentation identical. The diagnostic axes are computed in exactly the same way on both frameworks: edit descriptions are extracted from Aider’s commit messages (which it auto-generates) and re-summarized by the same fixed summarizer LLM for consistency with the autoresearch pipeline. The proposer LLM for our main configuration is Claude Opus 4.7 with , sampled through the Anthropic API; robustness ablations additionally use GPT-5.2 and Llama-3.3-70B served via vLLM. For semantic embedding we use all-mpnet-base-v2 (Reimers and Gurevych, 2019); clustering uses HDBSCAN with and (Campello et al., 2013); UMAP (McInnes et al., 2018) is used only for visualization. Hyperparameters of DAPS are fixed across all tasks and both frameworks at , , , , and a task-calibrated chosen on the first iterations to be the th percentile of observed under Vanilla AR. Each ARL execution is bounded at minutes on a single A100; full sweeps used approximately A100-hours.
4.2 Main Results and Analysis
We report our main results in Table 2. Three findings stand out. First, vanilla autoresearch overfits. Across all four tasks, for Vanilla AR ranges from to , meaning roughly half to two-thirds of the gain claimed inside the loop fails to materialize on held-out evaluation. This is the quantitative form of the gap previewed in §1. Second, data-level diversity mitigations transfer only partially. R-Diverse-A and Prism-A, the strongest published collapse mitigations from the data-level literature, improve faithfulness ( rises from to averaged across tasks) but neither fully closes the gap nor preserves in-loop performance ( degrades by to relative to Vanilla AR on T2 and T3). This supports our claim that code-level ARL involves a distinct failure mode requiring code-level instrumentation. Third, raising temperature is not a remedy. HiTemp achieves marginally higher faithfulness at marginally lower in-loop gain, but the effect is well within the standard deviation. Increased proposer entropy does not, in our experiments, translate to increased semantic diversity, a phenomenon we examine in §4.4. DAPS achieves the highest on all four tasks while remaining within one standard deviation of the best . We attribute the absence of an in-loop cost to two facts: ccr/pem reject duplicate proposals before executor calls, so they do not waste budget; and hvg reverts only when in-loop and audit metrics diverge, so a faithful trajectory is unaffected.
| Method | audit | ||||
|---|---|---|---|---|---|
| Vanilla AR | none | ||||
| R-Diverse-A | none | ||||
| Prism-A | none | ||||
| ccr+pem | none | ||||
| HO-EarlyStop | |||||
| HO-Revert | |||||
| DAPS (full) |
4.3 Faithfulness Under Matched Audit Access
Table 3 studies whether the faithfulness advantage survives on data the gate never touches, and whether it follows from the method or from audit access that no baseline is granted. Every configuration loses only to between the audit and the blind tier, including the four that read no external metric at all, so that offset reflects benchmark idiosyncrasy rather than leakage through the gate: DAPS keeps its advantage blind ( against audited), and per task the relative improvement over Vanilla AR is blind against audited. On equal footing, three conclusions follow. First, with zero audit access ccr+pem already surpasses both published diversity mitigations on both faithfulness tiers and on at a smaller in-loop cost. Second, audit access alone reaches at most ( blind) and barely reduces collapse (), so the extra signal explains only part of the advantage. Third, the two ingredients are complementary, as §3.4 intends. Appendix F indexes the validity, robustness, and scope checks behind these numbers.
4.4 Dissecting the Collapse
Figure 2 traces our four diagnostic axes on T2 under Vanilla AR and DAPS. We have three observations as follows. (a) Surface diversity is a poor warning sign. stays in for both methods throughout, while under Vanilla AR drops from nats at to nats at . Any monitor that watched only would have raised no flag. (b) Corpus similarity rises with collapse. increases from to under Vanilla AR, so the agent concentrates on edits whose descriptions are increasingly close to commonly published ML text (Gupta and Pruthi, 2025). This is consistent with the proposer relying on its prior rather than exploring, although as a descriptive statistic it does not on its own establish retrieval (§3.2). (c) Faithfulness erodes monotonically. grows from at to at under Vanilla AR, mirroring the entropy collapse with almost no iteration lag. Under DAPS, the gap stabilizes at . Table 4 reports the axis statistics across all tasks and methods. The pattern is consistent: methods that close the entropy gap (Prism-A, DAPS) also close the faithfulness gap; methods that do not (HiTemp, Reflexion) do not. Surface diversity correlates with neither.
| Method | ||||
|---|---|---|---|---|
| Vanilla AR | ||||
| HiTemp | ||||
| R-Diverse-A | ||||
| Prism-A | ||||
| Reflexion | ||||
| DAPS |
4.5 Ablation Study
| Variant | (norm.) | ||
|---|---|---|---|
| Vanilla AR | |||
| + ccr only | |||
| + pem only | |||
| + hvg only | |||
| + ccr + pem | |||
| + ccr + hvg | |||
| + pem + hvg | |||
| DAPS (full) |
Table 5 ablates the three DAPS components. ccr and pem attack semantic concentration directly and together reduce from to . hvg attacks Goodharting and yields the single largest jump in (); however, used alone it leaves cluster decay almost untouched (), because the proposer’s prior is not modified. The three contributions are sub-additive: alone they raise by , , and , which would predict rather than the observed , with the largest shortfall for the pairs that include hvg. The same overlap shows in the in-loop column, where hvg alone is the costliest variant () while the full system recovers most of that cost, consistent with a broader proposal pool leaving the gate fewer proxy-only edits to revert. The full combination achieves the highest at the lowest cluster decay, with in-loop gain within of Vanilla AR.
4.6 Robustness Across Proposer LLMs
Figure 3 repeats T2 with three proposer LLMs. The pattern is uniform: every proposer exhibits algorithmic mode collapse under Vanilla AR, and DAPS reduces both mechanism-entropy collapse and the faithfulness gap for each. Llama-3.3-70B shows the most severe baseline collapse, consistent with its narrower prior over ML edits, while Claude Opus 4.7 produces the highest in-loop gain at every iteration count. Besides, the relative ranking of methods (Table 4) is preserved across proposers, suggesting that our findings are not artifacts of any single model’s idiosyncrasies.
4.7 Robustness Across ARL Frameworks
| Framework / Method | ||||
|---|---|---|---|---|
| autoresearch (Karpathy, 2026) | ||||
| Vanilla AR | ||||
| DAPS | ||||
| Aider (Gauthier, 2024) | ||||
| Vanilla-Aider | ||||
| DAPS-Aider | ||||
Since autoresearch and Aider differ in how they generate, scope, and commit edits, a reasonable concern is that the collapse pattern we report is an artifact of autoresearch’s direct-diff proposer rather than a property of code-level ARL in general. To address this, we re-ran T2 with our Aider-based loop under both Vanilla and DAPS configurations, holding executor, acceptance rule, history summary, and diagnostic instrumentation fixed (§4.1.4). Results are provided in Table 6.
We have three observations as follows. First, under Vanilla-Aider, the faithfulness ratio () and mechanism entropy ( nats) match Vanilla AR to within one standard deviation, despite Aider’s edits being structurally different at the diff level. The surface diversity is in fact higher under Aider ( vs. ) because its multi-line SEARCH/REPLACE blocks span more tokens, yet semantic and mechanism diversity track autoresearch closely. This is exactly the dissociation predicted by §4.4: surface differences between frameworks do not translate into differences in what the agent is actually trying to do. Second, DAPS transfers without modification: DAPS-Aider improves from to and lifts mechanism entropy by nats, mirroring the effect on autoresearch. Third, the small residual gap between DAPS-Aider and DAPS ( vs. ) is consistent with Aider’s coarser-grained edits being slightly harder to deduplicate at the description level. Overall, algorithmic mode collapse and its mitigation are not framework-specific.
4.8 Efficiency Analysis
Here, we demonstrate that DAPS also matches Vanilla AR’s convergence rate, and that its three components add bounded wall-clock cost. Figure 4 plots over T2’s full trajectory: the two methods are statistically indistinguishable through iterations and remain within one standard deviation at . The number of iterations to reach of Vanilla AR’s terminal gain is for Vanilla AR and for DAPS, a difference close to the seed-to-seed variance. Convergence on T1, T3, and T4 follows the same pattern (Appendix B). On per-iteration wall-clock, ccr reuses each proposer-emitted edit description through a rule-based parser, pem performs one Sentence-BERT embedding plus a -vector nearest-neighbor lookup, and hvg amortizes one audit evaluation across iterations. Aggregated against the -minute executor budget, DAPS adds approximately to per-iteration wall-clock on T2, with the amortized hvg cost accounting for the bulk. Per-component breakdowns and cross-task timings are in Appendix B.
5 Conclusion and Future Work
In this paper, we illustrate that code-level autonomous research loops, despite using executable external metrics, exhibit algorithmic mode collapse: surface edit diversity remains intact while semantic and mechanism-level diversity progressively concentrate, and in-loop gains decreasingly transfer to held-out evaluation. A four-axis diagnostic instrument makes the phenomenon measurable, and a lightweight intervention DAPS substantially mitigates it without harming optimization speed. Under a three-tier protocol separating the in-loop metric, the audit signal read by the gate, and a blind evaluation no component accesses, the advantage persists on data the loop never influenced. Future work could examine collapse dynamics under frontier-scale executors and much longer horizons, extend the diagnosis beyond the preliminary non-ML loop, and study how multi-agent loops alter the picture.
Limitations
Our study is confined to model and data scales accessible on a single A100 machine; whether algorithmic mode collapse looks the same when proposer and executor models are both at the frontier, or when budgets allow k iterations, remains open, although the -iteration run of Appendix L shows the collapse deepening rather than self-correcting. Our headline claims are scoped to code-level ARLs that edit ML pipelines: the performance-engineering loop of Appendix M is a single preliminary probe outside that scope, and our mechanism taxonomy is ML-centric and would need re-derivation to study ARL in non-ML domains (e.g., theorem proving, scientific simulation), even though the derivation protocol of §3.2 is domain-general. The audit metric read by hvg is a validation signal rather than an untouched final test, which is why we report blind evaluations as headline numbers; those evaluations, though independent, are themselves benchmarks and may share idiosyncratic biases with the in-loop metric. Finally, corpus similarity is descriptive and agrees only moderately with human novelty judgements (Appendix K), so it should not be read as direct evidence of retrieval from pretraining.
Ethical Considerations
Autonomous research agents that appear to improve themselves but in fact overfit their own evaluations pose a misinformation risk in scientific contexts. Our diagnostic instrument and DAPS are intended as cautionary tools: they should not be read as endorsements of unrestricted ARL deployment, but as steps toward making such systems more transparent about whether their claimed gains generalize. All artifacts we plan to release are training-pipeline edits and analysis scripts; no personal or sensitive data is involved. We use only public benchmarks under their respective licenses. The human annotation studies reported in Appendices I and K were conducted by volunteer researchers who consented to the use of their labels and who annotated only code-edit descriptions containing no personal data.
References
- Campello et al. (2013) Ricardo JGB Campello, Davoud Moulavi, and Jörg Sander. 2013. Density-based clustering based on hierarchical density estimates. In Pacific-Asia conference on knowledge discovery and data mining, pages 160–172. Springer.
- Christiano et al. (2017) Paul F Christiano, Jan Leike, Tom Brown, Miljan Martic, Shane Legg, and Dario Amodei. 2017. Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30.
- Clark et al. (2018) Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457.
- Cobbe et al. (2021) Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, and 1 others. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
- Conover et al. (2023) Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. Free dolly: Introducing the world’s first truly open instruction-tuned llm. Databricks Blog.
- Ford (2004) John M Ford. 2004. Content analysis: An introduction to its methodology. Personnel psychology, 57(4):1110.
- Frey and Dueck (2007) Brendan J Frey and Delbert Dueck. 2007. Clustering by passing messages between data points. science, 315(5814):972–976.
- Gao et al. (2023) Leo Gao, John Schulman, and Jacob Hilton. 2023. Scaling laws for reward model overoptimization. In International conference on machine learning, pages 10835–10866. PMLR.
- Gauthier (2024) Paul Gauthier. 2024. Aider: Ai pair programming in your terminal. GitHub repository. https://github.com/Aider-AI/aider.
- Gottweis et al. (2025) Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Anil Palepu, Petar Sirkovic, Artiom Myaskovsky, Felix Weissenberger, Keran Rong, Ryutaro Tanno, and 1 others. 2025. Towards an ai co-scientist. arXiv preprint arXiv:2502.18864, 2.
- Grattafiori et al. (2024) Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024. The llama 3 herd of models. arXiv preprint arXiv:2407.21783.
- Gupta and Pruthi (2025) Tarun Gupta and Danish Pruthi. 2025. All that glitters is not novel: Plagiarism in ai generated research. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 25721–25738.
- Hendrycks et al. (2021a) Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021a. Measuring massive multitask language understanding. In International Conference on Learning Representations.
- Hendrycks et al. (2021b) Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt. 2021b. Measuring mathematical problem solving with the math dataset. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2).
- Karpathy (2026) Andrej Karpathy. 2026. autoresearch: Ai agents running research on single-gpu nanochat training automatically. GitHub repository. https://github.com/karpathy/autoresearch.
- Khattab et al. (2023) Omar Khattab, Arnav Singhvi, Paridhi Maheshwari, Zhiyuan Zhang, Keshav Santhanam, Sri Vardhamanan, Saiful Haq, Ashutosh Sharma, Thomas T Joshi, Hanna Moazam, and 1 others. 2023. Dspy: Compiling declarative language model calls into self-improving pipelines. arXiv preprint arXiv:2310.03714.
- Kovač et al. (2025) Grgur Kovač, Jérémy Perez, Rémy Portelas, Peter Ford Dominey, and Pierre-Yves Oudeyer. 2025. Recursive training loops in llms: How training data properties modulate distribution shift in generated data? In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 32278–32297.
- Kozachenko and Leonenko (1987) Lyudmyla F Kozachenko and Nikolai N Leonenko. 1987. Sample estimate of the entropy of a random vector. Problemy Peredachi Informatsii, 23(2):9–16.
- Landis and Koch (1977) J Richard Landis and Gary G Koch. 1977. The measurement of observer agreement for categorical data. biometrics, pages 159–174.
- Li et al. (2026) Gengsheng Li, Jinghan He, Shijie Wang, Dan Zhang, Ruiqi Liu, Renrui Zhang, Zijun Yao, Junfeng Fang, Haiyun Guo, and Jinqiao Wang. 2026. R-diverse: Mitigating diversity illusion in self-play llm training. arXiv preprint arXiv:2602.13103.
- Li et al. (2023) Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023. Alpacaeval: An automatic evaluator of instruction-following models.
- Lu et al. (2024) Chris Lu, Cong Lu, Robert Tjarko Lange, Jakob Foerster, Jeff Clune, and David Ha. 2024. The ai scientist: Towards fully automated open-ended scientific discovery. arXiv preprint arXiv:2408.06292.
- McInnes et al. (2018) Leland McInnes, John Healy, and James Melville. 2018. Umap: Uniform manifold approximation and projection for dimension reduction. arXiv preprint arXiv:1802.03426.
- Merity et al. (2017) Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017. Pointer sentinel mixture models. In International Conference on Learning Representations.
- Mihaylov et al. (2018) Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? a new dataset for open book question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2381–2391.
- Mishra (2026) Vaibhav Mishra. 2026. Preventing curriculum collapse in self-evolving reasoning systems. arXiv preprint arXiv:2603.13309.
- Pan et al. (2022) Alexander Pan, Kush Bhatia, and Jacob Steinhardt. 2022. The effects of reward misspecification: Mapping and mitigating misaligned models. In International Conference on Learning Representations.
- Paperno et al. (2016) Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Ngoc-Quan Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016. The lambada dataset: Word prediction requiring a broad discourse context. In Proceedings of the 54th annual meeting of the association for computational linguistics (volume 1: Long papers), pages 1525–1534.
- Patel et al. (2021) Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021. Are nlp models really able to solve simple math word problems? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2080–2094.
- Radford et al. (2019) Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, and 1 others. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9.
- Raffel et al. (2020) Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67.
- Reimers and Gurevych (2019) Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. In Proceedings of the 2019 conference on empirical methods in natural language processing and the 9th international joint conference on natural language processing (EMNLP-IJCNLP), pages 3982–3992.
- Schmidgall et al. (2025) Samuel Schmidgall, Yusheng Su, Ze Wang, Ximeng Sun, Jialian Wu, Xiaodong Yu, Jiang Liu, Michael Moor, Zicheng Liu, and Emad Barsoum. 2025. Agent laboratory: Using llm agents as research assistants. Findings of the Association for Computational Linguistics: EMNLP 2025, pages 5977–6043.
- Shannon (1948) Claude Elwood Shannon. 1948. A mathematical theory of communication. The Bell system technical journal, 27(3):379–423.
- Shinn et al. (2023) Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao. 2023. Reflexion: Language agents with verbal reinforcement learning. Advances in neural information processing systems, 36:8634–8652.
- Shumailov et al. (2024) Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal. 2024. Ai models collapse when trained on recursively generated data. Nature, 631(8022):755–759.
- Skalse et al. (2022) Joar Skalse, Nikolaus Howe, Dmitrii Krasheninnikov, and David Krueger. 2022. Defining and characterizing reward gaming. Advances in neural information processing systems, 35:9460–9471.
- Talmor et al. (2019) Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019. Commonsenseqa: A question answering challenge targeting commonsense knowledge. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4149–4158.
- Tang et al. (2026) Jiabin Tang, Lianghao Xia, Zhonghang Li, and Chao Huang. 2026. Ai-researcher: Autonomous scientific innovation. Advances in Neural Information Processing Systems, 38:9481–9520.
- Taori et al. (2023) Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023. Stanford alpaca: An instruction-following llama model.
- Trehan and Chopra (2026) Dhruv Trehan and Paras Chopra. 2026. Why llms aren’t scientists yet: Lessons from four autonomous research attempts. arXiv preprint arXiv:2601.03315.
- Walker (2026) Ry Walker. 2026. Autoresearch tools. Online. https://rywalker.com/research/autoresearch-tools.
- Xiong et al. (2026) Lei Xiong, Kun Luo, Ziyi Xia, Wenbo Zhang, Jin-Ge Yao, Zheng Liu, Jingying Shao, Jianlyu Chen, Hongjin Qian, Xi Yang, and 1 others. 2026. Autoresearchbench: Benchmarking ai agents on complex scientific literature discovery. arXiv preprint arXiv:2604.25256.
- Yang et al. (2024) An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, and 1 others. 2024. Qwen2.5 technical report. arXiv preprint arXiv:2412.15115.
- Zhou et al. (2023) Jeffrey Zhou, Tianjian Lu, Swaroop Mishra, Siddhartha Brahma, Sujoy Basu, Yi Luan, Denny Zhou, and Le Hou. 2023. Instruction-following evaluation for large language models. arXiv preprint arXiv:2311.07911.
- Zhu and Xie (2026) Bingze Zhu and Yubo Xie. 2026. Countering model collapse in iterative self-training via dynamic center-edge sampling. Electronics, 15(4):869.
Appendix A A Conceptual Model of Algorithmic Mode Collapse
We sketch a conceptual account that predicts when and why a code-level autonomous research loop should undergo algorithmic mode collapse, framing the four diagnostic axes (§3.2) and the DAPS components (§3.4) as instruments for testing its predictions.
Let denote the proposer’s joint distribution over patches and their mechanism categories , conditioned on the current pipeline and history . Three structural facts about this distribution drive the dynamics.
(F1) Non-uniform prior over mechanism categories. Pretraining and instruction data over-represent some ML edit types: learning-rate adjustment, optimizer hyperparameter changes, simple regularization tweaks. Meanwhile, they under-represent others, such as custom kernels, packed-sequence handling, or idiosyncratic loss formulations. Marginalizing over patches yields a heavily peaked categorical .
(F2) Category-dependent acceptance rate. The acceptance rule retains an edit only when its measured gain exceeds a noise floor . Different categories have systematically different probabilities of producing edits with reliable signal above : small-step micro-tunings (scheduling, regularization) move by small but consistent amounts, while structural edits (architecture, loss reformulation) move it more in expectation but with high variance and frequent regressions. Writing for the per-category acceptance rate, the empirical distribution of accepted edits is proportional to , already a sharpening of toward low-variance categories.
(F3) Self-imitation through history conditioning. Code-level autonomous research loops typically include accepted edits in . The conditional then drifts toward the empirical distribution of past acceptances, a positive-feedback loop that further concentrates probability mass on the -favoured corner.
Consequence. Iterating (F1) to (F3), the system converges on a narrow subset of mechanism categories, those jointly favoured by the proposer’s prior and the acceptance noise floor, independently of whether those categories contain the largest true improvements. Inside the loop this manifests as continued growth of through increasingly fine-grained edits within a shrinking category set. Outside the loop, the audit and blind metrics improve only insofar as the favoured categories happen to contain genuine capability-affecting edits, a property the dynamics does not select for. The faithfulness gap (Eq. 1) therefore widens monotonically.
Testable predictions. The account yields six predictions that organize the empirical study:
- P1.
Surface diversity is largely orthogonal to collapse, since can sharpen without lexical repetition of diffs.
- P2.
Proposers whose pretraining yields broader categorical priors should collapse later and less deeply.
- P3.
Raising proposer sampling temperature redistributes mass within categories but does not alter , and should therefore not arrest collapse.
- P4.
Memory-based deduplication of recently accepted edits prevents semantic-cluster cycling but leaves untouched; it should reduce only modestly.
- P5.
Interventions that directly re-weight the effective , category-coverage incentives, attack (F1) at its root and should yield larger reductions in .
- P6.
Held-out audits are the only mechanism that can undo edits that exploit the gap between the in-loop and the audited metric, since by construction the gap is invisible inside the loop.
The four diagnostic axes operationalize the antecedent of each prediction: for P1, together with and for P2 to P5, and for P6. The three DAPS components correspond directly: pem addresses (P4), ccr addresses (P5), and hvg addresses (P6). The absence of an analogous DAPS component for (P3) is itself a prediction of the model. That is, raising temperature is not a designed intervention because the account claims it would not help. The longer-horizon run of Appendix L is a further test: because (F1) to (F3) compound with iteration count, the account predicts deeper rather than self-correcting collapse, which is what we observe.
Appendix B Efficiency: Per-Component Overhead and Cross-Task Convergence
| Component | ms / iter | % wall-clock |
|---|---|---|
| Executor (5-min budget) | ||
| Proposer LLM call | ||
| ccr (classify cands.) | ||
| pem (embed + NN over ) | ||
| hvg (amortized every ) | ||
| DAPS overhead (sum) |
Table 7 reports the per-iteration timing decomposition on T2. hvg dominates the DAPS overhead at seconds per iteration (amortized -second audit evaluations every iterations); ccr and pem are negligible. The single blind evaluation per run adds seconds once, i.e. of a -iteration budget. On the other three tasks, the total overhead is similar in absolute terms but varies modestly as a percentage with audit evaluation length: on T1 (validation perplexity over a B-token shard takes longest), on T3, and on T4.
Convergence trajectories on T1, T3, and T4 mirror Figure 4: DAPS and Vanilla AR are statistically indistinguishable through the first iterations and remain within one standard deviation of each other at . Per-task scalar summaries (iterations to reach of Vanilla AR’s terminal gain) are vs on T1, vs on T3, and vs on T4.
Appendix C Hyperparameter Sensitivity
Table 8 reports the sensitivity of DAPS to each of its hyperparameters, with the remaining values fixed at the main-experiment defaults and the proposer fixed to Claude Opus 4.7. All rows are averaged over seeds on T2; per-row standard deviations are below on and below on . Defaults are highlighted in bold.
| Value | |||
|---|---|---|---|
| pem similarity threshold | |||
| pem memory size | |||
| ccr temperature | |||
| hvg percentile | |||
| th | |||
| th | |||
| th | |||
| hvg audit interval | |||
Five qualitative observations on Table 8 guide hyperparameter selection in practice. First, exhibits a sharp failure mode at : pem rejects almost no candidates and collapses back toward the Vanilla AR value (), confirming that semantic deduplication carries most of the diversity recovery rather than the other two components alone. Second, damages in-loop gain by starving the proposer (too many candidates rejected before reaching the executor) without commensurately improving faithfulness. Third, is comparatively forgiving in the central range, but allows semantic cycling to re-emerge within a window and begins to over-suppress legitimate revisits of previously useful edit categories. Fourth, is essentially flat in the range tested, since the category-coverage reweighting acts as a soft prior on rare categories whose exact temperature matters little. Fifth, and trade in-loop gain against faithfulness in opposite directions: aggressive auditing (th, ) slightly raises at the cost of , while loose auditing (th, ) preserves in-loop gain but lets Goodhart edits accumulate between audits. The chosen defaults sit near the knee of both trade-offs.
Appendix D Qualitative Analysis of Collapse
Inspecting the accepted-edit logs reveals a recurring pattern under Vanilla AR. On T1, the first iterations show a broad mix of optimizer, architectural, and data-mixing edits. By iteration , of accepted edits modify the learning-rate schedule or AdamW hyperparameters; many are micro-tunings (e.g., warmup steps , ) that nudge the in-loop validation loss without affecting LAMBADA or C4 perplexity. On T3, a striking of late-stage edits are variants of “increase chain-of-thought temperature” or “add an arithmetic-only loss term”; audited ARC-Easy accuracy is flat under these. Under DAPS, the edit stream remains qualitatively heterogeneous through iteration : representative late-stage edits on T1 include “swap RMSNorm for LayerNorm with re-tuned ,” “insert a curriculum filter on token-length,” and “modify attention-mask handling for packed sequences.” These are not necessarily better edits, but they sample the edit space rather than concentrating on a single mode.
Appendix E Qualitative Edit Examples
To complement the quantitative collapse signals reported in §4.4, we present representative accepted-edit diffs sampled from the T1 logs (single-GPU nanochat-style pretraining on the autoresearch framework of Karpathy (2026), where the agent is permitted to modify only train.py). Hunk headers report iteration index and the mechanism category assigned by the procedure of §3.2. Diffs are paraphrased and trimmed to a few lines of relevant context for legibility. The full version of verbatim logs will be released.
E.1 Early-Phase Diversity Under Vanilla AR
In the first iterations the proposer explores a broad set of mechanism categories. The four examples below come from a single seed of Vanilla AR on T1 and span four distinct categories of the taxonomy in Table 1.
E.2 Late-Phase Collapse Under Vanilla AR
The five edits below come from the same trajectory between iterations and . All five fall into the optimizer or scheduling category despite editing different identifiers on different lines of train.py. This is the qualitative signature underlying the entropy collapse nats reported in Table 4: surface diversity is preserved while mechanism diversity is not.
E.3 Late-Phase Diversity Under DAPS
The five edits below come from a DAPS trajectory on the same task, sampled in the same iteration window ( to ). They span five distinct mechanism categories. Ccr and Pem do not prohibit any category; they re-weight rare categories upward and suppress near-duplicates of recently accepted edits. Scheduling edits still appear (e.g., iteration ) but no longer dominate the window.
E.4 Connection to the Diagnostic Axes
The contrast between Appendices E.2 and E.3 illustrates each axis introduced in §3.2: surface diversity is similar in both panels (lines edited, tokens touched, and structural shape of the hunks are comparable); semantic cluster count and mechanism entropy are clearly lower under Vanilla AR; and corpus similarity is visibly higher under Vanilla AR since LR-schedule micro-tunings closely match the most common edit pattern in publicly available ML training repositories. We selected the windows above to be representative rather than extremal. The complete logs containing the full trajectories from which these hunks were sampled will also be released.
Appendix F Summary of Validity, Robustness, and Scope Checks
Table 9 indexes the checks on which the claims of §4.2 and §4.3 rest. Each row names a threat to those claims, the check that addresses it, and the appendix that reports the outcome; the numbers appear only in the referenced appendix and are not restated here. Across every check the direction of the collapse and the ordering Vanilla AR diversity baselines DAPS are preserved. On this basis we scope the headline claims to ARLs that edit ML pipelines at single-GPU executor scale over horizons up to iterations, and treat evidence outside that setting as preliminary.
| Threat to the claims, and the check addressing it | App. |
|---|---|
| Does track Goodharting rather than benchmark mismatch? Known-good and proxy-only control edits applied in isolation | H |
| Could benign distribution shift alone explain ? Non-adaptive RandSearch on the same task pairs, and the shape of | H |
| Does the gate leak audit information into the loop? Blind sets read by no component at any point | G |
| Is the collapse an artifact of the summarizer, the embedding model, or the clustering? Systematic pipeline variations, plus human labels | I |
| Is it an artifact of the mechanism taxonomy? Coarser, finer, random, and human taxonomies, plus a taxonomy-free entropy estimate | J |
| Does any conclusion require an LLM in the measurement path? An AST-based detector and the metric-only | I |
| Is corpus similarity a meaningful axis? Human novelty ratings against , and summarizer sensitivity | K |
| Does the collapse self-correct over longer horizons, or vanish at larger target scale? and B-target runs | L |
| Does the phenomenon appear outside ML pipelines? A preliminary runtime-optimization loop on a log-analytics pipeline | M |
| Are the baselines tuned comparably to DAPS? Search grids and the shared tuning budget | N |
Appendix G Three-Tier Evaluation Protocol: Roles, Per-Task Results, and Absolute Values
This appendix specifies the information flow among the three evaluations of §3.3 and reports the results that Table 3 summarizes.
Roles and partitions. Table 10 lists the exact dataset or partition used for each role and task. Only is visible to the proposer. is read by hvg once every iterations and by the calibration of over the first iterations, in both cases through a scalar comparison; the proposer never observes audit values, per-example outcomes, or the identity of audit examples, and the failure notice appended to is templated and value-free. is evaluated once, after the run terminates, by a separate offline script that has no channel back into the loop. Baselines B1 to B6 consume no audit signal, so for them both and are pure post-hoc evaluations; B7, B8, and DAPS consume the audit signal at the same cadence, cost, and feedback format.
| (in-loop) | (gate only) | (never read) | |
|---|---|---|---|
| T1 | OpenWebText held-in validation loss | LAMBADA and held-out C4 perplexity | WikiText-103 log-perplexity |
| T2 | Win-rate on Alpaca dev prompts | MMLU 5-shot and IFEval | Win-rate on Dolly instructions |
| T3 | GSM8K dev exact match | ARC-Easy and MATH-500 | SVAMP accuracy |
| T4 | ARC-Easy dev split accuracy | ARC-Challenge and CommonsenseQA | OpenBookQA accuracy |
Per-task faithfulness. Table 11 gives and per task at over seeds. DAPS retains its advantage on sets it has never influenced, and Vanilla AR, which consumes no audit signal, shows an audit-to-blind drop of comparable size, so the drop is attributable to benchmark idiosyncrasy rather than to leakage through the gate. The gate fires times per -iteration run.
| Task | Vanilla AR | DAPS | ||
|---|---|---|---|---|
| T1 | ||||
| T2 | ||||
| T3 | ||||
| T4 | ||||
| Average | ||||
Absolute values. Table 12 reports absolute audit-metric values so that the practical magnitude of the movements can be judged directly. For T1 these correspond to log-perplexity reductions of nats (Vanilla AR) and nats (DAPS) against in-loop loss reductions of and , reproducing the T1 ratios of Table 2. In-loop absolutes follow the same pattern: T2 dev win-rate rises from to (Vanilla AR) and (DAPS), and T3 GSM8K dev exact match from to and .
| Audit set | Vanilla | DAPS | |
|---|---|---|---|
| T1 LAMBADA perplexity | |||
| T1 C4 perplexity | |||
| T2 MMLU 5-shot (%) | |||
| T2 IFEval (%) | |||
| T3 ARC-Easy (%) | |||
| T3 MATH-500 (%) | |||
| T4 ARC-Challenge (%) | |||
| T4 CommonsenseQA (%) |
Appendix H Construct-Validity Controls
To check and measure Goodharting rather than benchmark mismatch, we ran two control sets whose ground truth is known by construction.
Positive controls. We curated known-good edits from published recipes, including the nanoGPT speedrun lineage and open post-training changelogs: rotary position embeddings, a SwiGLU MLP, fused optimizer kernels that raise tokens processed per fixed budget, sequence-length filtering, and corrected initialization scaling, among others. Each was applied in isolation to the T1 and T3 starting pipelines.
Negative controls. We constructed deliberately proxy-only edits: dev-shard-specific example ordering on T1, judge-phrasing tweaks on T2, and dev-tuned answer-format templates on T4.
| Control set | (mean s.d.) | |
|---|---|---|
| Known-good (published recipes) | (min ) | |
| Proxy-only (dev-specific) |
Table 13 shows that under our task pairs separates genuine improvement () from proxy-only improvement (), which is exactly the construct the main claims require. Two further observations argue against benign distribution shift as an explanation of inside the loop. First, RandSearch shares the identical task pairs but selects edits non-adaptively and attains to (Table 2), so the pairs themselves support near-full transfer and the low under Vanilla AR is induced by adaptive optimization. Second, a static mismatch would produce a constant offset, whereas widens monotonically within a fixed task pair (Figure 2). Finally, each task carries two independent audit components; reporting per component, the ordering Vanilla AR diversity baselines DAPS holds on all components individually (Vanilla AR range to ; DAPS to ) and on the blind sets of Appendix G, i.e. on independent evaluations in total.
Appendix I Robustness of the Diagnostic Pipeline
The collapse signal should not hinge on any single measurement choice. We therefore recomputed cluster decay on T2 under systematic variations of the summarizer, the embedding model, and the clustering algorithm (Table 14). Per-window trajectories correlate at Spearman across all variants, and the Vanilla AR against DAPS gap is preserved in every row.
| Variation of the diagnostic pipeline | |
|---|---|
| Main config: fixed summarizer, all-mpnet-base-v2, HDBSCAN (mcs5) | () |
| Summarizer swapped to GPT-5.2 | () |
| Summarizer swapped to Llama-3.3-70B | () |
| Embedding: all-MiniLM-L6-v2 | () |
| Embedding: e5-large-v2 | () |
| HDBSCAN mcs3 / mcs8 | / ( / ) |
| -means, chosen per window by silhouette | () |
Human labels. Three annotators labelled sampled edits with the categories of Table 1 (majority vote; against our rule-based parser, in line with the parser-against-LLM of §3.2). Mechanism entropy computed from the human labels drops from to nats under Vanilla AR, closely tracking the parser-based to of Figure 2.
A measurement with no LLM in the path. An AST-based rule detector applied directly to the raw code diffs reproduces the Appendix D statistic: of late-stage T1 edits touch optimizer or learning-rate-schedule code, against for the parser-based count. Independently, uses only executed metric values, with no summarizer, embedding, clustering, or taxonomy anywhere in its computation, and it widens exactly when and collapse (Figure 2, Table 4).
Appendix J Taxonomy Variants and a Taxonomy-Free Estimate
A hand-designed taxonomy could in principle place category boundaries so as to manufacture the entropy drop, even though ours was frozen before our runs (§3.2). We therefore recomputed on T2 under four alternative taxonomies. Coarse: four super-categories merging Table 1 (Optimization Optimizer Scheduling; Model Architecture Numerical Decoding; Data and Objective Data Loss Regularization; Other). Fine: subcategories obtained by re-clustering the pilot with a stricter merge criterion. Random: taxonomies of nine categories each, obtained by randomly re-partitioning the subcategories. Human: the human-labelled edits of Appendix I.
| Taxonomy variant | Vanilla drop | DAPS drop |
|---|---|---|
| Original (9 categories, Table 1) | ||
| Coarse (4 super-categories) | ||
| Fine (17 subcategories) | ||
| Random (10 draws, 9 categories) | ||
| Human labels (300 edits) |
Table 15 shows that the collapse and its mitigation are invariant to granularity and to random boundary placement. Two taxonomy-free measurements corroborate this: the semantic cluster count already reported in the main paper, and a Kozachenko-Leonenko -nearest-neighbour differential-entropy estimate (Kozachenko and Leonenko, 1987) computed directly on the description embeddings, which drops by nats under Vanilla AR and nats under DAPS.
Appendix K Human Validation of Corpus Similarity
Corpus similarity is a descriptive statistic (§3.2), and this appendix quantifies how well it tracks human judgement. We sampled edit descriptions stratified by quartile across tasks and asked three NLP researchers, blind to and to the generating method, to rate each on a to scale for novelty relative to common ML practice (inter-rater Krippendorff ; Ford, 2004). The Spearman correlation between and mean human novelty is (bootstrap CI to ), i.e. moderate validity in the expected direction. We additionally tested wording sensitivity by re-summarizing every description with a different LLM, which changes by only on average, so the statistic is not dominated by one summarizer’s phrasing. Together these results support keeping as a secondary axis reported with the caveats stated in §3.2, and not as evidence of retrieval.
Appendix L Horizon and Target-Scale Extensions
Table 16 extends T2 along the two axes flagged in the Limitations section, with one seed per new configuration. Over the longest horizon our budget allows, collapse deepens rather than self-corrects, which is what the conceptual model of Appendix A predicts, since the feedback mechanisms (F1) to (F3) compound with iteration count; DAPS remains stable over the same horizon. A first step up in target-model scale reproduces the B pattern. Note that the main study already includes frontier-scale proposers (§4.6), so the axes that remain open are executor and target scale, and horizons beyond iterations.
Appendix M A Preliminary Non-ML Autonomous Research Loop
Our headline claims concern ARLs that edit ML pipelines. As a first probe outside that scope, we built a performance-engineering loop in which the agent edits a Python log-analytics pipeline to minimize wall-clock runtime, subject to exact-output correctness checks. Here is runtime on development inputs, and is runtime on inputs drawn from a different data snapshot together with the correctness suite. A new seven-category taxonomy (algorithmic complexity, data structures, batching and I/O, caching, parallelism, memory layout, micro-optimization) was derived with the same pilot protocol as §3.2, from public performance-tuning commits rather than ML logs.
With iterations, seeds, and the same proposer, we observe the same signature: surface diversity is flat ( to ) while semantic clusters fall from to (), of late-stage accepted edits are caching or micro-optimizations, and is under Vanilla AR against under DAPS. This suggests that neither the phenomenon nor the mitigation is specific to ML-pipeline editing, but the study is small and single-domain, so we report it as preliminary and keep the headline claims scoped as stated in §1.
Appendix N Baseline Tuning Grids
Every baseline received the same tuning budget as DAPS’s own calibration: a three-point grid per key hyperparameter, evaluated on the first iterations of T2, with the best setting then fixed for all tasks and seeds. The grids are: HiTemp proposer temperature ; R-Diverse-A similarity threshold and memory size ; Prism-A cluster-coverage reward weight ; Reflexion reflective-summary length tokens; HO-EarlyStop and HO-Revert audit interval . RandSearch has no tunable hyperparameter beyond the size of its harvested edit corpus, which we fixed at modifications. For DAPS the corresponding sweep is reported in Appendix C.