Beneath the Diff: Diagnosing and Mitigating Algorithmic Mode Collapse in Code-Level Autonomous Research Loops
Code implementation of our submission "Beneath the Diff: Diagnosing and Mitigating Algorithmic Mode Collapse in Code-Level Autonomous Research Loops".
This repository contains:
- A reproducible code-level Autonomous Research Loop (ARL) framework that wraps a base training pipeline (a fork of Karpathy's
autoresearch) and is also compatible with Aider as an alternative proposer. - A four-axis diagnostic instrument for algorithmic mode collapse: surface diversity
S_t, semantic cluster countC_t, mechanism entropyH_t, and provenance similarityV_t, plus a faithfulness auditΔ_faith_t. - A drop-in mitigation, DAPS (Diversity-Aware Proposal Sampling), composed of Category-Coverage Reweighting (CCR), Persistent Edit Memory (PEM), and a Held-Out Validation Gate (HVG).
- Four NLP-relevant task instantiations (T1: GPT-2 pretraining; T2: instruction tuning; T3: GSM8K reasoning; T4: prompt optimization) and six baselines (vanilla AR, hi-temp, R-Diverse-A, PRISM-A, Reflexion, Random Search).
- Plotting / analysis scripts that reproduce all figures and tables in the paper.
# 1. clone & install (requires Python 3.10+ and a single NVIDIA GPU for real runs)
git clone https://github.com/<anonymized>/arl-mode-collapse.git
cd arl-mode-collapse
pip install -e .
# 2. (optional) download the provenance corpus (~1.1M arXiv cs.CL/cs.LG abstracts pre-2024)
python scripts/prepare_provenance_corpus.py
# 3. set your proposer-LLM credentials (Claude Opus 4.7 in the paper)
export ANTHROPIC_API_KEY=sk-...
# 4. dry-run on T2 with the mock executor (no GPU needed, ~2 minutes)
python scripts/run_experiment.py --task t2 --method daps --mock --iters 30
# 5. real run: 300 iterations of DAPS on T2 (Llama-3.2-1B + LoRA)
python scripts/run_experiment.py --task t2 --method daps --iters 300 --seed 0Outputs land in data/logs/<task>/<method>/<seed>/ and include:
trajectory.jsonl— one accepted edit per line (diff, description, category, metrics).diagnostics.csv— per-iteration values ofS_t,C_t,H_t,V_t,Δ_faith_t.pipeline_snapshots/— git history of the pipeline being edited.
To regenerate every figure/table in the paper from logged trajectories:
bash scripts/run_all_tasks.sh # ~6,200 A100-hours; not for the faint of heart
python scripts/plot_results.py --out figures/ # cheap; reads pre-computed logsarl-mode-collapse/
├── arl/ Core library
│ ├── loop.py The ARL loop (Algorithm 1 in the paper)
│ ├── proposer/ LLM proposers (Anthropic / OpenAI / vLLM) + Aider adapter
│ ├── executor/ Sandboxed pipeline executor with 5-min budget
│ ├── diagnostics/ S_t, C_t, H_t, V_t, faithfulness audit
│ ├── daps/ CCR, PEM, HVG
│ ├── tasks/ T1-T4 task definitions
│ ├── baselines/ Vanilla AR, HiTemp, R-Diverse-A, PRISM-A, Reflexion, RandSearch
│ └── utils/ Diff utilities, seeding, logging
├── configs/ YAML configs for each task / method
├── pipelines/ The training pipelines the agent edits
│ ├── t1_nanogpt/ Forked & trimmed from Karpathy's autoresearch
│ ├── t2_instrtune/ LoRA on Llama-3.2-1B + Alpaca-20k
│ ├── t3_reason/ Qwen-2.5-1.5B + GSM8K SFT
│ └── t4_prompt/ DSPy-style prompt program
├── scripts/ CLI entry points (run / plot / prepare)
├── tests/ pytest smoke tests for the loop & diagnostics
└── notebooks/ analysis.ipynb (reproduces all paper figures)
for task in t1 t2 t3 t4; do
for method in vanilla hitemp r_diverse prism reflexion randsearch daps; do
for seed in 0 1 2; do
python scripts/run_experiment.py --task $task --method $method --seed $seed --iters 300
done
done
done
python scripts/aggregate_results.py --out tables/table2.tex(T1 uses 5 seeds because per-trajectory variance is largest there.)
After running the trajectories above:
python scripts/plot_results.py --figure 1 --out figures/fig1_collapse.pdf
python scripts/plot_results.py --figure 2 --out figures/fig2_axes_t2.pdf
python scripts/plot_results.py --figure 3 --out figures/fig3_robustness.pdfpython scripts/run_ablation.py --task t2 --components ccr,pem,hvg --seeds 0,1,2python scripts/run_sensitivity.py --task t2 --grid configs/sensitivity_grid.yamlpython scripts/run_experiment.py --task t2 --method daps --proposer aider --seed 0Fixed across all tasks and frameworks; see configs/daps_default.yaml.
| Symbol | Value | Role |
|---|---|---|
τ_c |
1.0 | CCR temperature (rare-category reweighting) |
τ_m |
0.85 | PEM cosine-similarity threshold |
M |
200 | PEM FIFO buffer size |
K |
10 | HVG audit interval |
τ_h |
task-calib | HVG faithfulness threshold (80th percentile of ` |
N |
8 | candidates per proposal step (for CCR sampling) |
W |
20 | sliding window for S_t, C_t, H_t, V_t |
MIT. See LICENSE. The forked autoresearch training core in pipelines/t1_nanogpt/ retains its original MIT license from Karpathy (2026).
This codebase builds on Karpathy's autoresearch and reuses portions of its prepare.py and train.py. We thank the authors of R-Diverse (Li et al., 2026), PRISM (Mishra, 2026), Reflexion (Shinn et al., 2023), and Aider (Gauthier, 2024) for releasing their code, which made the baseline comparisons in §4.2 possible.