Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Beneath the Diff: Diagnosing and Mitigating Algorithmic Mode Collapse in Code-Level Autonomous Research Loops

Code implementation of our submission "Beneath the Diff: Diagnosing and Mitigating Algorithmic Mode Collapse in Code-Level Autonomous Research Loops".

This repository contains:

  1. A reproducible code-level Autonomous Research Loop (ARL) framework that wraps a base training pipeline (a fork of Karpathy's autoresearch) and is also compatible with Aider as an alternative proposer.
  2. A four-axis diagnostic instrument for algorithmic mode collapse: surface diversity S_t, semantic cluster count C_t, mechanism entropy H_t, and provenance similarity V_t, plus a faithfulness audit Δ_faith_t.
  3. A drop-in mitigation, DAPS (Diversity-Aware Proposal Sampling), composed of Category-Coverage Reweighting (CCR), Persistent Edit Memory (PEM), and a Held-Out Validation Gate (HVG).
  4. Four NLP-relevant task instantiations (T1: GPT-2 pretraining; T2: instruction tuning; T3: GSM8K reasoning; T4: prompt optimization) and six baselines (vanilla AR, hi-temp, R-Diverse-A, PRISM-A, Reflexion, Random Search).
  5. Plotting / analysis scripts that reproduce all figures and tables in the paper.

Quick start

# 1. clone & install (requires Python 3.10+ and a single NVIDIA GPU for real runs)
git clone https://github.com/<anonymized>/arl-mode-collapse.git
cd arl-mode-collapse
pip install -e .

# 2. (optional) download the provenance corpus (~1.1M arXiv cs.CL/cs.LG abstracts pre-2024)
python scripts/prepare_provenance_corpus.py

# 3. set your proposer-LLM credentials (Claude Opus 4.7 in the paper)
export ANTHROPIC_API_KEY=sk-...

# 4. dry-run on T2 with the mock executor (no GPU needed, ~2 minutes)
python scripts/run_experiment.py --task t2 --method daps --mock --iters 30

# 5. real run: 300 iterations of DAPS on T2 (Llama-3.2-1B + LoRA)
python scripts/run_experiment.py --task t2 --method daps --iters 300 --seed 0

Outputs land in data/logs/<task>/<method>/<seed>/ and include:

  • trajectory.jsonl — one accepted edit per line (diff, description, category, metrics).
  • diagnostics.csv — per-iteration values of S_t, C_t, H_t, V_t, Δ_faith_t.
  • pipeline_snapshots/ — git history of the pipeline being edited.

To regenerate every figure/table in the paper from logged trajectories:

bash scripts/run_all_tasks.sh                  # ~6,200 A100-hours; not for the faint of heart
python scripts/plot_results.py --out figures/  # cheap; reads pre-computed logs

Repository layout

arl-mode-collapse/
├── arl/                       Core library
│   ├── loop.py                The ARL loop (Algorithm 1 in the paper)
│   ├── proposer/              LLM proposers (Anthropic / OpenAI / vLLM) + Aider adapter
│   ├── executor/              Sandboxed pipeline executor with 5-min budget
│   ├── diagnostics/           S_t, C_t, H_t, V_t, faithfulness audit
│   ├── daps/                  CCR, PEM, HVG
│   ├── tasks/                 T1-T4 task definitions
│   ├── baselines/             Vanilla AR, HiTemp, R-Diverse-A, PRISM-A, Reflexion, RandSearch
│   └── utils/                 Diff utilities, seeding, logging
├── configs/                   YAML configs for each task / method
├── pipelines/                 The training pipelines the agent edits
│   ├── t1_nanogpt/            Forked & trimmed from Karpathy's autoresearch
│   ├── t2_instrtune/          LoRA on Llama-3.2-1B + Alpaca-20k
│   ├── t3_reason/             Qwen-2.5-1.5B + GSM8K SFT
│   └── t4_prompt/             DSPy-style prompt program
├── scripts/                   CLI entry points (run / plot / prepare)
├── tests/                     pytest smoke tests for the loop & diagnostics
└── notebooks/                 analysis.ipynb (reproduces all paper figures)

Reproducing the paper

Main table

for task in t1 t2 t3 t4; do
  for method in vanilla hitemp r_diverse prism reflexion randsearch daps; do
    for seed in 0 1 2; do
      python scripts/run_experiment.py --task $task --method $method --seed $seed --iters 300
    done
  done
done
python scripts/aggregate_results.py --out tables/table2.tex

(T1 uses 5 seeds because per-trajectory variance is largest there.)

Diagnostic plots

After running the trajectories above:

python scripts/plot_results.py --figure 1 --out figures/fig1_collapse.pdf
python scripts/plot_results.py --figure 2 --out figures/fig2_axes_t2.pdf
python scripts/plot_results.py --figure 3 --out figures/fig3_robustness.pdf

Ablation

python scripts/run_ablation.py --task t2 --components ccr,pem,hvg --seeds 0,1,2

Hyperparameter sensitivity

python scripts/run_sensitivity.py --task t2 --grid configs/sensitivity_grid.yaml

Aider robustness

python scripts/run_experiment.py --task t2 --method daps --proposer aider --seed 0

Hyperparameters (DAPS)

Fixed across all tasks and frameworks; see configs/daps_default.yaml.

Symbol Value Role
τ_c 1.0 CCR temperature (rare-category reweighting)
τ_m 0.85 PEM cosine-similarity threshold
M 200 PEM FIFO buffer size
K 10 HVG audit interval
τ_h task-calib HVG faithfulness threshold (80th percentile of `
N 8 candidates per proposal step (for CCR sampling)
W 20 sliding window for S_t, C_t, H_t, V_t

License

MIT. See LICENSE. The forked autoresearch training core in pipelines/t1_nanogpt/ retains its original MIT license from Karpathy (2026).

Acknowledgements

This codebase builds on Karpathy's autoresearch and reuses portions of its prepare.py and train.py. We thank the authors of R-Diverse (Li et al., 2026), PRISM (Mishra, 2026), Reflexion (Shinn et al., 2023), and Aider (Gauthier, 2024) for releasing their code, which made the baseline comparisons in §4.2 possible.

About

Code for our submission 'Beneath the Diff: Diagnosing and Mitigating Algorithmic Mode Collapse in Code-Level Autonomous Research Loops'

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages