Skip to content

Repository files navigation

Experiment Agent

Version License: CC BY-NC 4.0 Sponsor

繁體中文版

A Claude Code skill for executing, monitoring, interpreting, and verifying experiments in academic research.

What's new in v1.2.0

  • manage mode checks each study state file with a program, scripts/check_study_state.py. A study moves on to participant recruitment and data collection only when the program reports READY, and its data goes to analysis only when the ethics status is READY. If the program cannot run, the agent applies the rules by hand and says so, and it does not move a study into data collection.
  • After an IRB approval, the reconfirmation items must be answered again, later than the approval, before the study is READY.
  • Moving a study into data collection now needs Python 3.9 or later and PyYAML; see Requirements for human studies below.
  • Study state files are checked more strictly, so a file that v1.1.0 accepted can be reported INVALID. After upgrading, run the checker on each study state file before you resume the study (CHANGELOG.md, Compatibility).

Full notes: CHANGELOG.md

What It Does

  • Runs code experiments — executes scripts (Python, R, etc.), monitors for stalls/crashes in real-time, collects results
  • Manages human studies — plans protocols, checks IRB ethics, tracks data collection progress
  • Interprets statistics — reads p-values, effect sizes, CIs; checks 11 types of statistical fallacies (Simpson's Paradox, survivorship bias, etc.)
  • Verifies reproducibility — re-runs experiments and compares results

Why It Exists

Lu et al. (2026, Nature) demonstrated an Experiment Progress Manager for autonomous AI research. This skill brings the same execute-and-monitor capability to human-in-the-loop academic workflows — without the risks of full automation.

Modes

Mode What It Does
run Execute code + monitor process
manage Plan + track human studies
validate Statistical interpretation + reproducibility check
plan Socratic dialogue to design experiments

Quick Start

  1. Clone this repo into your project or .claude/skills/
  2. Start a Claude Code session
  3. Try: "Run my analysis: Rscript analysis.R"

ARS Compatibility

This skill works independently. It also integrates optionally with Academic Research Skills (ARS):

  • Reads ARS Stage 1 output (RQ Brief, Methodology Blueprint) to pre-populate experiment design
  • Produces Material Passport-compatible output, including an explicit verification status, for ARS Stage 2 consumption
  • ARS requires zero modification — the user bridges manually

When to use with ARS

In the ARS pipeline, experiment-agent fits between Stage 1 (RESEARCH) and Stage 2 (WRITE):

ARS Stage 1 RESEARCH  →  you get RQ Brief + Methodology Blueprint
        ↓
  [pause ARS pipeline]
        ↓
  experiment-agent     →  plan → run/manage → validate → get analyzed or verified results
        ↓
  [resume ARS pipeline]
        ↓
ARS Stage 2 WRITE     →  write paper using your experiment results

Use experiment-agent when your research requires running experiments (code or human studies) before writing. If your paper is purely based on literature review or secondary data analysis, you don't need this — go directly from ARS Stage 1 to Stage 2.

How to load

Step 1: Clone this repo alongside your ARS project (or anywhere on your machine):

cd ~/Projects
git clone https://github.com/Imbad0202/experiment-agent.git

Step 2: When you need to run experiments, open a Claude Code session in the experiment-agent directory:

cd ~/Projects/experiment-agent
claude

Step 3: Paste the relevant ARS Stage 1 output (RQ Brief, Methodology Blueprint) into the session. The agent will auto-detect the ARS headings and pre-populate your experiment plan.

Step 4: After your experiments are done and validated, copy the output (which includes a Material Passport header and verification status) back into your ARS session to continue Stage 2.

Requirements for human studies

manage mode checks each study's state file with scripts/check_study_state.py, which needs Python 3.9 or later and PyYAML:

python3 -m pip install pyyaml

If pip stops with an externally-managed-environment error, follow that message's instructions to install PyYAML for the python3 on your PATH. Claude Code asks for permission before it runs the checker; you can choose not to be asked again. Without the checker, manage mode still plans studies and runs the ethics checklist, but it will not move a study into data collection.

You can also add this skill to any project via .claude/skills/ symlink — see Claude Code docs for skill installation.

Safety

  • Only executes commands you specify — never auto-generates or modifies your code. The one exception: manage mode runs this skill's read-only study state checker
  • Never auto-retries crashed experiments
  • Never touches raw participant data
  • Statistical interpretation describes, never concludes
  • Full list: see SKILL.md Safety Rules

License

CC-BY-NC 4.0

Author

吳政宜 Edward Cheng-I Wu


Changelog

See CHANGELOG.md

About

Claude Code skill for experiment execution, monitoring, statistical interpretation, and reproducibility verification

Resources

Stars

199 stars

Watchers

0 watching

Forks

Releases

Sponsor this project

Packages

Contributors

Languages