This directory is the reproducible code release of SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment. SafeEvolve turns on-policy agent trajectories into two coupled artifacts���a global safety prompt and a hierarchical SkillBank—and then internalizes their use with harness-use SFT followed by harness-augmented GRPO.
SafeEvolve runs a closed-loop harness–policy co-evolution process:
- Agent–environment interaction. A policy model operates through a safety harness containing safety rules and retrieved skills. It issues tool calls, receives environment observations, and produces a complete trajectory.
- Safety harness evolution. Rollout evidence is grouped by risk and task metadata, distilled into prompt or SkillBank candidates, and accepted only after bounded edits pass safety and utility gates.
- Harness-augmented policy optimization. Curated multi-step rollouts are rejection-sampled for SFT, then GRPO trains the policy to use the evolved harness across benign, query, and injection tasks.
- Next-round feedback. New policy trajectories provide evidence for the next harness update, closing the continual evolution loop.
SafeEvolve couples agent–environment interaction, safety-harness evolution, and harness-augmented policy optimization.
assets/ prompts, SkillBanks, and compact evolution logs
configs/ LLaMA-Factory and Phase-2 templates
core/ refined Python implementation
env_server/ runtime catalog server and API contract
scripts_skillrl/ refined reward, SFT, and Phase-3 launchers
scripts/ short portable entry points
skillrl_seed_from_student_rollout/
seed_skill_from_utility_safety_teacher/
compact seed assets (raw rollouts omitted)
docs/ method, reproduction and source-mapping notes
tests/ unit and contract tests
The scripts_skillrl/python directory exposes ten core modules as symlinks to
core/skillrl/; the complete implementation remains under core/skillrl/.
For the runtime server and CPU-side utilities:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txtFor development and the included tests, install the additional test dependency:
python -m pip install -r requirements-dev.txtThe minimal environment-server-only install is:
python -m pip install -r env_server/requirements-runtime.txtGPU training is intentionally a separate dependency stack. Install a Slime
checkout together with compatible PyTorch, Ray, SGLang, CUDA, and Megatron-LM
versions, and install LLaMA-Factory separately when using SFT. The compatible
versions depend on the model and container; this release does not pin or
vendor those external repositories. Before training, set SLIME_ROOT,
MEGATRON_ROOT, LLAMA_FACTORY_ROOT, and START_MODEL_DIR.
cd safeevolve_code
python3 -m py_compile core/skillrl/*.py
python3 core/skillrl/runtime_skill_compiler.py \
--source assets/skillbanks/evolved_skill_bank.json \
--output /tmp/runtime_skill_bank.json \
--summary /tmp/runtime_skill_bank.summary.json
PYTHONPATH=core/skillrl \
python3 core/skillrl/runtime_retrieval.py \
--bank assets/skillbanks/runtime_skill_bank.json \
--metadata-json '{"scenario":"attacked","domain":"travel","task_type":"injection","attack_type":"important_instructions"}'The server requires an externally supplied runtime catalog because the internal catalog is generated, multi-gigabyte data. Set the two data paths and run:
export ENV_SERVER_CATALOG_FILE=/path/to/runtime_catalog.json
export ENV_SERVER_PAIRING_FILE=/path/to/semantic_local_pairing.jsonl
bash scripts/env_server/start.sh
bash scripts/env_server/smoke.shThe API and bundle schema are documented in
env_server/docs/runtime_api.md. The server
accepts only prebuilt runtime bundles; it never executes an LLM-generated task
draft directly.
Install the external Slime checkout and set SLIME_ROOT, MEGATRON_ROOT,
START_MODEL_DIR, SLIME_ENV_SERVER_BASE_URL, and SLIME_OUTPUT_ROOT.
Configure the strong proposer with SKILLRL_STRONG_MODEL_BASE_URL,
SKILLRL_STRONG_MODEL_MODEL, and SKILLRL_STRONG_MODEL_API_KEY (the key is
read from the environment and is never stored in this repository).
# Frozen-policy harness evolution
bash scripts/phase1/evolve_prompt.sh
bash scripts/phase1/evolve_skill.sh
# Variant is baseline, prompt, skill, or prompt_skill.
bash scripts/phase2/train_grpo.sh skillbaseline/pure RL uses legacy_utility_first; prompt/skill-augmented runs use
augmented_safety_first. The implementation keeps the sampled-context KL term
to the fixed reference policy (KL coefficient = 0.01) in the underlying GRPO
trainer. See configs/training/phase2_grpo.json.
Use only a compiled runtime bank for fresh rollout collection:
MODEL_FAMILY=qwen35 \
RUNTIME_SKILL_BANK=assets/skillbanks/runtime_skill_bank.json \
START_MODEL_DIR=/path/to/qwen35 \
SLIME_ENV_SERVER_BASE_URL=http://127.0.0.1:18080 \
bash scripts/sft/rollout.sh
bash scripts/sft/build_data.sh /path/to/run/rollout_debug /path/to/sft_data
MODEL_FAMILY=qwen35 \
LLAMA_FACTORY_ROOT=/path/to/LLaMA-Factory \
SFT_DATA_DIR=/path/to/sft_data \
bash scripts/sft/train.shThe model-family-specific templates are in configs/training/. They use full
SFT, one epoch, assistant/tool-call-only loss, and Qwen-native tool formats.
Set MODEL_NAME_OR_PATH/SFT_OUTPUT_DIR or edit the template for local paths.
The pure-Python reward, retrieval, compiler, parser-filter, and SFT-contract tests do not require GPUs or Slime:
python3 -m compileall -q core env_server
PYTHONPATH=core/skillrl pytest -q tests
for f in scripts_skillrl/bash/*.sh scripts/**/*.sh; do bash -n "$f"; doneGPU training and benchmark evaluation require the external environments named
above. Benchmark-specific evaluators should preserve their default system
prompt and append the runtime skill block; see scripts/evaluation/README.md.
Citation information will be added with the paper release.
This repository is released under the MIT License. See the LICENSE
file for the complete license text.
