Skip to content

Latest commit

 

History

History
185 lines (140 loc) · 6.81 KB

File metadata and controls

185 lines (140 loc) · 6.81 KB

AgentDoG icon

SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment

This directory is the reproducible code release of SafeEvolve: Harness-Policy Co-Evolution from Agent Experience for Safety Alignment. SafeEvolve turns on-policy agent trajectories into two coupled artifacts���a global safety prompt and a hierarchical SkillBank—and then internalizes their use with harness-use SFT followed by harness-augmented GRPO.

Overview

SafeEvolve runs a closed-loop harness–policy co-evolution process:

  1. Agent–environment interaction. A policy model operates through a safety harness containing safety rules and retrieved skills. It issues tool calls, receives environment observations, and produces a complete trajectory.
  2. Safety harness evolution. Rollout evidence is grouped by risk and task metadata, distilled into prompt or SkillBank candidates, and accepted only after bounded edits pass safety and utility gates.
  3. Harness-augmented policy optimization. Curated multi-step rollouts are rejection-sampled for SFT, then GRPO trains the policy to use the evolved harness across benign, query, and injection tasks.
  4. Next-round feedback. New policy trajectories provide evidence for the next harness update, closing the continual evolution loop.

SafeEvolve framework overview

SafeEvolve couples agent–environment interaction, safety-harness evolution, and harness-augmented policy optimization.

Repository map

assets/                         prompts, SkillBanks, and compact evolution logs
configs/                        LLaMA-Factory and Phase-2 templates
core/                           refined Python implementation
env_server/                     runtime catalog server and API contract
scripts_skillrl/                refined reward, SFT, and Phase-3 launchers
scripts/                        short portable entry points
skillrl_seed_from_student_rollout/
seed_skill_from_utility_safety_teacher/
                                compact seed assets (raw rollouts omitted)
docs/                           method, reproduction and source-mapping notes
tests/                          unit and contract tests

The scripts_skillrl/python directory exposes ten core modules as symlinks to core/skillrl/; the complete implementation remains under core/skillrl/.

Installation

For the runtime server and CPU-side utilities:

python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt

For development and the included tests, install the additional test dependency:

python -m pip install -r requirements-dev.txt

The minimal environment-server-only install is:

python -m pip install -r env_server/requirements-runtime.txt

GPU training is intentionally a separate dependency stack. Install a Slime checkout together with compatible PyTorch, Ray, SGLang, CUDA, and Megatron-LM versions, and install LLaMA-Factory separately when using SFT. The compatible versions depend on the model and container; this release does not pin or vendor those external repositories. Before training, set SLIME_ROOT, MEGATRON_ROOT, LLAMA_FACTORY_ROOT, and START_MODEL_DIR.

Quick start: inspect and compile a SkillBank

cd safeevolve_code
python3 -m py_compile core/skillrl/*.py

python3 core/skillrl/runtime_skill_compiler.py \
  --source assets/skillbanks/evolved_skill_bank.json \
  --output /tmp/runtime_skill_bank.json \
  --summary /tmp/runtime_skill_bank.summary.json

PYTHONPATH=core/skillrl \
python3 core/skillrl/runtime_retrieval.py \
  --bank assets/skillbanks/runtime_skill_bank.json \
  --metadata-json '{"scenario":"attacked","domain":"travel","task_type":"injection","attack_type":"important_instructions"}'

Environment server

The server requires an externally supplied runtime catalog because the internal catalog is generated, multi-gigabyte data. Set the two data paths and run:

export ENV_SERVER_CATALOG_FILE=/path/to/runtime_catalog.json
export ENV_SERVER_PAIRING_FILE=/path/to/semantic_local_pairing.jsonl
bash scripts/env_server/start.sh
bash scripts/env_server/smoke.sh

The API and bundle schema are documented in env_server/docs/runtime_api.md. The server accepts only prebuilt runtime bundles; it never executes an LLM-generated task draft directly.

Phase 1 and Phase 2 training

Install the external Slime checkout and set SLIME_ROOT, MEGATRON_ROOT, START_MODEL_DIR, SLIME_ENV_SERVER_BASE_URL, and SLIME_OUTPUT_ROOT. Configure the strong proposer with SKILLRL_STRONG_MODEL_BASE_URL, SKILLRL_STRONG_MODEL_MODEL, and SKILLRL_STRONG_MODEL_API_KEY (the key is read from the environment and is never stored in this repository).

# Frozen-policy harness evolution
bash scripts/phase1/evolve_prompt.sh
bash scripts/phase1/evolve_skill.sh

# Variant is baseline, prompt, skill, or prompt_skill.
bash scripts/phase2/train_grpo.sh skill

baseline/pure RL uses legacy_utility_first; prompt/skill-augmented runs use augmented_safety_first. The implementation keeps the sampled-context KL term to the fixed reference policy (KL coefficient = 0.01) in the underlying GRPO trainer. See configs/training/phase2_grpo.json.

SFT cold start

Use only a compiled runtime bank for fresh rollout collection:

MODEL_FAMILY=qwen35 \
RUNTIME_SKILL_BANK=assets/skillbanks/runtime_skill_bank.json \
START_MODEL_DIR=/path/to/qwen35 \
SLIME_ENV_SERVER_BASE_URL=http://127.0.0.1:18080 \
bash scripts/sft/rollout.sh

bash scripts/sft/build_data.sh /path/to/run/rollout_debug /path/to/sft_data

MODEL_FAMILY=qwen35 \
LLAMA_FACTORY_ROOT=/path/to/LLaMA-Factory \
SFT_DATA_DIR=/path/to/sft_data \
bash scripts/sft/train.sh

The model-family-specific templates are in configs/training/. They use full SFT, one epoch, assistant/tool-call-only loss, and Qwen-native tool formats. Set MODEL_NAME_OR_PATH/SFT_OUTPUT_DIR or edit the template for local paths.

Tests and checks

The pure-Python reward, retrieval, compiler, parser-filter, and SFT-contract tests do not require GPUs or Slime:

python3 -m compileall -q core env_server
PYTHONPATH=core/skillrl pytest -q tests
for f in scripts_skillrl/bash/*.sh scripts/**/*.sh; do bash -n "$f"; done

GPU training and benchmark evaluation require the external environments named above. Benchmark-specific evaluators should preserve their default system prompt and append the runtime skill block; see scripts/evaluation/README.md.

Citation

Citation information will be added with the paper release.

License

This repository is released under the MIT License. See the LICENSE file for the complete license text.