Skip to content

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Minimal, Local, Causal Explanations
for Jailbreak Success in Large Language Models

Shubham Kumar · Narendra Ahuja
University of Illinois Urbana-Champaign
Third Conference on Language Modeling (COLM) 2026

Paper PDF Project Page

We introduce LOCA, a method that gives Local, CAusal explanations of jailbreak success by identifying a minimal set of interpretable, intermediate representation changes that causally induce model refusal on an otherwise successful jailbreak request. LOCA uses token-specific activation patching along sparse-autoencoder concept directions, with a chat-template-aware token matching scheme between refused original prompts and successful jailbreaks. LOCA iteratively re-scores patches after each intervention, so later steps account for interaction effects. On Gemma, Llama, and Qwen chat models, LOCA induces refusal with far fewer patches than adapted baselines from Lee et al. (2025) and Yeo et al. (2025), and a localization analysis shows how causally important tokens shift from user instructions in early layers to post-instruction and punctuation tokens in middle layers.

This repository is the official implementation. It includes patch selection (LOCA and baselines), first-token evaluation, and scripts to reproduce Section 4.1 (main comparison) and Section 4.3 (localization analysis).

Repository layout

loca-jailbreaks/
  loca/                 # patch selection, evaluation, and analysis
  data/<model>/         # precomputed WhatFeatures inputs (responses, splits, refusal direction)
  data_preparation/     # scripts to regenerate data/ from scratch
  results/              # experiment outputs (gitignored)

Setup

Environment

Requires Python ≥ 3.10, a CUDA GPU, and HuggingFace access to the target chat models.

git clone https://github.com/skumar-ml/loca-jailbreaks.git
cd loca-jailbreaks
pip install -e .
# or: pip install -r requirements.txt && export PYTHONPATH=.

You also need:

  • Model weights from HuggingFace (meta-llama/Llama-3.1-8B-Instruct, google/gemma-2-2b-it, etc.; see loca/experiment_utils.py → get_model_id).
  • Sparse autoencoders (SAEs), downloaded automatically on first run via SAELens.
  • HarmBench grading (optional; only needed for HarmBench ASR recovery, not for Sections 4.1 or 4.3).

Cluster (SLURM)

Run GPU jobs with srun or sbatch.

srun --account=<account> --partition=gpuA40x4 \
  --gpus=1 --cpus-per-task=8 --mem=62g --time=04:00:00 \
  bash -lc '
set -eo pipefail
activate_jailbreak   # or: conda activate jailbreak_llms
cd /path/to/loca-jailbreaks
pip install -e . -q
python3 -m loca.run_pipeline --method all ...
'

Quick smoke test

python3 -m loca.run_pipeline \
  --method all \
  --intervention_layer 19 \
  --n_samples 1 \
  --seed 42 \
  --model_name Llama-3.1-8B-Instruct

Data

Precomputed inputs for the paper models live under data/<model_name>/:

File Description
jailbreak_responses.json Greedy-decoded responses to WhatFeatures jailbreaks
jailbreak_responses_graded.json HarmBench labels + refusal heuristics
train.json, val.json, test.json 70/10/20 stratified split
split_metadata.json Split provenance
refusal_direction.npy Post-instruction refusal direction (Lee et al. baseline)
refusal_train_metadata.json Refusal-direction training metadata

Supported models: Llama-3.1-8B-Instruct, gemma-2-2b-it, Qwen3-8B, gemma-3-27b-it.

To regenerate these files from the HuggingFace sevdeawesome/jailbreak_success dataset, see [data_preparation/COMMANDS.md](data_preparation/COMMANDS.md).

Results are written to results/<model_name>/ (gitignored).

Methods

CLI --method Paper reference Description
loca LOCA (this work) Token-specific, iterative causal greedy KL selection
yeo Yeo et al. (2025) Prob-margin baseline
arditi Lee et al. (2025) Refusal-direction baseline

Reproducing Section 4.1 — main comparison (Fig. 3)

Setup (from the paper): Filter the test set to original–jailbreak pairs where the original prompt is refused and the jailbreak succeeds (HarmBench). Run on all filtered test pairs (--n_samples -1). Apply up to K = 20 patches at every SAE layer. Report KL-AUC, LD-AUC, Minimal Patches (MP), and Refusal Rate (RR).

Model Filtered test pairs
Llama-3.1-8B-Instruct 225
gemma-2-2b-it 334

Use --n_samples N with a smaller N (and optional --seed) for faster debugging; the smoke test below uses N=1.

Llama-3.1-8B-Instruct

Valid SAE layers: 3, 7, 11, 15, 19, 23, 27.

# 1) Patch selection + first-token evaluation (LOCA + baselines)
python3 -m loca.run_pipeline \
  --method all \
  --intervention_layer 3,7,11,15,19,23,27 \
  --n_samples -1 \
  --model_name Llama-3.1-8B-Instruct

# 2) Aggregate metrics and plot Fig. 3-style curves
python3 -m loca.analyze_metrics \
  --intervention_layer 3,7,11,15,19,23,27 \
  --model_name Llama-3.1-8B-Instruct

Plots are saved under results/Llama-3.1-8B-Instruct/analysis/:

  • kl_auc_over_layers_*.pdf
  • ld_auc_over_layers_*.pdf
  • mp_over_layers_*.pdf (Minimal Patches)
  • refusal_rate_over_layers_*.pdf (Refusal Rate)

gemma-2-2b-it

Valid SAE layers: 0–24 (25 layers).

python3 -m loca.run_pipeline \
  --method all \
  --intervention_layer 0:24 \
  --n_samples -1 \
  --model_name gemma-2-2b-it

python3 -m loca.analyze_metrics \
  --intervention_layer 0:24 \
  --model_name gemma-2-2b-it

Each run writes per-dataset artifacts under:

results/<model>/dataset_<id>/steer_layer_<L>/
  loca/   patch_list.npy, run_metadata.json, eval/eval_metadata.json
  yeo/    ...
  arditi/ ...

Reproducing Section 4.3 — localization analysis (Fig. 5)

Setup (from the paper): Analyze which tokens LOCA selects, by location (instruction vs. post-instruction) and type (word vs. punctuation). For Llama, use intervention layers 3, 7, and 15; for Gemma, layers 5, 10, and 15 (Appendix E). Aggregate over all filtered test pairs (--n_samples -1), same as Section 4.1.

If you already ran the full Section 4.1 pipeline, LOCA patch_list.npy files exist at these layers and you can skip straight to analysis.

Llama-3.1-8B-Instruct

# If needed: ensure LOCA patch lists exist at layers 3, 7, 15
python3 -m loca.run_pipeline \
  --method loca \
  --intervention_layer 3,7,15 \
  --n_samples -1 \
  --model_name Llama-3.1-8B-Instruct

# Localization plots (aggregated over all datasets with patch lists at each layer)
for L in 3 7 15; do
  python3 -m loca.analyze_patch_tokens \
    --method loca \
    --intervention_layer "$L" \
    --model_name Llama-3.1-8B-Instruct
done

Figures are saved to results/Llama-3.1-8B-Instruct/analysis/patch_density_loca_steer<L>_top20.pdf (top row: cumulative % post-instruction tokens; bottom row: cumulative % word tokens).

gemma-2-2b-it (Appendix E)

for L in 5 10 15; do
  python3 -m loca.run_pipeline \
    --method loca \
    --intervention_layer "$L" \
    --n_samples -1 \
    --model_name gemma-2-2b-it

  python3 -m loca.analyze_patch_tokens \
    --method loca \
    --intervention_layer "$L" \
    --model_name gemma-2-2b-it
done

Other experiments (not required for Sections 4.1 / 4.3)

# HarmBench ASR recovery (requires vLLM + HarmBench classifier)
python3 -m loca.generate_swapped_responses --method all \
  --intervention_layer 3,7,11,15,19,23,27 \
  --n_samples -1 --model_name Llama-3.1-8B-Instruct
python3 -m loca.grade_harmbench --method all \
  --intervention_layer 3,7,11,15,19,23,27 \
  --model_name Llama-3.1-8B-Instruct
python3 -m loca.analyze_harmbench \
  --intervention_layer 3,7,11,15,19,23,27 \
  --model_name Llama-3.1-8B-Instruct

Section 4.2 ablations (Base-LOCA, Token-LOCA) and the Section 4.4 case study are described in the paper but not included in this release.

Notes

  • Metric plots use first-token proxies (KL, logit difference) as in the paper; full-response refusal is not regenerated at each patch step.
  • Neuronpedia can be used to interpret SAE feature IDs from patch_list.npy (see paper Appendix G for the case study).

Citation

@inproceedings{kumar2026loca,
  title={Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models},
  author={Kumar, Shubham and Ahuja, Narendra},
  booktitle={Third Conference on Language Modeling (COLM)},
  year={2026},
  url={https://arxiv.org/abs/2605.00123}
}

About

Official code for LOCA (published in COLM 2026).

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages