Shubham Kumar
·
Narendra Ahuja
University of Illinois Urbana-Champaign
Third Conference on Language Modeling (COLM) 2026
We introduce LOCA, a method that gives Local, CAusal explanations of jailbreak success by identifying a minimal set of interpretable, intermediate representation changes that causally induce model refusal on an otherwise successful jailbreak request. LOCA uses token-specific activation patching along sparse-autoencoder concept directions, with a chat-template-aware token matching scheme between refused original prompts and successful jailbreaks. LOCA iteratively re-scores patches after each intervention, so later steps account for interaction effects. On Gemma, Llama, and Qwen chat models, LOCA induces refusal with far fewer patches than adapted baselines from Lee et al. (2025) and Yeo et al. (2025), and a localization analysis shows how causally important tokens shift from user instructions in early layers to post-instruction and punctuation tokens in middle layers.
This repository is the official implementation. It includes patch selection (LOCA and baselines), first-token evaluation, and scripts to reproduce Section 4.1 (main comparison) and Section 4.3 (localization analysis).
loca-jailbreaks/
loca/ # patch selection, evaluation, and analysis
data/<model>/ # precomputed WhatFeatures inputs (responses, splits, refusal direction)
data_preparation/ # scripts to regenerate data/ from scratch
results/ # experiment outputs (gitignored)
Requires Python ≥ 3.10, a CUDA GPU, and HuggingFace access to the target chat models.
git clone https://github.com/skumar-ml/loca-jailbreaks.git
cd loca-jailbreaks
pip install -e .
# or: pip install -r requirements.txt && export PYTHONPATH=.You also need:
- Model weights from HuggingFace (
meta-llama/Llama-3.1-8B-Instruct,google/gemma-2-2b-it, etc.; seeloca/experiment_utils.py→get_model_id). - Sparse autoencoders (SAEs), downloaded automatically on first run via SAELens.
- HarmBench grading (optional; only needed for HarmBench ASR recovery, not for Sections 4.1 or 4.3).
Run GPU jobs with srun or sbatch.
srun --account=<account> --partition=gpuA40x4 \
--gpus=1 --cpus-per-task=8 --mem=62g --time=04:00:00 \
bash -lc '
set -eo pipefail
activate_jailbreak # or: conda activate jailbreak_llms
cd /path/to/loca-jailbreaks
pip install -e . -q
python3 -m loca.run_pipeline --method all ...
'python3 -m loca.run_pipeline \
--method all \
--intervention_layer 19 \
--n_samples 1 \
--seed 42 \
--model_name Llama-3.1-8B-InstructPrecomputed inputs for the paper models live under data/<model_name>/:
| File | Description |
|---|---|
jailbreak_responses.json |
Greedy-decoded responses to WhatFeatures jailbreaks |
jailbreak_responses_graded.json |
HarmBench labels + refusal heuristics |
train.json, val.json, test.json |
70/10/20 stratified split |
split_metadata.json |
Split provenance |
refusal_direction.npy |
Post-instruction refusal direction (Lee et al. baseline) |
refusal_train_metadata.json |
Refusal-direction training metadata |
Supported models: Llama-3.1-8B-Instruct, gemma-2-2b-it, Qwen3-8B, gemma-3-27b-it.
To regenerate these files from the HuggingFace sevdeawesome/jailbreak_success dataset, see [data_preparation/COMMANDS.md](data_preparation/COMMANDS.md).
Results are written to results/<model_name>/ (gitignored).
CLI --method |
Paper reference | Description |
|---|---|---|
loca |
LOCA (this work) | Token-specific, iterative causal greedy KL selection |
yeo |
Yeo et al. (2025) | Prob-margin baseline |
arditi |
Lee et al. (2025) | Refusal-direction baseline |
Setup (from the paper): Filter the test set to original–jailbreak pairs where the original prompt is refused and the jailbreak succeeds (HarmBench). Run on all filtered test pairs (--n_samples -1). Apply up to K = 20 patches at every SAE layer. Report KL-AUC, LD-AUC, Minimal Patches (MP), and Refusal Rate (RR).
| Model | Filtered test pairs |
|---|---|
Llama-3.1-8B-Instruct |
225 |
gemma-2-2b-it |
334 |
Use --n_samples N with a smaller N (and optional --seed) for faster debugging; the smoke test below uses N=1.
Valid SAE layers: 3, 7, 11, 15, 19, 23, 27.
# 1) Patch selection + first-token evaluation (LOCA + baselines)
python3 -m loca.run_pipeline \
--method all \
--intervention_layer 3,7,11,15,19,23,27 \
--n_samples -1 \
--model_name Llama-3.1-8B-Instruct
# 2) Aggregate metrics and plot Fig. 3-style curves
python3 -m loca.analyze_metrics \
--intervention_layer 3,7,11,15,19,23,27 \
--model_name Llama-3.1-8B-InstructPlots are saved under results/Llama-3.1-8B-Instruct/analysis/:
kl_auc_over_layers_*.pdfld_auc_over_layers_*.pdfmp_over_layers_*.pdf(Minimal Patches)refusal_rate_over_layers_*.pdf(Refusal Rate)
Valid SAE layers: 0–24 (25 layers).
python3 -m loca.run_pipeline \
--method all \
--intervention_layer 0:24 \
--n_samples -1 \
--model_name gemma-2-2b-it
python3 -m loca.analyze_metrics \
--intervention_layer 0:24 \
--model_name gemma-2-2b-itEach run writes per-dataset artifacts under:
results/<model>/dataset_<id>/steer_layer_<L>/
loca/ patch_list.npy, run_metadata.json, eval/eval_metadata.json
yeo/ ...
arditi/ ...
Setup (from the paper): Analyze which tokens LOCA selects, by location (instruction vs. post-instruction) and type (word vs. punctuation). For Llama, use intervention layers 3, 7, and 15; for Gemma, layers 5, 10, and 15 (Appendix E). Aggregate over all filtered test pairs (--n_samples -1), same as Section 4.1.
If you already ran the full Section 4.1 pipeline, LOCA patch_list.npy files exist at these layers and you can skip straight to analysis.
# If needed: ensure LOCA patch lists exist at layers 3, 7, 15
python3 -m loca.run_pipeline \
--method loca \
--intervention_layer 3,7,15 \
--n_samples -1 \
--model_name Llama-3.1-8B-Instruct
# Localization plots (aggregated over all datasets with patch lists at each layer)
for L in 3 7 15; do
python3 -m loca.analyze_patch_tokens \
--method loca \
--intervention_layer "$L" \
--model_name Llama-3.1-8B-Instruct
doneFigures are saved to results/Llama-3.1-8B-Instruct/analysis/patch_density_loca_steer<L>_top20.pdf (top row: cumulative % post-instruction tokens; bottom row: cumulative % word tokens).
for L in 5 10 15; do
python3 -m loca.run_pipeline \
--method loca \
--intervention_layer "$L" \
--n_samples -1 \
--model_name gemma-2-2b-it
python3 -m loca.analyze_patch_tokens \
--method loca \
--intervention_layer "$L" \
--model_name gemma-2-2b-it
done# HarmBench ASR recovery (requires vLLM + HarmBench classifier)
python3 -m loca.generate_swapped_responses --method all \
--intervention_layer 3,7,11,15,19,23,27 \
--n_samples -1 --model_name Llama-3.1-8B-Instruct
python3 -m loca.grade_harmbench --method all \
--intervention_layer 3,7,11,15,19,23,27 \
--model_name Llama-3.1-8B-Instruct
python3 -m loca.analyze_harmbench \
--intervention_layer 3,7,11,15,19,23,27 \
--model_name Llama-3.1-8B-InstructSection 4.2 ablations (Base-LOCA, Token-LOCA) and the Section 4.4 case study are described in the paper but not included in this release.
- Metric plots use first-token proxies (KL, logit difference) as in the paper; full-response refusal is not regenerated at each patch step.
- Neuronpedia can be used to interpret SAE feature IDs from
patch_list.npy(see paper Appendix G for the case study).
@inproceedings{kumar2026loca,
title={Minimal, Local, Causal Explanations for Jailbreak Success in Large Language Models},
author={Kumar, Shubham and Ahuja, Narendra},
booktitle={Third Conference on Language Modeling (COLM)},
year={2026},
url={https://arxiv.org/abs/2605.00123}
}