Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Validity-Aware Jailbreak Evaluation for Large Language Models

SEAV (Sequential Epistemic and Action-Level Validation) reproducibility code.

Authors: Qilong Wu, Sahil Wadhwa, Pranab Mohanty, Giri Iyengar, Varun Chandrasekaran

Venue: EMNLP 2026 (Main Conference)

Paper: ACL Anthology link TBA

Setup

pip install -r requirements.txt

Environment Variables

export GEMINI_API_KEY="..."        # For Gemini models (default backend)
export OPENROUTER_API_KEY="..."    # For OpenRouter models (alternative)
export OPENAI_API_KEY="..."        # For OpenAI-compatible baselines
export TAVILY_API_KEY="..."        # For Tavily web search

Quick Demo

python demo/run_demo.py

Running SEAV

python experiments/run_methods.py \
    --method seal \
    --dataset jailbreakqr \
    --model gemini-3-flash-preview \
    --jailbreakqr-path <path-to-dataset.json>

See experiments/run_methods.py --help for all options, methods, and datasets.

Datasets

  • OrdSense: Included as datasets/ordsense/order_variants_with_responses.jsonl and published as a public, ungated Hugging Face dataset (137 records, with original, alternative-valid, and wrong-order variants). See the Dataset Card, datasets/ordsense/README.md, and paper Appendix B.2.

    Hugging Face Dataset

    from datasets import load_dataset
    
    dataset = load_dataset("QilongWu/OrdSense", split="test")

    The software's MIT license does not cover upstream-derived OrdSense text; review the Hugging Face Dataset Card terms before use.

  • SD, JQR-Binary: Not included due to licensing restrictions on source data. See paper for construction details.

  • JBB: Publicly available via pip install datasets and load_dataset("JailbreakBench/JBB-Behaviors").

  • GPTFuzz, WildGuardMix, UltraSafety: Publicly available; see paper references for download links.

Package Structure

seav/           SEAV pipeline (4 nodes: extraction, verification, ordering, judgment)
baselines/      Baseline evaluation methods
jades/          JADES baseline
experiments/    Experiment runner and metrics computation
demo/           Quick demo script
datasets/       OrdSense ordering-sensitivity evaluation dataset

Citation

@inproceedings{wu2026seav,
  title = {Validity-Aware Jailbreak Evaluation for Large Language Models},
  author = {Wu, Qilong and Wadhwa, Sahil and Mohanty, Pranab and Iyengar, Giri and Chandrasekaran, Varun},
  booktitle = {Proceedings of EMNLP 2026},
  year = {2026}
}

License

The software is released under the MIT License. Dataset files retain their source terms and are not covered by the software license; see the dataset card in each dataset directory.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages