Skip to content

Latest commit

 

History

8 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

CCA: Context Compilation Architecture

Official code repository for the EMNLP 2026 Findings paper:

Compile, Don't Memorize: A Context Compilation Architecture (CCA) for In-Context Learning Jinhu Qi, Minda Hu, Wentao Zhang, Weiqiang Jin, Yanyu Chen, Junli Wang, Irwin King The Chinese University of Hong Kong · Macao Polytechnic University · Xi'an Jiaotong University · University of Science and Technology of China

📄 Paper (arXiv) · 📊 OpenReview

CCA architecture — Vanilla ICL "Fixed-Plate" (left) vs CCA "Movable Type" (right)

Left (red, "Fixed-Plate"): vanilla ICL pushes the long context through a single LLM forward pass — one overlooked rule fails the whole response. Right (green, "Movable Type"): CCA compiles the context into a typed IR (rules, exact terms, output schema), generates executable Python verifiers from the IR, drafts an answer with the IR as an explicit checklist, and regenerates with violation feedback if verifiers flag ≥2 violations.


What is CCA?

Standard in-context learning (ICL) pushes long, rule-dense prose through an LLM in a single forward pass, and one overlooked rule fails the whole response. On CL-bench, even strong open-weights models pass only 12–16% of tasks under this "read-and-reason" paradigm.

CCA compiles context instead of memorising it. The paper's central novelty is a typed intermediate representation (IR) with fixed slots:

rules.{must_do, must_not, conditional}
output_spec
available_tools
data_profile

Any prose context is compiled once into this typed IR. Executable Python verifiers and a violation-gated correction loop follow as downstream consequences.

Result on CL-bench (1,899 tasks, 4 open base models): Kimi K2.5 lifts from 15.4% → 21.4% (+6.0pp, p<0.0001 vs Self-Refine), with gains concentrated on rule-dense sub-categories.

The runner is backend-agnostic: the same code runs on OpenAI (GPT-5.x), Anthropic (Claude series), or AWS Bedrock. Everything is set from the command line — you never edit a file to change the model or key.


Repository layout

.
├── code/
│   ├── cca_core.py          # CCA pipeline library (Compiler → CodeGen → Reasoner-1 → Reasoner-2)
│   ├── run_cca.py           # CLI inference runner (this is what you run)
│   ├── eval.py              # LLM-as-judge grader (binary pass/fail per CL-bench rubric)
│   ├── config.py            # put your API keys here once (keys + paper-matching defaults)
│   └── requirements.txt
├── data/
│   └── CL-bench.jsonl       # the benchmark — 1,899 tasks, bundled (86 MB)
├── EXPERIMENT_DATA.md       # full paper results (4 methods × 5 models × 1,899 tasks)
├── LICENSE                  # Apache 2.0
└── README.md                # this file

The benchmark data is bundled in data/CL-bench.jsonl (1,899 tasks); you don't need to download anything separately. (Same file also released at github.com/Tencent-Hunyuan/CL-bench under data/.)


1. Install

git clone https://github.com/TonyQJH/cca-emnlp2026.git
cd cca-emnlp2026
python3 -m venv .venv && source .venv/bin/activate
pip install -r code/requirements.txt

boto3 is only needed for the Bedrock backend; if you only run OpenAI/Anthropic you can skip it.

2. Set your API keys (once)

Open code/config.py and fill in:

OPENAI_API_KEY    = "sk-..."        # for --provider openai  and the eval.py judge
ANTHROPIC_API_KEY = "sk-ant-..."    # for --provider anthropic

run_cca.py and eval.py read these automatically, so you never have to pass --api-key. (A --api-key on the command line, or the OPENAI_API_KEY / ANTHROPIC_API_KEY env vars, still override config.py.) config.py also pins the paper-matching defaults (temperature 0.0, output cap 8192, judge gpt-5.1) — leave those as-is for a faithful comparison.

⚠️ config.py holds secrets. Do not commit real keys — the file is in the repo so users know where to put keys, but treat the copy on disk as sensitive.


3. Run CCA inference

One model at a time. Output is a JSONL file (one graded-ready record per task), with automatic resume — re-run the exact same command to continue an interrupted run. (API key is read from config.py, so it is not on the command line.)

OpenAI (GPT-5.x):

cd code
python run_cca.py \
  --provider openai --model gpt-5.1 \
  --input ../data/CL-bench.jsonl \
  --output ../outputs/cca_gpt51.jsonl \
  --workers 8

Swap --model gpt-5.1 for gpt-5.2, gpt-5.4, gpt-5.5, etc. — use the exact model id your account exposes.

Anthropic (Claude):

cd code
python run_cca.py \
  --provider anthropic --model claude-opus-4-6 \
  --input ../data/CL-bench.jsonl \
  --output ../outputs/cca_claude.jsonl \
  --workers 8

Use whatever Claude model id your account exposes for --model.

Smoke-test on a slice first — always confirm your key/model work (and eyeball the cost) before committing to all 1,899 tasks. Make a small input and point the runner at it:

head -n 40 ../data/CL-bench.jsonl > ../data/smoke.jsonl
python run_cca.py --provider openai --model gpt-5.1 \
  --input ../data/smoke.jsonl --output ../outputs/smoke.jsonl

The full run is just the same command with --input ../data/CL-bench.jsonl; it resumes, so you can stop and restart it freely.

Useful flags

Flag Meaning Default
--provider openai / anthropic / bedrock (required)
--model exact API model id (required)
--input CL-bench JSONL path (required)
--output output JSONL path (resumes if exists) (required)
--workers parallelism: how many context groups run concurrently 8
--api-key overrides config.py / env var config.py
--temperature sampling temperature (0.0 = reproducible, matches paper) config.py (0.0)
--max-output-tokens answer / reasoner output cap (the graded output) config.py (8192)
--max-compiler-tokens Compiler IR output cap (kept larger so verbose models don't truncate the IR) config.py (16384)
--max-context-chars compiler input cap (chars) provider default
--max-reasoner-context reasoner context cap (chars) provider default
--base-url OpenAI-compatible proxy/Azure endpoint none
--region AWS region (bedrock only, ignored otherwise) us-east-1

--workers is throughput, not a task cap: --workers 8 runs 8 context groups at once. The runner always processes the whole --input file — to run on fewer tasks, feed it a smaller input (see the smoke-test slice above). Tune --workers to your account's rate limits.


4. Evaluate (grade pass/fail)

CL-bench grades with a strong LLM-as-judge under all-or-nothing rubric scoring (a task passes only if every rubric criterion is met). The judge is GPT-5.1 for every model under test (set in config.py), so pass rates stay comparable no matter which model produced the answers — do not change it.

cd code
python eval.py \
  --input ../outputs/cca_gpt51.jsonl \
  --output ../outputs/cca_gpt51_graded.jsonl \
  --workers 8

eval.py reads the judge key + model from config.py, also resumes, and prints the headline number at the end:

📈 Solving Rate: 0.2140 (406/1899)

📂 Scores by context_category:
   Domain Knowledge Reasoning: total=663, score_1=..., rate=...
   Rule System Application:    total=566, ...
   Procedural Task Execution:  total=471, ...
   Empirical Discovery & Simulation: total=199, ...

(Judge via a non-OpenAI endpoint: add --base-url https://your-openai-compatible-endpoint/v1.)


5. Where to see results

  • Overall pass rate + per-category breakdown — printed by eval.py at the end (re-run on a finished *_graded.jsonl to reprint without re-grading).
  • Per-task detail — each line of *_graded.jsonl has score (0/1), grading_rationale, requirement_status, plus the original model_output, rubrics, and metadata.
  • Per-task CCA cost/behaviour — each inference line carries a cca_meta block: tokens_compiler_*, tokens_codegen_*, tokens_reasoner_1/2, tokens_total_per_task, n_violations, corrected, strategy.

Method (CCA)

Four stages; the first two run once per context (amortized over every question that shares that context), the last two run once per task:

  1. Compiler — context → typed JSON IR (rules.{must_do, must_not, conditional}, output_spec, available_tools, data_profile).
  2. CodeGen — IR → up to three Python modules: rule_checker, format_validator, and (for data-rich contexts) data_analyzer.
  3. Reasoner-1 — drafts an answer using the IR as an explicit checklist (plus the cached data_analyzer summary when applicable).
  4. Reasoner-2 (correction) — if the draft-time verifiers flag ≥ 2 concrete violations, regenerate with the violation list as feedback; otherwise keep the draft.

Paper results (Pass Rate, full 1,899 tasks)

These are the numbers from the paper's main table (AWS Bedrock open-weights models; see EXPERIMENT_DATA.md for per-domain / per-sub-category / per-stage-token detail).

Method Kimi K2.5 GLM-5 DeepSeek-V3.2 Qwen3-Next-80B Avg
Vanilla 15.4% 16.1% 15.0% 11.9% 14.6%
ReadAgent-P 14.7% 13.0% 13.1% 11.2% 13.0%
Ctx2Skill 15.5% 15.2% 15.0% 11.9% 14.4%
CCA (ours) 21.4% 21.2% 17.7% 12.4% 18.2%

Δ vs Vanilla: +6.0 (Kimi) / +5.1 (GLM-5) / +2.7 (DeepSeek) / +0.5 (Qwen3) pp — all p<0.01 by paired McNemar test.

The runner above targets OpenAI/Anthropic so you can reproduce the method on GPT-5.x and Claude; the paper's headline table used open-weights base models via Bedrock.


Notes

  • Inference temperature defaults to 0.0 (greedy) for reproducibility — this matches the paper's main table. Pass --temperature 1.0 to study sampling robustness (paper Appendix I).
  • The Compiler emits a long JSON IR. Its output cap (--max-compiler-tokens, default 16384) is kept above the answer cap (--max-output-tokens, default 8192) so verbose models don't truncate the IR mid-JSON — a truncated IR would silently collapse CCA to plain prompting. The graded answer still uses the 8192 cap.
  • The Anthropic backend uses streaming, which is required once --max-compiler-tokens is large (the SDK refuses long non-streaming calls).
  • The Compiler can rarely emit an unparseable IR; the runner falls back to direct prompting for those tasks so no task is dropped.
  • Graded *.jsonl outputs are git-ignored (regenerate via the steps above).

Citation

If you use this work, please cite:

@inproceedings{qi2026cca,
  title     = {Compile, Don't Memorize: A Context Compilation Architecture (CCA) for In-Context Learning},
  author    = {Qi, Jinhu and Hu, Minda and Zhang, Wentao and Jin, Weiqiang and Chen, Yanyu and Wang, Junli and King, Irwin},
  booktitle = {Findings of the Association for Computational Linguistics: EMNLP 2026},
  year      = {2026},
  url       = {https://openreview.net/forum?id=fkHzqnhdIY}
}

License

Apache License 2.0 — see LICENSE.


Contact

  • Jinhu Qi (PhD student, first author) — jhqi25@cse.cuhk.edu.hk
  • Prof. Irwin King (supervisor) — king@cse.cuhk.edu.hk
  • Department of Computer Science and Engineering, The Chinese University of Hong Kong

Issues and questions are welcome via GitHub Issues.


Acknowledgements

Supported in part by RGC grants and CUHK internal funding — see the paper for full acknowledgement.

About

Compile, Don't Memorize: A Context Compilation Architecture (CCA) for In-Context Learning — EMNLP 2026 Findings

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages