Code repository for "Slow to See, Slow to Suppress: Understanding the Effects of Modality in Context-Memory Conflicts" -- published in EMNLP Findings 2026
We investigate how vision-language models (VLMs) handle context-memory conflicts; that is, situations in which the model is given information in context that differs from what was stored parametrically during training. We document asymmetric biases: models tend to prefer in-context information when modalities align, but prefer parametric information otherwise. We relate this asymmetry to the late resolution of visual entities, and show that the longer processing time associated with resolving visual entities prevents the suppression of the model's usual factual recall mechanism, thus resulting in more parametric answers. Chain-of-thought reasoning does not appear to resolve the gap, but increasing the amount of visual information in the context does show an effect. These results illustrate the complexity of ensuring consistent behavior as models become increasingly multimodal and retrieval-augmented.
| Step | Script | What it does |
|---|---|---|
| 0 | src/download_data.py |
Pulls the dataset from Hugging Face and lays it out on disk. |
| 1 | src/ctxt_mem_experiment.py |
The core behavioral experiment: entity recognition, context-memory conflict, CoT and visual-context variants. |
| 2 | src/judge.py |
LLM-as-a-judge scoring of a script's output CSV, plus a summary figure. |
| 3 | src/mlp_ablation.py + src/plot_mlp_ablation.py |
MLP ablation sweep + Figure 3 (log-probability margin vs. ablated layer). |
| 4 | src/attention_block.py + src/plot_attention_block.py |
Attention-masking sweep + Figure 4 (margin vs. blocked layer, with the causal-mediation "clean MLP" control). |
| 5 | src/backpatch.py + src/backpatch_scorer.py + src/plot_backpatch.py |
Back-patching sweep + Figure 5 (% parametric vs. source layer). |
All example commands below use Qwen/Qwen2.5-VL-7B-Instruct; swap in google/gemma-3-12b-it or a Ministral 3 checkpoint as needed (see Environments).
Tested models include: Qwen/Qwen2.5-VL-7B-Instruct, Qwen/Qwen2.5-VL-32B-Instruct, Qwen/Qwen2.5-VL-72B-Instruct, google/gemma-3-12b-it, google/gemma-3-27b-it, mistralai/Ministral-3-8B-Instruct-2512,mistralai/Ministral-3-14B-Instruct-2512
Model support is split across two environments because Ministral 3 needs a very recent (pre-release) transformers build that conflicts with the pinned versions used for Qwen2.5-VL and Gemma-3:
# Qwen2.5-VL and Gemma-3
conda env create -f environment_qwen_gemma.yml
conda activate cmc-qwen-gemma
# Ministral 3
conda env create -f environment_mistral.yml
conda activate cmc-mistralIf you dont have/want to use conda, use the matching pip requirements file instead:
pip install -r requirements-qwen-gemma.txt # or requirements-mistral.txtThe judge model (meta-llama/Meta-Llama-3-8B-Instruct, used by judge.py and backpatch_scorer.py) works fine under the cmc-qwen-gemma environment.
Every script imports its paths from src/paths.py. By default everything lives under the repo root (data/, results/, results/interpretability/, figures/), created automatically on first import. Override any of them with environment variables if you'd rather point at cluster scratch space, e.g.:
export CMC_DATA_ROOT=/scratch/$USER/cmc_data
export CMC_RESULTS_DIR=/scratch/$USER/cmc_results
export CMC_HF_CACHE_DIR=/scratch/$USER/hf_cache # where model weights get cachedcd src
python download_data.pyThis pulls all three domains (people, artwork, building) from aparaselli/slow-to-see-slow-to-suppress on the Hugging Face Hub -- images are embedded directly in the dataset, so there's no separate scraping/downloading step -- and writes them back out to data/MLLMKC_data/ as {people,artwork,building}_knowledge.json plus an image*/ folder per domain, matching exactly what the experiment scripts expect. Use --domains people to grab just one, or --data_root to write elsewhere.
Each entry has a ground-truth knowledge string, an mis_knowledge dict of plausible-but-false alternatives for each attribute (birth year / nationality / career for people, artist / year / location for artwork, creator / year / location for buildings), and both open-ended and multiple-choice query variants. See the paper's Appendix A for exactly how the misinformation and MCQ distractors were constructed.
ctxt_mem_experiment.py runs a single model over a single domain, feeding it a (possibly conflicting) text context alongside the entity presented either as an image or as a name, and records the model's response, its top decoded token, and the sequence log-probability of the generated continuation.
The conflict experiment only evaluates entities the model can actually recognize and already knows something about (otherwise "context vs. memory conflict" isn't well-defined). Run this once per model/domain before anything else:
python ctxt_mem_experiment.py \
--model_name Qwen/Qwen2.5-VL-7B-Instruct \
--bench_type people \
--run_entity_id True \
--entity_modality visionThis writes results/Qwen7B_identify_entity_id_people.csv, which the main experiment reads to filter to known entities (--filter_known_entities True, the default). Repeat for --bench_type artwork and --bench_type building.
python ctxt_mem_experiment.py \
--model_name Qwen/Qwen2.5-VL-7B-Instruct \
--bench_type people \
--entity_modality vision # or: textKey flags:
| Flag | Meaning |
|---|---|
--entity_modality {vision,text} |
Whether the queried entity is shown as an image or referred to by name. |
--bench_type {people,artwork,building} |
Which domain to run. |
--mock_RAG |
True (default) injects the misleading context; False runs with no context (pure parametric recall). |
--ctxt_type {text,vision,vision+text} |
How the retrieved context is presented -- see visual-context prompting below. |
--CoT |
True enables chain-of-thought prompting (JSON entity/reasoning/answer output). |
--control |
Use a fixed, unrelated control context (e.g. "Peter Parker...") instead of a plausible misinformation context, as a sanity baseline. |
--faithful_context |
Use the entity's true knowledge as context, instead of misinformation -- another sanity baseline. |
Output is written to results/{ModelTag}_FRQ_{RAG|no_RAG}_{VISION|TEXT}[...]_Experiment_{people|artwork|building}_Results.csv (auto-named by get_default_out_path(); pass --out_path to override).
python ctxt_mem_experiment.py \
--model_name Qwen/Qwen2.5-VL-7B-Instruct \
--bench_type people --entity_modality vision --CoT TrueSection 6.2 of the paper: instead of supplying the retrieved context as text, supply it as an image of the (possibly different) context entity, e.g. an image of Taylor Swift plus "The person pictured is a novelist" instead of "Taylor Swift is a novelist."
# Context image = the same image used for the query entity
python ctxt_mem_experiment.py \
--model_name Qwen/Qwen2.5-VL-7B-Instruct \
--bench_type people --entity_modality text \
--ctxt_type vision --vis_ctxt_diff False
# Context image = a different image of the same entity
python ctxt_mem_experiment.py \
--model_name Qwen/Qwen2.5-VL-7B-Instruct \
--bench_type people --entity_modality text \
--ctxt_type vision --vis_ctxt_diff True--ctxt_type vision+text keeps the original text caption alongside the context image, rather than replacing the entity name with "The person pictured."
judge.py classifies each response as favoring the parametric answer, the (misleading) contextual answer, or neither, using meta-llama/Meta-Llama-3-8B-Instruct as the judge. The judge prompts are unchanged from the original research code (src/classify_response.py).
python judge.py \
--csv "results/Qwen7B_FRQ_RAG_VISION_Experiment_People_Results.csv" \
--mode ctxt_mem \
--plot --group_by Category--csv accepts globs and multiple paths. --mode fact_recall instead scores CORRECT/INCORRECT against a single ground-truth answer (used to pre-filter for baseline factual recall before running a conflict experiment). Add --plot to save a bar chart of judge verdicts to figures/judge_summary.png (optionally broken down --group_by one or more columns, e.g. Category or entity_modality).
The causal scripts below (mlp_ablation.py, attention_block.py, backpatch.py) write CSVs with slightly different column names (Parametric_ans/Contextual_ans/New_Answer instead of Ground_Truth/Mis_Answer_Label/Response); point judge.py at those with --parametric_col Parametric_ans --contextual_col Contextual_ans --response_col New_Answer (backpatch_scorer.py already does this internally for backpatch output, so you don't need to run judge.py separately on those).
Zero-ablates MLP outputs over a sliding window of layers at the entity token positions, and measures how the log-probability margin between the parametric and contextual answer shifts.
# Clean (no-ablation) baseline -- needed once per model/domain/modality for the plot's dashed reference lines
python mlp_ablation.py \
--model_name Qwen/Qwen2.5-VL-7B-Instruct --dataset celeb \
--entity_modality vision --ablate_component mlp \
--exp_type ctxt_mem_confl --k_min -1 --k_max -1 --dst_start 0
# Sliding 5-layer window sweep, one call per starting layer
for DST in $(seq 0 23); do
python mlp_ablation.py \
--model_name Qwen/Qwen2.5-VL-7B-Instruct --dataset celeb \
--entity_modality vision --ablate_component mlp \
--exp_type ctxt_mem_confl --window_size 5 --dst_start $DST --k_min 4 --k_max 4
done
# repeat both loops for --entity_modality textThen plot:
python plot_mlp_ablation.py --models qwen --datasets celeb --out_name mlp_ablation_qwen(celeb is the internal dataset tag for the people domain here -- see Note on naming below.) --max_dst lets you override the swept range per model if you didn't use the paper's default (24 for Qwen, 44 for Gemma, 30 for Ministral 3).
Masks attention from entity tokens to the retrieved-context tokens, layer by layer, to see how much of the parametric-vs-contextual preference is attention-mediated. --keep_original_mlps runs the causal-mediation control that isolates the direct effect of the attention mask from its downstream effect on MLPs.
for DST in $(seq 0 27); do
# standard masking
python attention_block.py \
--model_name Qwen/Qwen2.5-VL-7B-Instruct --dataset celeb \
--entity_modality vision --block_ctxt_only \
--max_block_layer $DST
# causal-mediation control
python attention_block.py \
--model_name Qwen/Qwen2.5-VL-7B-Instruct --dataset celeb \
--entity_modality vision --block_ctxt_only --keep_original_mlps \
--max_block_layer $DST
done
# repeat for --entity_modality textThen plot (also reads the mlp_ablation.py clean-baseline CSVs from step 3 as the "Clean Baseline" reference line):
python plot_attention_block.py --models qwen --datasets celebTests whether visual entities are simply slower to resolve: patches the visual-token residual stream from a later layer back into layer 0, and checks whether that's enough to let the (now earlier-resolved) entity representation attend to and suppress the conflicting context.
for SRC in $(seq 14 27); do # Qwen's swept range; see plot_backpatch.py DEFAULT_SRC_RANGE for Gemma/Ministral
OUT="results/interpretability/FRQ_backpatch_entity_resid_ctxt_mem_confl_qwen_celeb_vision_src${SRC}_tgt0_w1_stride1.csv"
python backpatch.py \
--model_name Qwen/Qwen2.5-VL-7B-Instruct --dataset celeb \
--entity_modality vision --exp_type ctxt_mem_confl \
--source_start $SRC --target_start 0 --window_size 1 --layer_stride 1 \
--max_new_tokens 10 --out_csv "$OUT"
python backpatch_scorer.py --csv "$OUT"
donethen:
python plot_backpatch.py --models qwen --datasets celebA ready-to-submit SLURM version of this sweep (Gemma-3 on the artwork domain, matching the paper's Figure 5) is at examples/sbatch/backpatch_score_gemma_artwork.sh.
src/
paths.py # all path configuration -- override via env vars, see Setup
util.py # str2bool
make_inputs.py # builds the actual multimodal prompt (shared by every script)
classify_response.py # judge prompts (verbatim)
data_filtering_utils.py # filters RAG output down to "model knew this before conflict" rows
download_data.py
ctxt_mem_experiment.py
judge.py
mlp_ablation.py, plot_mlp_ablation.py
attention_block.py, plot_attention_block.py
backpatch.py, backpatch_scorer.py, plot_backpatch.py
examples/sbatch/
backpatch_score_gemma_artwork.sh
environment_qwen_gemma.yml, environment_mistral.yml
requirements-qwen-gemma.txt, requirements-mistral.txt
TBD