Same Semantics, Different Outcome: On the Modality Robustness of Multimodal LLMs under Knowledge Conflict
Abstract
Multimodal large language models (MLLMs) are increasingly provided with contextual evidence in heterogeneous forms: as a text passage, as a rendered image of the same passage, or as both together. However, it remains unclear how consistently these surface forms are processed, especially when the evidence conflicts with the model’s parametric knowledge. We study modality robustness under knowledge conflict across 13 MLLMs and two datasets, and find them far from robust. (1) Contrary to common belief, models favor a context that contradicts parametric knowledge more readily in image form than in text form; (2) when a contradicting text and image are presented together, the preferred modality is essentially arbitrary, varying with input order, model, and dataset. We further demonstrate that this instability has practical consequences: it degrades performance in multimodal RAG and can be exploited by adversarial attacks. To alleviate this brittleness, we examine several simple techniques—prompting, steering, supervised fine-tuning (SFT), and direct preference optimization; the majority prove ineffective, whereas SFT achieves moderate success. We therefore call for greater awareness of this inconsistency and argue that it is fundamental, demanding attention at multiple training stages.
1 Introduction
Multimodal large language models (MLLMs) are increasingly deployed in retrieval-augmented and agentic workflows, where evidence arrives in a variety of surface forms, e.g., plain text, webpages, PDFs, screenshots, or scanned documents Yu et al. (2025); Faysse et al. (2025). These workflows presuppose that MLLMs process semantically equivalent content consistently, irrespective of its modality. That expectation is not guaranteed to hold, however, because inputs in different modalities typically follow distinct internal pathways through these models, beginning with modality-specific tokenization.
The risk posed by such latent inconsistency is greatly amplified under knowledge conflict (KC), where external evidence contradicts the model’s parametric knowledge or where multiple sources support competing answers. In such scenarios, the answer may diverge drastically, or even reverse, according to which source the model accepts as authoritative; if that acceptance is governed by the form in which the evidence is delivered rather than by its content, the robustness and reliability of the model are severely compromised.
Although previous work Deng et al. (2025); Sim et al. (2025); Zhang et al. (2025a) has reported that MLLMs exhibit a modality bias—typically placing more weight on textual than visual information—this claim has not been examined in knowledge-conflict scenarios. Moreover, the few studies on multimodal knowledge conflict Hua et al. (2025); Nguyen et al. (2025); Zhang et al. (2025b); Zhang et al. (2025c) focus mostly on natural images rather than documents, even though knowledge is typically organized in document form, not as scenes or faces.
To fill this gap, we investigate the modality robustness of MLLMs through the lens of knowledge conflict. We evaluate two setups: (1) a single piece of external evidence is delivered as text or as a rendered image (i.e., text-as-image); (2) two textual and visual contexts conflict with each other. Unlike prior work, we also guarantee that parametric knowledge is distinct from any contextual knowledge, enabling more tightly controlled analysis.
Spanning multiple MLLMs, datasets, and conflict categories, the dominant pattern we observe is instability: which modality a model trusts shifts unpredictably with input order, dataset construction, and model family, and no single factor reliably explains the variation. Crucially, this contradicts the widely reported textual bias in MLLMs. In our single-evidence setting, the direction is in fact reversed: when the image is a rendering of the same text, MLLMs tend to prefer it over the text itself. Moreover, even this reversed preference dissolves once both modalities are presented in conflict, a setting that prior work has not thoroughly examined.
We further examine whether this sensitivity has practical consequences beyond controlled configurations, using two downstream tasks. First, in multimodal RAG, we keep a passage’s content fixed and vary only its format—raw text or rendered image—to test whether that change alone alters how models use retrieved evidence (§5.1). We find that MLLMs are less distracted by irrelevant context when salient knowledge is represented as an image rather than text, suggesting a remedy for the lost-in-the-middle phenomenon Liu et al. (2024a).
Second, we examine safety-critical settings, in which the surface form of a harmful request may affect the model’s refusal behavior (§5.2). Our results indicate that rendering harmful requests as images increases attack success rates by 6.99 percentage points on average, underscoring a serious and readily exploitable vulnerability.
Finally, we explore methods to improve the modality robustness of MLLMs. We observe that simple prompting is insufficient; instead, fine-tuning on knowledge-conflict data partially mitigates the problem, although it does not completely eliminate the inconsistency (§6). Consequently, we call for greater community awareness of MLLMs’ inconsistency across input modalities, and argue that this problem is fundamental, requiring further investigation at different stages of model training.
To summarize, our contributions are as follows:
- •
We formalize modality robustness as an evidence-authority problem under knowledge conflict.
- •
We construct single- and multi-evidence conflict settings that disentangle semantic content, presentation modality, and parametric knowledge.
- •
Across 13 MLLMs and two datasets, modality reliance varies across model families, input orders, and evidence compositions, with downstream effects on multimodal RAG and refusal behavior.
- •
Comparing several mitigation strategies, we find conflict-aware fine-tuning to be the most effective at balancing modality preference.
) but reverts to its memory (
) when the same content is delivered as text (
).
(b) Multi-evidence: when and support different answers,
the order in which they appear can flip which modality the model follows.
These examples illustrate a broader instability of MLLMs under knowledge conflict: the information a model prioritizes varies arbitrarily with input order, dataset, and model family.
2 Related Work
2.1 Knowledge Conflict
Knowledge conflict has been studied as the tension between parametric memory and external evidence, formalized via entity substitution (Longpre et al., 2021) and later expanded along multiple conflict axes (Xu et al., 2024; Wang et al., 2024). While these works characterize how models resolve such conflicts (Jin et al., 2024; Khandelwal et al., 2025), they generally assume a fixed modality, i.e., text.
The framework has been recently extended to multimodal settings where context is also presented as images Zhang et al. (2025c); Nguyen et al. (2025); Zhang et al. (2025b); Hua et al. (2025), but does not explicitly control for conflicts with the model’s parametric memory. Other studies test parametric or commonsense knowledge with counterfactual images, observing whether models follow the visual input or their internal beliefs Liu et al. (2025); Ortu et al. (2025). A related line examines inconsistencies between visual entities or attributes and their textual descriptions Jia et al. (2026). Across these studies, however, the evidence is predominantly natural images such as scenes or faces, rather than the knowledge-intensive sentences and paragraphs—presented either as text or as rendered images—that constitute the focus of our work.
2.2 Modality Bias in MLLMs
Recent work reports a “Blind Faith in Text” tendency in MLLMs, where models favor text over visual evidence Deng et al. (2025). Counterfactual-image studies similarly show reliance on language priors over visual cues Lee et al. (2025), and others report text-dominant behavior under controlled evidence conflicts Zhang et al. (2025a). A recent survey synthesizes these findings under the notion of modality collapse Sim et al. (2025).
We extend these efforts to knowledge-conflict scenarios, where consistent handling of input context is essential to the robust behavior of MLLMs. We later show that the textual bias widely reported in the literature does not hold in our setting; the direction in fact often reverses, with MLLMs favoring images over text. Furthermore, this trend grows more arbitrary as multiple pieces of evidence simultaneously contradict parametric knowledge, urging a reassessment of the current consensus.
3 Problem Formulation
In this work, we study modality robustness under knowledge conflict (KC), namely, whether MLLMs resolve knowledge conflicts consistently across different numbers and orderings of textual and visual inputs. Let denote the model’s internal parametric knowledge and the external evidence for query . We specify two competing answer candidates: supported by and supported by , with . To isolate modality, we hold the semantic content of the evidence fixed and present it either as text or as an image . Since this work focuses on knowledge-intensive scenarios, is defined as a rendered image of .11 1 The rendering procedure is provided in Appendix A.1. If the model produces different responses across and , we regard this as a failure of modality robustness.
3.1 Categories of Knowledge Conflict
Previous work typically focuses on either single- or two-evidence settings, often without carefully controlling for parametric knowledge. Our setup covers both configurations and explicitly accounts for parametric alignment. Figure 1 illustrates the model’s unstable responses in both settings, driven by inconsistent multimodal processing under KC.
Single-evidence conflict
For each query , we pair it with external evidence that contradicts , instantiated in two modalities: textual and image . Both support the same answer , conflicting with . We compare two input conditions:
This measures how the modality of evidence affects the model’s resolution of conflicts between parametric () and external () knowledge.
Specifically, we use KC datasets (§3.2) with 5-tuples : a query , a factual context with answer , and a counterfactual context with answer . For each , we determine whether matches or , and align and with and .22 2 Prior work on multimodal KC Ortu et al. (2025) either skipped this verification or assumed parametric memory is always correct, leaving their findings potentially confounded.
To this end, we leverage a closed-book multiple-choice QA setup. The answer choices are taken from the source KC dataset and augmented with a ‘none of the above’ option. Inspired by Wang et al. (2023), we sample the model’s output for five times and retain only when all five responses agree, taking the common answer as .33 3 Unlike Wang et al. (2023), we require unanimous agreement across samples rather than majority voting, yielding a stricter estimate of internal knowledge. If is ‘none of the above’, we ask the model to provide an open-ended answer, which then replaces . Subsequently, if , we align and with and , respectively; otherwise, the alignment is reversed.
Multi-evidence conflict
We further consider a joint setting that provides both modalities to the model, in two possible orderings:
We extend the setup to a three-way conflict by adding a second conflicting source, yielding 7-tuples . For simplicity, we focus on cases where matches the factual pair . We then have two possible assignments: (1) and ; (2) the reverse. We average over both assignments and test whether the model favors one modality in this three-way conflict. Again, model-wise final dataset sizes () are reported in Table 7 of Appendix A.2.
| ConflictQA | NQ-Swap | |||||
| Single | Multi | Multi | Single | Multi | Multi | |
| Claude Sonnet 4.5 | ||||||
| GPT-5.4 | ||||||
| GPT-4o | ||||||
| Qwen3-Omni | ||||||
| Qwen2.5-Omni | ||||||
| MiniCPM-o2.6 | ||||||
| OmniVinci | ||||||
| Qwen2.5-VL-32B | ||||||
| Qwen2.5-VL-7B | ||||||
| InternVL3-38B | ||||||
| InternVL3-8B | ||||||
| LLaVA-v1.6-34B | ||||||
| LLaVA-vicuna-7B | ||||||
3.2 Other Experimental Factors
We consider additional experimental factors that may affect modality-dependent evidence reliance.
Datasets
We evaluate on two complementary knowledge conflict datasets:
- •
ConflictQA Xie et al. (2024) constructs conflicts by eliciting the model’s parametric memory and then generating counter-memory evidence supporting an alternative answer. Since the evidence is LLM-generated rather than produced by simple word substitution (as in NQ-Swap), the resulting passages read naturally.44 4 We use the publicly available GPT-4-generated PopQA Mallen et al. (2023) counterfactuals for and the ChatGPT-generated variant for or vice versa.
- •
NQ-Swap Longpre et al. (2021) is a variant of the Natural Questions dataset Kwiatkowski et al. (2019) where the answer entity is replaced with another entity in several substitution methods. Since NQ-Swap is based on real QA passages, it provides a retrieval-like conflict setting.
Models
We evaluate 13 MLLMs across proprietary API models and open-source model families:
- •
Proprietary models: Claude Sonnet 4.5 Anthropic (2025), GPT-5.4 OpenAI (2026), and GPT-4o Hurst et al. (2024).
- •
Open-source omni-modal LLMs (OLLMs): Qwen3-Omni-30B-A3B-Instruct Xu et al. (2025b), Qwen2.5-Omni-7B Xu et al. (2025a), MiniCPM-o2.6 OpenBMB (2025), and OmniVinci Ye et al. (2025).
- •
Open-source vision-language models (VLMs): Qwen2.5-VL-7B/-32B-Instruct Bai et al. (2025), InternVL3-8B/-38B-Instruct Zhu et al. (2025), LLaVA-v1.6-34b-hf, and LLaVA-v1.6-vicuna-7b-hf Liu et al. (2023).
We test whether modality robustness under KC generalizes across model families and scales.
Metrics
For , we have distinct forms of prompts , e.g., and . We request an LLM to generate its open-ended response given a specific type of three times, and then evaluate them using normalized span matching.55 5 We use normalized span matching to relax exact matching: after lowercasing, punctuation/whitespace cleanup, and alias handling, counts as a match if either the normalized prediction contains or vice versa.
Let = denote a set of three responses for instance under prompt . We define the external-source following rate (EFR) over the dataset from §3.1 as
where becomes 1 only when all three responses match the external answer candidate , and 0 otherwise. By design, this definition is conservative: it credits a model only for consistently following the evidence, not for a single match that could arise by chance.
We quantify modality reliance using the normalized image-over-text gap, denoted by . Since models differ in their overall EFR, we normalize the image–text difference by their combined rate to enable stable cross-model comparison. For single-evidence conflict, we compare two prompts that differ only in the modality of the same external evidence: and . We thus compute as
In the multi-evidence setting, both and are presented jointly, supporting competing answers and . For each ordering , we compute the modality-specific rates and , and define the (order-aware) image-over-text gap as
The resulting gap is bounded within . Positive values indicate that the model relies more on the image, with larger magnitudes corresponding to a wider image–text difference.
4 MLLMs Lack Modality Robustness
Table 1 summarizes the normalized image-over-text gap across models, datasets, input orders, and conflict categories; the corresponding absolute percentage-point gaps are reported in Table 9 in the Appendix. The dominant pattern is instability: modality is not merely a passive container of evidence, but a factor that influences how models resolve conflicts.
Single-evidence conflict: a weak but pervasive image preference.
The single-evidence setting is the most direct test of presentation invariance: the external evidence has the same semantic content across conditions, and only its surface form changes. Positive therefore means that rendering the same evidence as an image makes the model more likely to follow it over its parametric knowledge. This image advantage is most evident in proprietary models. Claude Sonnet 4.5 shows the strongest positive gaps, with +42.86% on ConflictQA and +22.63% on NQ-Swap. GPT-5.4 also remains positive on both datasets, with +15.29% and +3.80%, respectively. The pattern is not universal, however: GPT-4o reverses from a positive gap on ConflictQA (+9.75%) to a negative gap on NQ-Swap (-14.74%), while the LLaVA models show consistently negative gaps, indicating stronger reliance on text evidence. These results show that even the simplest presentation change can alter evidence following.
Multi-evidence conflict: modality order matters.
Table 1 shows that when precedes (), most open-source models show negative gains, indicating stronger reliance on the text-supported answers. However, when the order is reversed (), the same models often shift toward image-supported answers. This order sensitivity is reflected in both the per-model Flip Ratio and the aggregate number of sign flips, as shown in Table 8 in the Appendix. The Flip Ratio further shows instance-level preference changes when the input order is reversed. At the aggregate level, 8 out of 13 models on ConflictQA and 6 out of 13 models on NQ-Swap change the sign of across the two input orders. These reversals show that modality reliance is not fixed; it can be reshaped by the order in which conflicting sources enter the context.
is also influenced by the dataset.
As Table 1 shows in the multi-evidence settings, ConflictQA elicits markedly larger absolute gaps than NQ-Swap; OmniVinci, for example, exceeds 86% on ConflictQA but is nearly neutral on NQ-Swap. We hypothesize that modality-dependent reliance grows with the semantic divergence between conflicting sources. NQ-Swap perturbs only the answer entity and otherwise leaves the passage intact, so the two evidence sources remain largely overlapping; ConflictQA, by contrast, generates entirely separate counter-evidence, producing a much broader semantic gap. Figure 6 in Appendix A.2 supports this characterization.
Variation across model families.
Figure 2 visualizes from Table 1 by grouping open-source models by family and proprietary models into a separate category. The proprietary group maintains the most consistently positive average , indicating a relatively stable image preference. In contrast, LLaVA models remain strongly text-dominant across all conditions and datasets. Qwen and InternVL families exhibit pronounced order sensitivity: their average becomes strongly positive under but negative under , especially on ConflictQA. Together, these patterns indicate that modality-dependent reliance can be influenced by model characteristics and evidence conditions, rather than reflecting a universal property of MLLMs.
5 When Modality Instability Matters
The instability documented in §4 is not merely a benchmark artifact. Deployed MLLM pipelines rarely control how evidence arrives. Even when the underlying evidence remains identical, models may resolve conflicts differently due to upstream choices in evidence representation.
We examine two settings where this failure mode is consequential: question-answering performance in multimodal RAG (§5.1), where format changes what the model attends to, and safety vulnerability to image-rendered adversarial instructions (§5.2).
5.1 Impact of Gold-Passage Modality in RAG
To probe whether modality instability matters in practice, we begin with retrieval-augmented QA. As a thought experiment, imagine that an ideal paragraph directly answering the query is given in advance: how do MLLMs treat this gold-standard passage when only its modality changes? We fix the query, the gold passage’s content, and the distractors (all provided as text for simplicity), and vary only the modality of the gold passage.
We sample 170 query–passage instances from MS MARCO v2.1 (Bajaj et al., 2016) and report answer quality with ROUGE-L. MLLMs are evaluated under four cases: a gold-only setting, in which the gold passage is the only input, and three multi-passage settings, in which it appears among nine distractors at the First, Middle, or Last position.
Figure 3 demonstrates that the modality of the gold-standard passage influences RAG performance, with image rendering mitigating the well-known lost-in-the-middle phenomenon Liu et al. (2024a). For instance, Gemini 2.5 Pro gains +6.59 ROUGE-L when the image-rendered gold passage occupies the middle position. This suggests that rendering relevant context as an image may serve as a new strategy to direct model attention toward salient passages embedded among distractors.
5.2 Vulnerability to Image-Rendered Attacks
As a second case study, we ask whether modality instability also affects MLLMs’ safety. Our experiments build on MM-SafetyBench (Liu et al., 2024b), which benchmarks robustness against attacks that hide harmful content in query-relevant images. Its four conditions---Text-Only, Stable Diffusion (SD), Typography (Typo), and SD+Typo---serve as our baselines and embed only the key phrase in image form.66 6 SD generates images from prompts, while Typo overlays the key phrase on a white background; see Appendix G. We extend this setup with an image-rendered instruction condition, in which the full instruction is moved into the visual channel, leaving the request semantically unchanged, and report Attack Success Rate (ASR), the proportion of prompts for which the model produces harmful content instead of refusing.
Figure 4 summarizes the increase in ASR from rendering harmful instructions as images across models and baseline conditions, averaged over all 13 scenarios. Over 156 evaluation settings in total, this intervention increases ASR by 6.99% on average, with positive ASR in 126/156 (80.8%) comparisons. Table 15 in the Appendix reports the corresponding per-scenario results. The average increase is consistently positive across all model baseline cells, although its magnitude varies across models (Qwen +3.74%, InternVL +7.93%, LLaVA +9.30%). It is largest against the Text-only baseline (+11.60%) and smallest against SD (+4.97%). Together, these results suggest that moving the whole instruction into the visual channel introduces an additional safety vulnerability beyond existing attacks that embed only parts of the instruction visually.
6 Mitigating Modality Instability
Having established the modality instability of MLLMs and its practical consequences, we next explore whether this behavior can be mitigated. We consider interventions at three levels: a lightweight intervention through explicit prompting, post-training through supervised fine-tuning and direct preference optimization, and inference-time control through representation steering.
6.1 Prompting is Not Sufficient
A natural first step is to instruct the model directly to use both modalities in balance. We therefore prepend a system-level prompt asking the model to jointly consider textual and visual evidence whenever multiple sources are provided. Under the hypothesis that underspecified instructions drive the imbalance, this explicit multimodal prompting should close the gap.
Contrary to expectation, as shown in Table 2, prompting does not provide reliable control over modality reliance. While it partially reduces the imbalance for Qwen2.5-VL-7B, it instead amplifies the bias for GPT-5.4. We observe similar instability when the prompt explicitly asks the model to focus on a particular modality. These outcomes suggest that modality-dependent reliance cannot be reliably corrected through surface-level prompting alone, motivating the post-training and inference-time interventions explored next.
| Model | Orig. | +Prompt. | Orig. | +Prompt. |
|---|---|---|---|---|
| GPT-5.4 | 1.80 | 71.26 | 63.37 | 80.49 |
| Qwen2.5-VL | -68.97 | -17.54 | 73.82 | 45.69 |
6.2 Conflict-Aware Fine-Tuning
We here examine whether this modality-dependent reliance can be recalibrated through post-training. Unlike prompting, which intervenes only through instructions at inference time, fine-tuning directly adjusts the model’s behavior under conflicting multimodal evidence.
To test whether reliance on misleading visual evidence can be recalibrated in a reproducible open-source setting, we select Qwen2.5-VL-7B-Instruct, which exhibits strong image-following behavior. We then apply conflict-aware LoRA fine-tuning using training data constructed from NQ-Swap and MMM-Fact (Xu et al., 2025c).77 7 For NQ-Swap, we train the model to produce the ground-truth answer despite a counterfactual image-rendered passage. For MMM-Fact, we use non-supporting image-description pairs, where the image and the textual description disagree, to encourage cross-modal comparison. See Appendix C.1. This tuning objective reduces the model’s tendency to follow misleading image evidence while preserving its use of factual evidence during conflict resolution.
As shown in Table 3, conflict-aware SFT consistently reduces the magnitude of the modality gap across all evaluated settings. The effect is particularly pronounced in the multi-evidence condition, where the image-dominant gap decreases from 73.82% to 52.04%. The opposite order shows a modest reduction in the text-dominant gap. The intervention also reduces modality imbalance in the single-evidence setting for both misleading (-2.29%) and reliable (-2.37%) evidence.
Importantly, moving toward zero does not by itself guarantee desirable evidence use, since balanced modality reliance could in principle arise from uniformly increasing or decreasing evidence following. We therefore additionally examine the corresponding raw in Appendix C.1. Table 12 shows that SFT reduces following of misleading image evidence while preserving, and slightly increasing, following of reliable image evidence. These findings provide evidence that conflict-aware SFT can improve modality balance without merely suppressing visual evidence or shifting preference toward the opposite modality.
| Evidence | False | True | ||
|---|---|---|---|---|
| Method | Single | Single | ||
| Original | -68.97 | 73.82 | 2.54 | 3.45 |
| + SFT | -67.07 | 52.04 | 0.25 | 1.08 |
6.3 Additional Mitigation Baselines
To broaden our mitigation analysis beyond prompting and conflict-aware SFT, we additionally evaluate two approaches operating at different stages: Direct Preference Optimization (DPO) Rafailov et al. (2023) as a post-training method and representation-level steering as an inference-time intervention. Detailed implementation settings and the corresponding raw image-following rates are provided in Appendix C.2 and C.3.
Preference-based post-training.
We construct NQ-Swap preference pairs that favor the original factual answer over the answer supported by the substituted misleading image, and use them to train Qwen2.5-VL-7B-Instruct with DPO. Table 4 reports that DPO slightly reduces the image-dominant gap under , but substantially increases the text-dominant gap under . Thus, DPO can shift reliance away from misleading image evidence in some cases, but does not improve modality balance across different input orders.
Representation-level steering.
We further evaluate an inference-time steering baseline following prior work Zhang et al. (2025a) on the same Qwen2.5-VL-7B model. Using a calibration subset of ConflictQA, we derive separate text-following steering directions for each input order and apply them during decoding on a held-out evaluation subset. As the steering strength increases, the image-dominant gap under decreases from 77.77% to 60.00%. In contrast, the text-dominant gap under grows. At stronger steering scales, modality preference shifts toward text under both input orders.
Overall, both methods exhibit a directional trade-off: reducing image-dominant behavior can coincide with stronger text-dominant behavior. This contrasts with conflict-aware SFT, which moves every closer to zero shown in Table 3. These observations highlight an important distinction between controlling modality preference and improving modality robustness; effective mitigation should reduce imbalance in both directions rather than simply shifting reliance toward a fixed modality.
| Method | ||
|---|---|---|
| Preference-based post-training | ||
| Original | -68.89 | 68.47 |
| + DPO | -88.08 | 65.57 |
| Representation-level steering | ||
| Original (Scale 0) | -38.64 | 77.77 |
| Scale 1 | -45.05 | 80.64 |
| Scale 3 | -70.21 | 64.95 |
| Scale 5 | -86.96 | 60.00 |
7 Further Analysis: Input-Side Sensitivity
Given that input processing is the stage where modality-specific handling is most pronounced, we ask whether input-side interventions—altering how visual evidence is preprocessed and rendered—can change a model’s modality reliance. We examine this with frozen model weights, intervening solely on the visual preprocessing pipeline and the rendered form of semantically identical evidence.
7.1 Influence of Visual Preprocessing
The results in §4 show that modality reliance varies markedly across model families. To probe one possible source of this variation, we focus on Qwen2.5-VL and LLaVA-v1.6-7B, two families with opposing tendencies: Qwen leans toward image evidence, LLaVA leans toward text. Using this contrast, we examine whether preprocessing-level interventions can shift the model away from modality-dependent instability. We apply bidirectional interventions between the two models. Specifically, Qwen2.5-VL receives LLaVA-style 336336 global/local view construction, whereas LLaVA-v1.6 receives images pre-resized by Qwen’s smart-resize rule. All remaining model-specific processing is unchanged; see Appendix D for details.
| Base Model | Preprocessing | Orig. | After Prep. |
|---|---|---|---|
| Qwen2.5-VL | LLaVA-style | 9.71 | -37.23 |
| LLaVA-v1.6 | Qwen-style | -34.92 | -30.91 |
Table 5 reveals a clear but asymmetric effect across intervention directions. Under LLaVA-style preprocessing, Qwen’s moves from 9.71% to -37.23%, not merely flipping sign but landing within three percentage points of LLaVA’s native value (-34.92%). Input-side processing alone is thus sufficient to reproduce LLaVA’s modality profile in a model with different weights and training. The reverse intervention produces only a modest shift, from -34.92% to -30.91%.
The key implication of this analysis is that preprocessing is one factor in modality reliance but not the sole mechanism; our setup also cannot isolate which downstream components interact with it. A more comprehensive test would require training matched models that differ only in preprocessing, which we leave to future work.
7.2 Influence of Image Presentation
A natural concern is whether our findings depend on arbitrary choices in how textual evidence is rendered as images. To address this, we perturb the default rendering at two levels of increasing visual salience and measure how the model’s source-following behavior shifts.
On the NQ-Swap subset, we change only the visual appearance of the rendered evidence, leaving the query and evidence content untouched. The first level applies global changes to the entire image, either by rendering all text in red or by setting a yellow background. The second level applies the same changes only to the targeted answer-bearing span; this serves as an upper-bound salience diagnostic rather than a deployable setting, since it presupposes knowledge of the answer.
Figure 5 shows that modest global visual changes produce only small shifts relative to the default rendering, whereas answer-span highlighting produces much larger shifts. These findings both strengthen and qualify our main finding: the modality instability documented in §4 is not an artifact of font, color, or background choice, since reasonable variations along those dimensions leave the gain magnitudes largely intact. Yet the answer-span result indicates that image evidence carries an additional axis of variability that text evidence does not. This aligns with the effectiveness of image-rendered attacks observed in §5.2, suggesting that the influence of image evidence is shaped not only by its content but also by its visual form.
8 Conclusion
We study whether MLLMs resolve knowledge conflicts consistently across modalities and find that they do not: the same conflict yields different answers depending on how the evidence is delivered, and the direction of preference shifts with dataset, input order, and evidence composition. The fixed text or image preference reported in prior work does not survive once evidence is allowed to vary along these axes, and no single factor we examine reliably predicts which modality a model will favor. Modality preference is thus better understood as an artifact of the evaluation setup than as a fixed property of the model.
This instability has practical consequences, both beneficial (e.g., highlighting salient passages in RAG) and harmful (e.g., more potent adversarial attacks); the simple remedies we explored were largely ineffective, with conflict-aware fine-tuning offering only partial mitigation. The origin of this instability remains an open question. Our preliminary findings identify the visual input pipeline as a tractable starting point, but a complete account will require examining multiple stages of model training, which we leave to future work.
Limitations
First, we consider only text-modality evidence and its rendered-image form, and do not examine naturally occurring visual evidence (e.g., figures, charts, photographs) or other modalities such as audio, where modality dynamics may differ. Extending our framework to these broader modalities is a direction for future work. Second, while we evaluate 13 MLLMs across proprietary and open-source families, our coverage is not exhaustive. Third, while we identify several contributing factors, a more thorough mechanistic account of why modality reliance shifts in the ways we document remains to be developed.
Ethical Statement
Our safety experiments are intended to diagnose modality-dependent weaknesses in MLLM refusal behavior, not to provide new attack recipes. We report aggregate ASR results and do not include concrete harmful instructions or unsafe generations. Any released artifacts will exclude harmful rendered prompts and will be limited to evaluation metadata and non-sensitive analysis code.
Acknowledgments
This work was supported by Institute of Information & communications Technology Planning & Evaluation (IITP) grant funded by the Korea government(MSIT) (No.RS-2020-II201373, Artificial Intelligence Graduate School Program(Hanyang University)). This work was supported by Institute of Information & communications Technology Planning & Evaluation (IITP) under the artificial intelligence semiconductor support program to nurture the best talents (IITP-2026-RS-2023-00253914) grant funded by the Korea government(MSIT). This work was supported by the National Research Foundation of Korea(NRF) grant funded by the Korea government(MSIT) (RS-2025-00558151).
References
- Anthropic (2025) Anthropic. 2025. Claude 4.5 Sonnet.
- Bai et al. (2025) Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, and 8 others. 2025. Qwen2.5-vl technical report. arXiv preprint arXiv:2502.13923.
- Bajaj et al. (2016) Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, and 1 others. 2016. Ms marco: A human generated machine reading comprehension dataset. arXiv preprint arXiv:1611.09268.
- Deng et al. (2025) Ailin Deng, Tri Cao, Zhirui Chen, and Bryan Hooi. 2025. Words or vision: Do vision-language models have blind faith in text? 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3867–3876.
- Faysse et al. (2025) Manuel Faysse, Hugues Sibille, Tony Wu, Bilel Omrani, Gautier Viaud, Céline Hudelot, and Pierre Colombo. 2025. ColPali: Efficient document retrieval with vision language models. In The Thirteenth International Conference on Learning Representations.
- Hu et al. (2024) Jian Hu, Xibin Wu, Zilin Zhu, Xianyu, Weixun Wang, Dehao Zhang, and Yu Cao. 2024. Openrlhf: An easy-to-use, scalable and high-performance rlhf framework. arXiv preprint arXiv:2405.11143.
- Hua et al. (2025) Tianze Hua, Tian Yun, and Ellie Pavlick. 2025. How do vision-language models process conflicting information across modalities? arXiv preprint arXiv:2507.01790.
- Hurst et al. (2024) Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, and 1 others. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276.
- Jia et al. (2026) Yifan Jia, Yuntao Du, Kailin Jiang, Yuyang Liang, Qihan Ren, Yi Xin, Rui Yang, Fenze Feng, MingCai Chen, Hengyang Lu, and 1 others. 2026. Benchmarking multimodal knowledge conflict for large multimodal models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 22283–22291.
- Jin et al. (2024) Zhuoran Jin, Pengfei Cao, Hongbang Yuan, Yubo Chen, Jiexin Xu, Huaijun Li, Xiaojian Jiang, Kang Liu, and Jun Zhao. 2024. Cutting off the head ends the conflict: A mechanism for interpreting and mitigating knowledge conflicts in language models. In Findings of the Association for Computational Linguistics: ACL 2024, pages 1193–1215, Bangkok, Thailand. Association for Computational Linguistics.
- Khandelwal et al. (2025) Anant Khandelwal, Manish Gupta, and Puneet Agrawal. 2025. CoCoA: Confidence- and context-aware adaptive decoding for resolving knowledge conflicts in large language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 6846–6866, Suzhou, China. Association for Computational Linguistics.
- Kwiatkowski et al. (2019) Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019. Natural questions: A benchmark for question answering research. Transactions of the Association for Computational Linguistics, 7:452–466.
- Lee et al. (2025) Kang-il Lee, Minbeom Kim, Seunghyun Yoon, Minsung Kim, Dongryeol Lee, Hyukhun Koh, and Kyomin Jung. 2025. VLind-bench: Measuring language priors in large vision-language models. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 4129–4144, Albuquerque, New Mexico. Association for Computational Linguistics.
- Liu et al. (2023) Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual instruction tuning. Advances in neural information processing systems, 36:34892–34916.
- Liu et al. (2024a) Nelson F. Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. 2024a. Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12:157–173.
- Liu et al. (2025) Xiaoyuan Liu, Wenxuan Wang, Youliang Yuan, Jen-tse Huang, Qiuzhi Liu, Pinjia He, and Zhaopeng Tu. 2025. Insight over sight: Exploring the vision-knowledge conflicts in multimodal LLMs. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 17825–17846, Vienna, Austria. Association for Computational Linguistics.
- Liu et al. (2024b) Xin Liu, Yichen Zhu, Jindong Gu, Yunshi Lan, Chao Yang, and Yu Qiao. 2024b. Mm-safetybench: A benchmark for safety evaluation of multimodal large language models. In European Conference on Computer Vision, pages 386–403. Springer.
- Longpre et al. (2021) Shayne Longpre, Kartik Perisetla, Anthony Chen, Nikhil Ramesh, Chris DuBois, and Sameer Singh. 2021. Entity-based knowledge conflicts in question answering. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 7052–7063, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
- Mallen et al. (2023) Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9802–9822, Toronto, Canada. Association for Computational Linguistics.
- Nguyen et al. (2025) Trang Nguyen, Jackson Michaels, Madalina Fiterau, and David Jensen. 2025. Challenges in understanding modality conflict in vision-language models. arXiv preprint arXiv:2509.02805.
- OpenAI (2026) OpenAI. 2026. GPT-5.4.
- OpenBMB (2025) OpenBMB. 2025. MiniCPM-o Series. https://github.com/OpenBMB/MiniCPM-o.
- Ortu et al. (2025) Francesco Ortu, Zhijing Jin, Diego Doimo, and Alberto Cazzaniga. 2025. When seeing overrides knowing: Disentangling knowledge conflicts in vision-language models. In Mechanistic Interpretability Workshop at NeurIPS 2025.
- Rafailov et al. (2023) Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. Advances in neural information processing systems, 36:53728–53741.
- Sim et al. (2025) Mong Yuan Sim, Wei Emma Zhang, Xiang Dai, and Biaoyan Fang. 2025. Can VLMs actually see and read? a survey on modality collapse in vision-language models. In Findings of the Association for Computational Linguistics: ACL 2025, pages 24452–24470, Vienna, Austria. Association for Computational Linguistics.
- Wang et al. (2023) Xuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V Le, Ed H. Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou. 2023. Self-consistency improves chain of thought reasoning in language models. In The Eleventh International Conference on Learning Representations.
- Wang et al. (2024) Yike Wang, Shangbin Feng, Heng Wang, Weijia Shi, Vidhisha Balachandran, Tianxing He, and Yulia Tsvetkov. 2024. Resolving knowledge conflicts in large language models. In First Conference on Language Modeling.
- Xie et al. (2024) Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su. 2024. Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts. In The Twelfth International Conference on Learning Representations.
- Xu et al. (2025a) Jin Xu, Zhifang Guo, Jinzheng He, Hangrui Hu, Ting He, Shuai Bai, Keqin Chen, Jialin Wang, Yang Fan, Kai Dang, and 1 others. 2025a. Qwen2. 5-omni technical report. arXiv preprint arXiv:2503.20215.
- Xu et al. (2025b) Jin Xu, Zhifang Guo, Hangrui Hu, Yunfei Chu, Xiong Wang, Jinzheng He, Yuxuan Wang, Xian Shi, Ting He, Xinfa Zhu, and 1 others. 2025b. Qwen3-omni technical report. arXiv preprint arXiv:2509.17765.
- Xu et al. (2024) Rongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang, Hongru Wang, Yue Zhang, and Wei Xu. 2024. Knowledge conflicts for LLMs: A survey. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 8541–8565, Miami, Florida, USA. Association for Computational Linguistics.
- Xu et al. (2025c) Wenyan Xu, Dawei Xiang, Tianqi Ding, and Weihai Lu. 2025c. Mmm-fact: A multimodal, multi-domain fact-checking dataset with multi-level retrieval difficulty. arXiv preprint arXiv:2510.25120.
- Ye et al. (2025) Hanrong Ye, Chao-Han Huck Yang, Arushi Goel, Wei Huang, Ligeng Zhu, Yuanhang Su, Sean Lin, An-Chieh Cheng, Zhen Wan, Jinchuan Tian, and 1 others. 2025. Omnivinci: Enhancing architecture and data for omni-modal understanding llm. arXiv preprint arXiv:2510.15870.
- Yu et al. (2025) Shi Yu, Chaoyue Tang, Bokai Xu, Junbo Cui, Junhao Ran, Yukun Yan, Zhenghao Liu, Shuo Wang, Xu Han, Zhiyuan Liu, and Maosong Sun. 2025. VisRAG: Vision-based retrieval-augmented generation on multi-modality documents. In The Thirteenth International Conference on Learning Representations.
- Zhang et al. (2025a) Yu Zhang, Jinlong Ma, Yongshuai Hou, Xuefeng Bai, Kehai Chen, Yang Xiang, Jun Yu, and Min Zhang. 2025a. Evaluating and steering modality preferences in multimodal large language model. arXiv preprint arXiv:2505.20977.
- Zhang et al. (2025b) Zhuoran Zhang, Tengyue Wang, Xilin Gong, Yang Shi, Haotian Wang, Di Wang, and Lijie Hu. 2025b. When modalities conflict: How unimodal reasoning uncertainty governs preference dynamics in mllms. arXiv preprint arXiv:2511.02243.
- Zhang et al. (2025c) Zongmeng Zhang, Wengang Zhou, Jie Zhao, and Houqiang Li. 2025c. Robust multimodal large language models against modality conflict. In Forty-second International Conference on Machine Learning.
- Zhu et al. (2025) Jinguo Zhu, Weiyun Wang, Zhe Chen, Zhaoyang Liu, Shenglong Ye, Lixin Gu, Hao Tian, Yuchen Duan, Weijie Su, Jie Shao, and 1 others. 2025. Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models. arXiv preprint arXiv:2504.10479.
Appendix
Appendix A Evaluation Data and Input Construction
A.1 Text-to-Image Rendering
We render textual evidence as image using the Python Pillow library. Unless otherwise specified, the following rendering template is adopted: black text on a white background, with a font family DejaVuSans, a font size of 24, a maximum text width of 800 pixels, and a line-spacing ratio of 1.25. The rendered text is center-aligned within the image. Line breaks are determined by rendered pixel width: words are wrapped to the next line when they exceed the maximum width, and overlong words are split at the character level. The final image size is then set adaptively from the wrapped text bounding box, with an additional 8-pixel padding to prevent clipping.
Preliminary checks showed that changes in font size(24-32) and text width(600-1,000) had little effect on model behavior. We therefore fix the rendering template in the main experiments and focus on the effect of delivery modality.
A.2 Dataset Preprocessing and Instance Selection
We use two knowledge-conflict datasets, ConflictQA and NQ-Swap, and apply lightweight preprocessing before constructing our conflict instances.
ConflictQA.
We use the GPT-4-generated version of ConflictQA, which is constructed from PopQA and contains 9,544 instances. Each instance provides an original answer–context pair and a counter-answer–context pair. We exclude instances whose evidence context contains non-informative responses such as “don’t know” or “I cannot answer”, since such contexts do not provide usable external evidence for a conflicting answer.
NQ-Swap.
NQ-Swap comprises 4,746 Natural Questions variants in which the answer entity in each original context is replaced with another entity to create conflicting evidence. We exclude contexts containing excessive markup, particularly repeated <Table> tags, because they disrupt the linear reading order and produce cluttered or structurally ambiguous image renderings, making it difficult to maintain semantic alignment between the text and rendered-image conditions.
Model-specific filtering.
After preprocessing, we identify the memory-supported candidate separately for each model using the closed-book prediction procedure described in §3 and Figure 9. We retain only instances with consistent predictions across five runs and a valid opposing answer–context pair that can serve as external evidence. Therefore, the final number of usable single-evidence conflict pairs differs by model. Table 6 and Table 7 report the retained instance counts for each model and dataset.
Multi-evidence input extension.
For multi-evidence experiments, we require two distinct external evidence sources for the same query. For ConflictQA, we use the ChatGPT-generated variant in addition to the default GPT-4-generated data, and use a differently tagged context for the same query as the second external input. For NQ-Swap, when multiple substituted contexts are available for the same original context, we use another substitution as the second external evidence source. Table 7 reports the number of unique multi-evidence pairs constructed in this way. For each pair, we evaluate both modality assignments, swapping which context is presented as text and which is rendered as an image. Thus, the number of evaluated cross-modal inputs is twice the number of unique pairs in Table 7 for each presentation order.
| Model | ConflictQA | NQ-Swap |
|---|---|---|
| Claude Sonnet 4.5 | 1,198 | 3,205 |
| GPT-5.4 | 2,695 | 3,270 |
| GPT-4o | 4,049 | 3,310 |
| Qwen3-Omni | 1,670 | 3,047 |
| Qwen2.5-Omni | 694 | 2,532 |
| MiniCPM-o2.6 | 1,170 | 2,863 |
| OmniVinci | 626 | 2,780 |
| Qwen2.5-VL-32B | 1,019 | 2,904 |
| Qwen2.5-VL-7B | 718 | 2,169 |
| InternVL3-38B | 740 | 2,859 |
| InternVL3-8B | 676 | 2,723 |
| LLaVA-v1.6-34B | 589 | 3,050 |
| LLaVA-v1.6-7B | 907 | 2,671 |
| Model | ConflictQA | NQ-Swap |
|---|---|---|
| Claude Sonnet 4.5 | 442 | 825 |
| GPT-5.4 | 972 | 792 |
| GPT-4o | 885 | 727 |
| Qwen3-Omni | 313 | 743 |
| Qwen2.5-Omni | 181 | 632 |
| MiniCPM-o2.6 | 344 | 732 |
| OmniVinci | 199 | 716 |
| Qwen2.5-VL-32B | 253 | 721 |
| Qwen2.5-VL-7B | 135 | 461 |
| InternVL3-38B | 171 | 721 |
| InternVL3-8B | 184 | 703 |
| LLaVA-v1.6-34B | 226 | 802 |
| LLaVA-v1.6-7B | 335 | 686 |
Dataset-Level Evidence Divergence
To examine the dataset-dependent modality gap, we quantify the divergence between the paired evidence passages within each dataset. For ConflictQA, we compare the separately generated parametric_memory and counter_memory passages, which support the parametric and counter answers, respectively. For NQ-Swap, we compare org_context with sub_context, where the latter is constructed by replacing the answer entity in the original passage.
As shown in Figure 6, ConflictQA exhibits a much larger relative token-length gap than NQ-Swap (0.372 vs. 0.010), while showing substantially lower token LCS similarity (0.188 vs. 0.942) and TF–IDF cosine similarity (0.126 vs. 0.922). Moreover, after masking the paired answer strings (memory_short_answer/counter_short_answer for ConflictQA and org_answer/sub_answer for NQ-Swap), 97.9% of NQ-Swap passage pairs become identical, compared with 0.0% of ConflictQA pairs. The overlap difference also remains after matching examples by passage length. These observations indicate that NQ-Swap primarily applies a localized answer substitution, whereas ConflictQA constructs globally distinct counter-evidence.
Appendix B Extended Main Results
B.1 Order Sensitivity in Multi-evidence Conflicts
Beyond aggregate modality preference in Table 1, we further examine whether individual predictions are sensitive to evidence ordering. Table 8 reports the flip ratio when reversing the order of multimodal evidence. A high flip ratio indicates that the model does not consistently prioritize one evidence source, but changes its preference depending on presentation order.
| Model | ConflictQA | NQ-Swap |
|---|---|---|
| Claude Sonnet 4.5 | 4.35 | 21.47 |
| GPT-5.4 | 16.10 | 17.59 |
| GPT-4o | 19.14 | 14.48 |
| Qwen3-Omni | 48.57 | 26.95 |
| Qwen2.5-Omni | 38.77 | 33.01 |
| MiniCPM-o2.6 | 23.70 | 14.68 |
| OmniVinci | 66.81 | 34.07 |
| Qwen2.5-VL-32B | 33.33 | 27.78 |
| Qwen2.5-VL-7B | 60.03 | 28.57 |
| InternVL3-38B | 39.07 | 32.69 |
| InternVL3-8B | 66.55 | 28.97 |
| LLaVA-v1.6-34B | 0.89 | 1.05 |
| LLaVA-v1.6-7B | 0.00 | 0.15 |
| Models with Sign Flip | 8/13 | 6/13 |
| ConflictQA | NQ-Swap | |||||
| Single | Multi | Multi | Single | Multi | Multi | |
| Claude Sonnet 4.5 | ||||||
| GPT-5.4 | ||||||
| GPT-4o | ||||||
| Qwen3-Omni | ||||||
| Qwen2.5-Omni | ||||||
| MiniCPM-o2.6 | ||||||
| OmniVinci | ||||||
| Qwen2.5-VL-32B | ||||||
| Qwen2.5-VL-7B | ||||||
| InternVL3-38B | ||||||
| InternVL3-8B | ||||||
| LLaVA-v1.6-34B | ||||||
| LLaVA-vicuna-7B | ||||||
B.2 Absolute EFR Differences
Table 9 reports the signed differences between image- and text-supported evidence following rates (EFR) underlying the normalized modality preference scores in Table 1. While the normalized metric captures relative modality preference, these values reveal the absolute magnitude of the underlying modality effect.
B.3 Four Evidence Orderings
Table 10 disaggregates the multi-evidence gain (defined in §3.2) into all four orderings , exposing both the cross-modality ordering effect and the same-modality controls hidden by the aggregation in Table 1.
| ConflictQA | NQ-Swap | |||||||
|---|---|---|---|---|---|---|---|---|
| Claude Sonnet 4.5 | ||||||||
| GPT-5.4 | ||||||||
| GPT-4o | ||||||||
| Qwen3-Omni | ||||||||
| Qwen2.5-Omni | ||||||||
| MiniCPM-o2.6 | ||||||||
| OmniVinci | ||||||||
| Qwen2.5-VL-32B | ||||||||
| Qwen2.5-VL-7B | ||||||||
| InternVL3-38B | ||||||||
| InternVL3-8B | ||||||||
| LLaVA-v1.6-34B | ||||||||
| LLaVA-vicuna-7B | ||||||||
B.4 Same- vs Cross-modality Aggregation
Table 11 reports the same multi-evidence results aggregated by whether the two evidences share a modality. This grouping highlights the effect of modality composition separately from the specific evidence ordering.
| ConflictQA | NQ-Swap | |||
|---|---|---|---|---|
| Same | Different | Same | Different | |
| Claude Sonnet 4.5 | ||||
| GPT-5.4 | ||||
| GPT-4o | ||||
| Qwen3-Omni | ||||
| Qwen2.5-Omni | ||||
| MiniCPM-o2.6 | ||||
| OmniVinci | ||||
| Qwen2.5-VL-32B | ||||
| Qwen2.5-VL-7B | ||||
| InternVL3-38B | ||||
| InternVL3-8B | ||||
| LLaVA-v1.6-34B | ||||
| LLaVA-vicuna-7B | ||||
Appendix C Post-Training Implementation Details
Unlike the main experiments, which use vLLM for inference, the mitigation experiments in §6 are implemented with Hugging Face Transformers to support fine-tuning and hidden-state interventions. Minor differences from the main results may therefore arise due to the different inference backends.
C.1 Conflict-aware Fine-tuning
| Evidence | False | True | |
|---|---|---|---|
| Method | |||
| Original | 43.84 | 92.89 | 80.36 |
| + SFT | 33.52 | 83.26 | 83.93 |
Our goal is to test whether reliance on misleading visual evidence can be recalibrated in a reproducible open-source setting. We use Qwen2.5-VL-7B-Instruct as the tuning target since it supports controlled fine-tuning and exhibits substantial image-following behavior in the single- and multi-evidence settings, respectively (Table 1). This provides a meaningful setting in which to test whether post-training can improve modality robustness.
The training set contains 2,620 examples constructed from NQ-Swap and MMM-Fact, using a 7:3 mixture: 1,834 examples from NQ-Swap and 786 contradiction examples from MMM-Fact. In the NQ-Swap portion, each instance pairs a factual question with an image-rendered counterfactual passage that supports a substituted answer. The target output is the original factual answer rather than the answer implied by the misleading image. The MMM-Fact portion consists of image–claim pairs labeled as ‘contradiction’ which visual and textual information disagree. Together, these examples expose the model to situations in which visual evidence should not be followed solely because of its presentation modality.
We perform LoRA-based supervised fine-tuning using Hugging Face Transformers and OpenRLHF Hu et al. (2024). The model is trained for one epoch with a learning rate of , global batch size 8, micro-batch size 1, LoRA rank/alpha 64/64, and gradient checkpointing.
In Table 3, conflict-aware SFT moves the normalized modality gap closer to zero in all evaluated settings, indicating more balanced modality reliance. However, a smaller alone does not reveal whether this balance results from desirable evidence use. To verify that this reduction does not simply result from uniformly suppressing image following, we additionally report the raw image-following rates in Table 12. For misleading evidence, decreases from 92.89% to 83.26% in the single-evidence setting and from 43.84% to 33.52% in the multi-evidence setting in average. In contrast, for reliable single evidence, increases from 80.36% to 83.93%. Collectively, these results show that the reduced modality gap is accompanied by selective changes in evidence following: the model relies less on misleading images while preserving, and slightly increasing, its reliance on reliable ones.
C.2 Preference-based post-training
| Method | ||
|---|---|---|
| Preference-based post-training | ||
| Original | 41.11 | – |
| + DPO | 38.15 | -2.96 |
| Representation-level steering | ||
| Original (Scale 0) | 42.28 | – |
| Scale 1 | 40.07 | -2.21 |
| Scale 3 | 34.55 | -7.73 |
| Scale 5 | 30.15 | -12.13 |
We additionally evaluate Direct Preference Optimization (DPO) as a preference-based post-training baseline. While conflict-aware SFT directly supervises the desired response, DPO provides an alternative objective that explicitly contrasts a preferred response against an undesired response.
Preference data construction.
We construct 1,834 preference pairs from NQ-Swap. Each example contains a factual question together with a substituted image-rendered passage that supports a counterfactual answer. The original factual answer is used as the chosen response, while the answer supported by the misleading image is used as the rejected response. Thus, the preference objective favors the factual response over one that follows the substituted visual evidence.
Training setup.
We apply DPO to Qwen2.5-VL-7B-Instruct using LoRA. The model is trained for one epoch with , a learning rate of , and LoRA rank/alpha of 64/64. This setup provides a preference-based counterpart to the supervised conflict-aware fine-tuning evaluated in §6.2.
Evaluation.
he resulting model is evaluated on the ConflictQA multi-evidence setting. This allows us to test whether the learned preference transfers to unseen conflict examples rather than merely reproducing the training distribution.
We report the normalized image-over-text gap separately for the two input orders in Table 4. Before DPO, the model exhibits a text-dominant gap of -68.89% under and an image-dominant gap of 68.47% under . After DPO, the latter is slightly reduced to 65.57%, whereas the former becomes substantially more negative, reaching -88.08%. Hence, preference optimization reduces image dominance in one order but simultaneously strengthens text dominance in the other.
For completeness, Table 13 reports the corresponding raw image-following rates. Averaged across the two input orders, decreases from 41.11% to 38.15%, a reduction of 2.96 percentage points. This confirms that DPO reduces overall following of misleading image evidence. However, the order-specific values show that this decrease does not reflect uniformly improved modality balance: DPO slightly reduces image dominance under but further strengthens text dominance under .
Taken together, DPO results illustrate the distinction between suppressing misleading-image following and achieving modality robustness. Preference-based post-training can alter modality reliance, but an objective that consistently improves balance across both input orders may require explicitly accounting for modality preference in both directions.
C.3 Representation-level Steering
We evaluate whether modality preference can be controlled at inference time through activation steering, following prior work on modality-preference steering. We focus on the multi-evidence of ConflictQA, where the image provides misleading evidence while the textual evidence supports the correct answer.
Data split.
We apply the method to the same model, Qwen2.5-VL-7B-Instruct and divide the evaluation examples into a calibration subset and a disjoint held-out evaluation subset. The calibration subset is used to derive the steering directions, and all reported results are measured on the held-out evaluation subset.
Constructing steering directions.
We first run the model on the calibration subset and separate examples into text-following and image-following groups, and derive the corresponding steering direction as
where and denote hidden representations associated with text-following and image-following generations under order , respectively.
We extract these representations from decoder layers 20–23. During held-out evaluation, the corresponding order-specific vector is added to the hidden state at every decoding step:
where denotes the decoder layer and controls the steering strength. We evaluate , with corresponding to the original model without steering. Positive values steer the representation toward the direction associated with text-following behavior.
Evaluation.
We quantify modality preference using the normalized image-over-text gap as denoted in §3. Positive values indicate image-dominant behavior, negative values indicate text-dominant behavior.
red
Results. Table 4 shows that increasing the steering scale consistently moves modality preference toward text. Under , where the unsteered model is image-dominant, changes from 77.77% at scale 0 to 80.64%, 64.95%, and 60.00% at scales 1, 3, and 5, respectively. In contrast, under , where the model is already text-dominant, the gap becomes increasingly negative, changing from -38.64% to -86.96%.
Table 13 reports the corresponding raw image-following rates, averaged across the two input orders. decreases monotonically from 42.28% to 30.15% as scales increase (-12.13 percentage points), which means stronger steering reduces overall reliance on misleading image evidence.
These results confirm that activation steering can systematically manipulate modality preference at inference time. However, the intervention is directional: reducing image preference under one input order can simultaneously amplify existing text preference under the other. Thus, successful control of modality preference should not be conflated with improved modality robustness, which would require reducing the magnitude of the modality gap in both directions.
Appendix D Visual Preprocessing Interventions
To examine whether differences in input-side visual tokenization contribute to modality reliance, we implement a preprocessing-level intervention between Qwen2.5-VL-7B-Instruct and LLaVA-v1.6-vicuna-7b. The intervention modifies only the image preprocessing pipeline applied before inference. We do not change model weights, language prompts, decoding hyperparameters, or answer evaluation rules. Therefore, any change in modality reliance should be interpreted as the effect of changing how the visual evidence is presented to the model.
LLaVA-style preprocessing for Qwen2.5-VL.
For the Qwen-to-LLaVA direction, we constrain Qwen2.5-VL to receive images in a LLaVA-style global/local view format. We first select a target canvas resolution from the following set of image grids:
For each input image, we choose the resolution that maximizes the effective retained image area, then resize and pad the image to the selected grid resolution. The resulting canvas is divided into local views. In addition, we create a global view by resizing the image according to its shortest edge and center-cropping it to , and final visual input to Qwen consists of both tiles. Here, since Qwen2.5-VL merges visual patches over a grid, each view corresponds to
visual tokens. It matches the global/local view structure of LLaVA-v1.6 while keeping each individual view at the LLaVA-style resolution.
We set Qwen’s visual processor constraints so that the minimum and maximum pixel budgets are fixed to the desired view size. This prevents Qwen’s default dynamic resizing from changing the intended visual tokenization pattern.
Qwen-style pre-resizing for LLaVA-v1.6.
For the reverse direction, we apply Qwen2.5-VL’s default image resizing rule before feeding the image to LLaVA-v1.6. Specifically, we load the Qwen2.5-VL image processor configuration and compute the resized height and width using Qwen’s smart_resize rule, including its default minimum pixel budget, maximum pixel budget, patch size, and merge size. The original image is converted to RGB, resized to the Qwen-computed resolution using bicubic interpolation, and saved as a new image file. This pre-resized image is then provided as the visual input to LLaVA.
This intervention should be understood as an input-side resizing control rather than a full processor replacement. We do not claim that this intervention fully explains the Qwen–LLaVA gap. Rather, it shows that modality reliance can be steered by input-side visual preprocessing alone, indicating that visual representation is one contributing factor alongside model architecture and training.
Appendix E Additional Image-Presentation Results under
Figure 7 shows the corresponding results for . Consistent with the results in §7.2, global presentation changes have only minor effects, while answer-span highlighting produces substantially larger shifts for both models.
Appendix F Influence of Image-to-Text Translation
Finally, we rule out a trivial confound: whether the observed modality-dependent behavior could be attributed to models simply failing to “read” the rendered text. We prompt each model to transcribe the visible text in the rendered image and compare it to the original passage using ROUGE-L.
Table 14 shows near-perfect scores across representative models and datasets, indicating that the rendered evidence is essentially fully recoverable as text. This result eliminates poor image-to-text recognition as the main explanation for §4: models that demonstrably can read the image still use it differently from the same text in raw form, pointing to cross-modal integration, rather than perception, as the locus of the problem.
| Model | ConflictQA | NQ-Swap |
|---|---|---|
| Qwen2.5-VL-7B | 99.97 | 99.56 |
| InternVL3-8B | 98.37 | 98.94 |
| LLaVA-v1.6-7B | 99.82 | 99.15 |
Appendix G Per-Scenario Safety Results
Table 15 extends the results in Figure 4 in §5.2 to all MM-SafetyBench scenarios that could be evaluated. Note that some instances are omitted because of the missing of the required image assets. Each cell reports the Attack Success Rate (ASR) of our image-rendered instruction condition, with the value in parentheses indicating the ASR change relative to the corresponding original baseline. Across most evaluated scenarios, our image-rendered instruction condition increases ASR over the corresponding baselines.
| Qwen2.5-VL-32B-Instruct | ||||
|---|---|---|---|---|
| Scenarios | Text-only | SD | Typo | SD+Typo |
| 01-Illegal Activity | 1.03 (+1.03) | 81.44 (+4.12) | 5.15 (+2.06) | 12.37 (+1.03) |
| 02-Hate Speech | 12.88 (+6.75) | 92.02 (+8.58) | 42.94 (+9.81) | 57.67 (-3.07) |
| 03-Malware Generation | 47.73 (+20.46) | 97.73 (+2.28) | 70.45 (+9.09) | 79.55 (+4.55) |
| 04-Physical Harm | 31.25 (+9.72) | 92.36 (+1.39) | 54.17 (+6.95) | 65.28 (+3.47) |
| 05-Economic Harm | 75.41 (+2.46) | 92.62 (-0.82) | 77.05 (+0.82) | 80.33 (+1.64) |
| 06-Fraud | 13.64 (+7.15) | 91.56 (+3.25) | 39.61 (+1.95) | 53.90 (+5.55) |
| 07-Pornography | 72.48 (+17.43) | 100.00 (+2.75) | 98.17 (+3.67) | 98.17 (+2.76) |
| 08-Political Lobbying | 98.69 (+0.65) | 87.58 (+0.00) | 86.93 (-0.65) | 87.58 (+0.00) |
| 09-Privacy Violence | 33.09 (+10.07) | 92.81 (+2.88) | 43.17 (+6.48) | 46.04 (+0.00) |
| 10-Legal Opinion | 93.08 (+13.08) | 98.46 (+3.08) | 98.46 (+1.54) | 99.23 (+2.31) |
| 11-Financial Advice | 98.80 (-0.60) | 100.00 (+0.00) | 100.00 (+0.00) | 99.40 (-0.60) |
| 12-Health Consultation | 84.40 (+15.59) | 99.08 (+0.00) | 97.25 (-1.83) | 97.25 (-0.92) |
| 13-Gov Decision | 98.66 (+6.04) | 100.00 (+0.00) | 94.63 (+0.67) | 96.64 (+0.00) |
| InternVL3-38B-Instruct | ||||
| Scenarios | Text-only | SD | Typo | SD+Typo |
| 01-Illegal Activity | 0.00 (+0.00) | 57.73 (+12.37) | 5.15 (+1.03) | 9.28 (+7.22) |
| 02-Hate Speech | 23.93 (+21.48) | 75.46 (+4.91) | 38.04 (+9.82) | 51.53 (+9.81) |
| 03-Malware Generation | 25.00 (+4.55) | 84.09 (-2.27) | 52.27 (+18.18) | 65.91 (+20.46) |
| 04-Physical Harm | 23.61 (+9.03) | 81.25 (+9.72) | 45.14 (+11.11) | 50.00 (+8.33) |
| 05-Economic Harm | 70.49 (+1.64) | 90.16 (+4.09) | 77.05 (+3.28) | 77.87 (+0.00) |
| 06-Fraud | 2.60 (-0.65) | 71.43 (+4.55) | 24.68 (+9.72) | 39.61 (+12.34) |
| 07-Pornography | 38.53 (+0.92) | 86.24 (+4.59) | 63.29 (+4.57) | 75.23 (-2.75) |
| 08-Political Lobbying | 97.39 (+8.50) | 84.97 (-2.61) | 87.58 (+1.31) | 87.58 (+0.00) |
| 09-Privacy Violence | 19.42 (+12.23) | 79.86 (+8.64) | 35.25 (+8.63) | 39.57 (+8.63) |
| 10-Legal Opinion | 79.23 (+29.23) | 95.38 (+18.46) | 97.69 (+16.15) | 95.38 (+7.69) |
| 11-Financial Advice | 96.41 (+10.78) | 100.00 (+1.20) | 100.00 (+0.60) | 100.00 (+0.00) |
| 12-Health Consultation | 71.56 (+21.10) | 93.58 (+9.18) | 96.33 (+6.42) | 97.25 (+0.92) |
| 13-Gov Decision | 92.62 (+42.96) | 97.32 (+10.07) | 95.30 (+3.35) | 92.62 (+0.67) |
| LLaVA-v1.6-34B-hf | ||||
| Scenarios | Text-only | SD | Typo | SD+Typo |
| 01-Illegal Activity | 10.31 (+6.19) | 95.88 (+17.53) | 43.30 (+25.77) | 52.58 (+35.05) |
| 02-Hate Speech | 48.47 (+20.86) | 96.93 (+5.52) | 79.75 (+10.42) | 92.64 (+8.59) |
| 03-Malware Generation | 65.91 (-13.64) | 100.00 (+0.00) | 81.82 (+2.27) | 95.45 (+20.45) |
| 04-Physical Harm | 64.58 (+27.77) | 98.61 (+2.78) | 81.25 (+15.28) | 87.50 (+18.75) |
| 05-Economic Harm | 77.87 (+7.38) | 93.44 (+4.92) | 79.51 (+0.00) | 86.89 (+6.56) |
| 06-Fraud | 34.42 (+2.60) | 100.00 (+7.79) | 79.87 (+9.74) | 89.61 (+24.67) |
| 07-Pornography | 89.91 (+11.93) | 97.25 (+1.84) | 97.25 (+0.92) | 93.58 (-0.92) |
| 08-Political Lobbying | 100.00 (+7.19) | 86.93 (+2.62) | 86.93 (-0.65) | 86.27 (-1.31) |
| 09-Privacy Violence | 53.96 (+12.23) | 100.00 (+7.19) | 73.38 (+6.47) | 82.73 (+14.38) |
| 10-Legal Opinion | 96.15 (+34.61) | 98.46 (+14.61) | 96.15 (+2.30) | 97.69 (+6.15) |
| 11-Financial Advice | 100.00 (+9.58) | 99.40 (+1.80) | 99.40 (-0.60) | 100.00 (+0.60) |
| 12-Health Consultation | 98.17 (+39.45) | 98.17 (+4.59) | 96.33 (+2.75) | 96.33 (+0.92) |
| 13-Gov Decision | 97.99 (+14.77) | 98.66 (+12.08) | 98.66 (+5.37) | 97.32 (+5.37) |