arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.00796v2 [cs.CL] 03 Sep 2026

SFAD: Speculative Factuality-Aware Decoding

Guanqiao Chen Affiliation: MBZUAI Affiliation: University of Science and Technology of China    Di Wang Affiliation: Provable Responsible AI and Data Analytics (PRADA) Lab Affiliation: King Abdullah University of Science and Technology*Corresponding author.    Lijie Hu Affiliation: MBZUAI
Abstract

As one of the most critical challenges in large language models, contextual faithfulness directly determines their reliability in knowledge-intensive applications. This task is particularly challenging as it requires balancing factual consistency with generation efficiency. Contrastive decoding methods require dual forward passes (with and without context) to compare model outputs, doubling inference computational overhead, while post-training alignment demands extensive reinforcement learning with substantial computational overhead. To address this challenge, we present SFAD, a speculative decoding framework that enhances contextual faithfulness without inference degradation. We first construct ConFide, a preference dataset with fine-grained atomic perturbations, to train a context-faithful draft model via Direct Preference Optimization. During inference, Epistemic Friction detects potential hallucinations by quantifying distributional tension weighted by specialist certainty. When friction exceeds the threshold, Asymmetric Logit Steering refines the target distribution through residual-based logit injection; otherwise, standard speculation proceeds. Extensive experiments demonstrate that SFAD substantially improves faithfulness while achieving 2.48×2.48\times speedup, offering a practical solution for efficient LLMs.

1 Introduction

Refer to caption
Figure 1: The ConFide pipeline for factuality-aware preference data construction.

Large Language Models (LLMs) Guo et al. (2025a); Achiam et al. (2023) have achieved remarkable success by leveraging vast parametric knowledge acquired during pretraining Petroni et al. (2019); Roberts et al. (2020). However, this internal knowledge is inherently static and constrained by the training distribution, rendering it prone to becoming outdated or incomplete Cheng et al. (2024b); Li et al. (2026b); Zhang et al. (2025a); Cheng et al. (2024a). To address this limitation, retrieval-augmented generation (RAG) Qin et al. (2025); Guu et al. (2020); Cheng et al. (2025a) and external tool integration have become increasingly prevalent, intensifying the demand for models to prioritize provided context over parametric priors. Yet LLMs frequently exhibit hallucinations Huang et al. (2025a); Ji et al. (2023); Cheng et al. (2025b) due to knowledge conflicts, where models favor internal knowledge over task-specific external contexts Xie et al. (2024); Chen et al. (2022); Zhang et al. (2025b), undermining their contextual faithfulness in knowledge-intensive applications.

Refer to caption
Figure 2: SFAD framework.

Enhancing contextual faithfulness while maintaining inference efficiency presents significant challenges. Decoding-based methods Shi et al. (2024); Xu (2023); Khandelwal et al. (2025); Yang et al. (2025c) contrast logits from contextualized inputs against uncontextualized baselines to amplify evidence-aligned signals. However, this requires dual forward, effectively doubling computational cost and halving generation speed. Post-training methods Yan et al. (2025); Wang et al. (2025c); Yang et al. (2025b); Zhou et al. (2026) employ reinforcement learning to align models with contextual evidence, but typically require extensive compute and large-scale preference data. These limitations motivate the need for approaches that enhance faithfulness without compromising efficiency or requiring substantial resources.

To resolve these problems, we propose SFAD, a speculative decoding framework that unifies contextual faithfulness enhancement with speculative acceleration. Our key insight is to leverage a context-faithful draft model as a factuality sentinel: by training a smaller model to prioritize contextual evidence, we enable it to detect and correct hallucinations in the larger target model during speculative decoding. Specifically, we first construct ConFide, a preference dataset that improves contextual faithfulness by generating diverse hard-negative samples through atomic decomposition and controllable perturbations. Training the draft model on ConFide and ConFiQA Bi et al. (2025) via Direct Preference Optimization instills strong contextual loyalty. During inference, SFAD introduces Epistemic Friction, a metric that quantifies distributional tension weighted by specialist certainty to dynamically detect knowledge conflicts. When friction signals potential hallucinations, Asymmetric Logit Steering selectively refines the target distribution through residual-based injection; otherwise, standard speculative decoding proceeds unmodified. This adaptive mechanism maintains the 2∼3×2\sim 3\times speedup of speculative decoding while substantially improving contextual faithfulness.

Our contributions are threefold: (1) Dataset: we introduce ConFide, a fine-grained preference dataset designed for training context-faithful draft models through atomic-level perturbations; (2) Framework: we propose SFAD, the first speculative decoding framework specifically designed for hallucination mitigation, which features Epistemic Friction for conflict detection and Asymmetric Logit Steering for selective correction; and (3) Evaluation: extensive experiments demonstrate that SFAD achieves substantial faithfulness improvements across diverse benchmarks while delivering a 2.48×2.48\times speedup, approaching the performance of models 5×5\times larger with minimal computational overhead.

2 Related Work

2.1 Hallucinations in LLMs

HallucinationsHuang et al. (2025b) occur when LLM outputs appear plausible yet deviate from factual or contextual knowledge Kaddour et al. (2023); Ji et al. (2023). These are typically bifurcated into factuality hallucinations, which contradict real-world facts based on internal parametric knowledge Min et al. (2023); Wei et al. (2024), and faithfulness hallucinations, characterized by inconsistency with the provided input or grounding documents Wang et al. (2026); Wan et al. (2023); Hu et al. (2025); Guo et al. (2025b). Mitigation strategies span the model lifecycle: training-phase efforts focus on data curation and grounding Pan et al. (2024), while inference-stage methods employ confidence estimation Huang et al. (2025c), knowledge retrieval Feng et al. (2024), and editing Ali et al. (2025); Wang et al. (2025a). Despite these advancements, LLMs remain prone to inconsistent outputs in context-rich tasks such as RAG and summarization Xu et al. (2024); Li and Yu (2025). Our work leverages the Speculative Decoding framework to ensure context-consistent generation.

2.2 Speculative Decoding

Speculative decoding accelerates LLM inference by verifying draft tokens generated by a smaller model against the target LLM Leviathan et al. (2023); Chen et al. (2023). Drafting strategies have evolved from independent models Xia et al. (2023) to architectural extensions such as recurrent feature utilization Li et al. (2024) and diffusion-based generation Li et al. (2026a). Efficiency has been further optimized through tree-based parallel verification Miao et al. (2024), sparse Mixture-of-Experts (MoE) acceleration Huang et al. (2025d), and specialized reinforcement learning system enhancements Chen et al. (2026). Recent research has also tailored SD for diverse and complex scenarios, including long-context retrieval-augmentation Sun et al. (2024); Sadhukhan et al. (2025); Yang et al. (2026); Chen et al. (2025), multi-sample inference tasks Li et al. (2025), and memory-efficient inference via quantized caching Tiwari et al. (2025). Beyond performance, safety-awareness has been integrated into the speculative framework to ensure alignment Wang et al. (2025d). Unlike these performance-driven or safety-centric works, SFAD is the first speculative framework to target hallucination mitigation.

3 Aligning Draft Models for Context-Faithfulness via ConFide

To enhance the draft model’s generalization across diverse hallucination types, we construct ConFide, a fine-grained preference dataset that systematically diversifies error patterns in negative samples and stylistic variations in positive samples. As illustrated in Figure 1, our pipeline leverages atomic fact decomposition and controllable perturbations to produce high-quality contrastive pairs.

3.1 Data Construction Pipeline

Source Data and Problem Formulation. We build upon the LLM-AggreFact Tang et al. (2024a) and CG2C Lei et al. (2025) datasets, which provide fact-intensive samples covering diverse hallucination patterns suitable for atomic decomposition. Our initial dataset is defined as 𝒟={(x(i),yref(i),l(i))}i=1N\mathcal{D}=\{(x^{(i)},y_{\text{ref}}^{(i)},l^{(i)})\}_{i=1}^{N}, where xx is the source context, yrefy_{\text{ref}} is the candidate response, and l∈{0,1}l\in\{0,1\} is the binary faithfulness label. Our goal is to construct a preference dataset 𝒫={(x,yw,yl)}\mathcal{P}=\{(x,y_{w},y_{l})\} for DPO training, where ywy_{w} is faithful to context xx and yly_{l} contains specific hallucinations.

Atomic Fact Decomposition. We decompose each response into verifiable atomic facts via a decomposition function fdecf_{\text{dec}}:

𝒜=fdec​(y)={a1,a2,…,am},\mathcal{A}=f_{\text{dec}}(y)=\{a_{1},a_{2},\dots,a_{m}\}, (1)

where each atomic fact aja_{j} represents a minimal verifiable claim.

Negative Sample Generation. We apply a Controllable Perturbation Mechanism to transform atomic facts into corrupted versions. Given aj=(sj,rj,oj)a_{j}=(s_{j},r_{j},o_{j}) as a subject-relation-object triple, we define three perturbation operators:

ϕ⁡(aj)={(sj,rj,oj′),τent(sj,rj,oj+ϵ),τnum(sj,¬rj,oj),τrel\phi(a_{j})=\begin{cases}(s_{j},r_{j},o^{\prime}_{j}),&\tau_{\text{ent}}\\ (s_{j},r_{j},o_{j}+\epsilon),&\tau_{\text{num}}\\ (s_{j},\neg r_{j},o_{j}),&\tau_{\text{rel}}\end{cases} (2)

where τent\tau_{\text{ent}}, τnum\tau_{\text{num}}, and τrel\tau_{\text{rel}} denote entity swap, numerical distortion, and relation inversion, respectively. Here, oj′o^{\prime}_{j} is a confusing entity drawn from context xx, ϵ\epsilon is calibrated noise preserving numerical plausibility, and ¬rj\neg r_{j} denotes logical negation of relation rjr_{j}. The perturbed facts are reconstructed into a fluent response yl=frec​(𝒜′)y_{l}=f_{\text{rec}}(\mathcal{A}^{\prime}), yielding hallucinated yet linguistically natural negative samples.

Positive Sample Refinement. Winning samples ywy_{w} are constructed based on the faithfulness label ll. If l=1l=1, we apply paraphrasing πpara\pi_{\text{para}} to generate semantically equivalent but syntactically varied sentences. If l=0l=0, a teacher model (GPT-4o Hurst et al. (2024)) corrects yrefy_{\text{ref}} against context xx to produce a strictly faithful ywy_{w}. This ensures winning samples reflect genuine factual alignment rather than surface-level patterns.

Preference Dataset Construction. Pairing the perturbed yly_{l} with the refined ywy_{w} yields the final preference dataset 𝒫\mathcal{P}, where negative samples exhibit fine-grained, fluent hallucinations and positive samples maintain strict contextual fidelity.

3.2 Draft Model Optimization via DPO

We optimize the draft model πθ\pi_{\theta} via the DPO objective (Rafailov et al., 2023), which directly maximizes the preference margin between faithful responses ywy_{w} and hallucinated responses yly_{l} without requiring an explicit reward model. The gradient signal encourages πθ\pi_{\theta} to concentrate probability mass on tokens strictly verifiable against the context xx, controlled by hyperparameter β\beta relative to the frozen reference model πref\pi_{\text{ref}}.

4 SFAD: Speculative Factuality-Aware Decoding

We formulate Speculative Factuality-Aware Decoding (SFAD) as a dynamic inference framework designed to decouple token verification from distribution correction. Figure 2 illustrates the overall pipeline. Let MM and mm denote the target (generalist) and draft (specialist) models with parameters θM\theta_{M} and θm\theta_{m}, respectively. At each time step tt, given the prefix context x<tx_{<t}, the models produce logit vectors 𝐳M,t,𝐳m,t∈ℝ|𝒱|\mathbf{z}_{M,t},\mathbf{z}_{m,t}\in\mathbb{R}^{|\mathcal{V}|}. Our framework is governed by a five-step formalism that adaptively steers the generation process when the specialist detects potential factual inconsistencies. Theoretical analysis and proofs are provided in Appendix E.

4.1 Specialist Certainty Measure

To prevent noise injection from an uncertain draft model, we first quantify its internal confidence. Unlike standard entropy measures, we require a normalized metric that penalizes high-entropy (uncertain) distributions. We define the Specialist Certainty κt\kappa_{t} as:

κt=(1−ℍ(ℙm(⋅|x<t))log⁡|𝒱|)γ,\displaystyle\kappa_{t}=\left(1-\frac{\mathbb{H}(\mathbb{P}_{m}(\cdot|x_{<t}))}{\log|\mathcal{V}|}\right)^{\gamma}, (3)
where ​ℙm=Softmax​(𝐳m,t).\displaystyle\text{where }\mathbb{P}_{m}=\text{Softmax}(\mathbf{z}_{m,t}).

where ℍ⁡(⋅)\mathbb{H}(\cdot) denotes the Shannon entropy, |𝒱||\mathcal{V}| is the vocabulary size, and γ≥1\gamma\geq 1 is a sharpening coefficient. Here, κt→1\kappa_{t}\to 1 implies the specialist is highly confident in its prediction, serving as a necessary precondition for intervention.

4.2 Epistemic Friction Coefficient

To distinguish factual conflicts from benign linguistic diversity, we propose the Epistemic Friction ℱt\mathcal{F}_{t}, which captures the distributional tension between the generalist and the specialist. It is formulated as the Jensen-Shannon (JS) divergence weighted by the specialist’s certainty:

ℱt=𝒟JS(ℙM(⋅|x<t)∥ℙm(⋅|x<t))⋅κt\mathcal{F}_{t}=\mathcal{D}_{\text{JS}}\left(\mathbb{P}_{M}(\cdot|x_{<t})\,\|\,\mathbb{P}_{m}(\cdot|x_{<t})\right)\cdot\kappa_{t} (4)

ℱt\mathcal{F}_{t} acts as a detector for confident hallucinations. High friction occurs if and only if the models disagree significantly (𝒟JS\mathcal{D}_{\text{JS}} is high) and the specialist is factually convinced (κt\kappa_{t} is high).

4.3 Adaptive Gating Mechanism

To ensure computational efficiency and avoid over-correction, SFAD employs a soft gating scalar λt∈[0,1]\lambda_{t}\in[0,1] to govern the intensity of the steering. It is derived via a shifted sigmoid activation:

λt=σ⁡(β⋅(ℱt−τ))=11+exp⁡(−β⁡(ℱt−τ))\lambda_{t}=\sigma\left(\beta\cdot(\mathcal{F}_{t}-\tau)\right)=\frac{1}{1+\exp\left(-\beta(\mathcal{F}_{t}-\tau)\right)} (5)

where τ\tau is the friction threshold and β\beta controls the transition sharpness. This creates a switching regime: when ℱt≪τ\mathcal{F}_{t}\ll\tau, λt→0\lambda_{t}\to 0 (Standard Speculation); when ℱt≫τ\mathcal{F}_{t}\gg\tau, λt→1\lambda_{t}\to 1 (Steering Mode).

4.4 Asymmetric Steering with Contextual Plausibility Masking

When intervention is triggered, we modify the target logits. To ensure that the specialist only intervenes when its predictions are linguistically plausible within the current sequence, we introduce a Contextual Plausibility Mask (CPM). We define the plausible set 𝒱C​P​C\mathcal{V}_{CPC} as:

𝒱C​P​C={v∈𝒱:ℙm​(v|x<t)≥η​ptmax}\mathcal{V}_{CPC}=\left\{v\in\mathcal{V}:\mathbb{P}_{m}(v|x_{<t})\geq\eta p_{t}^{\max}\right\} (6)

where ptmaxp_{t}^{\max} is the maximum probability assigned by ℙm(⋅|x<t)\mathbb{P}_{m}(\cdot|x_{<t}), and η∈(0,1)\eta\in(0,1) is a plausibility threshold.

Inspired by previous work Bachmann et al. (2025), which demonstrated that alignment-based verification in speculative decoding rejects many high-quality tokens due to distribution mismatches, we extend this insight beyond token-level acceptance. Instead of solely adapting verification schemes, we intervene at the logit level to inject factuality-aware corrections, ensuring both efficiency and reduced hallucinations.The corrected logit vector 𝐳t∗\mathbf{z}^{*}_{t} is computed as:

𝐳t∗=𝐳M,t+λt⋅ReLU​(𝐳m,t−𝐳M,t)⋅𝕀⁡(x∈𝒱C​P​C)\mathbf{z}^{*}_{t}=\mathbf{z}_{M,t}+\lambda_{t}\cdot\text{ReLU}(\mathbf{z}_{m,t}-\mathbf{z}_{M,t})\cdot\mathbb{I}(x\in\mathcal{V}_{CPC}) (7)

The ReLU​(⋅)\text{ReLU}(\cdot) operator enforces unidirectional knowledge injection, allowing the specialist to boost correct entities without penalizing the generalist’s linguistic fluency. The 𝕀⁡(x∈𝒱C​P​C)\mathbb{I}(x\in\mathcal{V}_{CPC}) term acts as a safety guard, ensuring the specialist only steers the target toward tokens that maintain "compositional plausibility."

4.5 Hybrid Decoding Policy

Finally, the next token xtx_{t} is determined by a hybrid policy that switches between the corrected distribution and standard speculative verification:

xt∼{Softmax​(𝐳t∗)if ​ℱt≥τ(Steering Path)Verify​(x~t,ℙM)if ​ℱt<τ(Fast Path)x_{t}\sim\begin{cases}\text{Softmax}(\mathbf{z}^{*}_{t})&\text{if }\mathcal{F}_{t}\geq\tau\quad\text{(Steering Path)}\\ \text{Verify}(\tilde{x}_{t},\mathbb{P}_{M})&\text{if }\mathcal{F}_{t}<\tau\quad\text{(Fast Path)}\end{cases} (8)

In the Steering Path, we sample from the steered logits, effectively overriding the draft. In the Fast Path, we perform standard rejection sampling using the draft token x~t\tilde{x}_{t} and the original target distribution ℙM\mathbb{P}_{M}, maintaining the speed guarantees of speculative decoding.

5 Experiments

5.1 Experimental Setup

Models and Configurations. We implement the SFAD framework using the Qwen3 family Yang et al. (2025a). Specifically, Qwen3-1.7B is utilized as the draft model, which is fine-tuned via Direct Preference Optimization to serve as the expert model.The target model is Qwen3-14B.

Data Construction for DPO. The DPO training set for the draft model consists of approximately 36K curated samples, combining two complementary sources. First, we directly adopt 18K instances from ConFiQA Bi et al. (2025), a benchmark focusing on multi-hop knowledge conflicts. Second, we construct our ConFide dataset using the atomic perturbation pipeline described in Section 3. Specifically, ConFide is synthesized from two source datasets: 12K samples from LLM-AggreFact Tang et al. (2024a), and 6K samples from the CG2C dataset Lei et al. (2025).

Baselines. We compare SFAD against two categories of baselines to demonstrate its superiority in both contextual faithfulness and inference efficiency. The first category consists of decoding-level methods applied to our target model, Qwen3-14B: Greedy Decoding, Context-Aware Decoding (CAD) Xu (2023), COIECD Yuan et al. (2024), and AdaCAD Wang et al. (2025b). The second category includes a frontier model, Llama-3.1-70B-Instruct MetaAI (2024), which has 5×5\times the parameters of Qwen3-14B, serving as a high-performance reference point.

Evaluation Tasks and Datasets. We evaluate across five task categories: (1) Factual Retrieval on HotpotQA Yang et al. (2018), PopQA Mallen et al. (2023), and TriviaQA Joshi et al. (2017); (2) Abstractive Faithfulness via summarization on TofuEval Tang et al. (2024b) and XSum Narayan et al. (2018); (3) Extended Generation on CLAPNQ Rosenthal et al. (2025), ExpertQA Malaviya et al. (2024), and HAGRID Kamalloo et al. (2023); (4) Knowledge Conflicts using 200 held-out instances from LLM-AggreFact; and (5) General Capabilities on GSM8K Cobbe et al. (2021) and Just-Eval Lin et al. (2024).

Quality and Efficiency Metrics. For generative quality, we use standard metrics tailored to each task: Exact Match (EM) for retrieval tasks; AlignScore, BERT-P, and ROUGE-L for summarization; and FaithScore (computed via MiniCheck Tang et al. (2024a)) for long-form QA. For knowledge conflict analysis, we adopt ConFiQA Bi et al. (2025) metrics: Context-faithful Frequency (PcP_{c}), Original Factual Frequency (PoP_{o}), and Memory Reliance (MRM_{R}). To quantify the inference acceleration of SFAD, we define the Average Token Generation Acceleration (ATGA), measured in an end-to-end, context-aware setting:

ATGA=Avg. token generation time w/o SFADAvg. token generation time w/ SFAD\text{ATGA}=\frac{\text{Avg. token generation time w/o SFAD}}{\text{Avg. token generation time w/ SFAD}} (9)

5.2 Main Results

We present comprehensive evaluations across three categories of context-intensive tasks: Foundation QA, Summarization, and Long-Form QA.

SFAD demonstrates consistent metric improvements across all evaluated benchmarks, optimizing both task accuracy and inference efficiency. While standard decoding-level interventions typically incur latency penalties due to redundant forward passes, SFAD integrates a speculative decoding framework that streamlines the generation process. This efficiency is driven by our design’s focus on enhancing contextual consistency, which allows the model to leverage evidentiary alignment to accelerate token production. In Foundation QA, the method mitigates the internal knowledge limitations of the base model, a capability that extends to long-form generation where SFAD addresses the inherent trade-off between lexical fluency and factual grounding. By prioritizing context-aligned tokens, SFAD substantially narrows the performance gap between the 14B base model and the 5×5\times larger Llama-3.1-70B frontier model. These results underscore the capacity of SFAD to maintain sustained evidentiary grounding and long-range coherence with reduced computational costs, avoiding the need for task-specific fine-tuning or increased parameter overhead.

Table 1: Performance comparison on Foundation QA. Rel. Latency denotes the inference time relative to the Qwen3-14B vanilla baseline.
Method TriviaQA HotpotQA PopQA Rel. Latency
Vanilla Baseline
Qwen3-14B 53.87 41.77 78.21 1.00×\times
Decoding-level Baselines (Qwen3-14B)
CAD 41.43 39.51 71.29 2.00×\times
AdaCAD 82.11 45.63 77.39 2.15×\times
COIECD 83.07 45.63 76.29 2.40×\times
Frontier Model
Llama-3.1-70B 90.20 56.11 86.11 4.85×\times
SFAD (Ours) 85.12 52.19 86.39 0.82×\times
Table 2: Summarization performance on XSum and TofuEval. SFAD achieves superior factuality while reducing inference overhead.
XSum TofuEval
Method R-L BERT-P AlignScore AlignScore Rel. Latency
Vanilla Baseline
Qwen3-14B 13.67 91.67 72.86 59.84 1.00×\times
Decoding-level Baselines (Qwen3-14B)
CAD 14.59 93.65 84.34 83.23 2.00×\times
AdaCAD 14.91 94.29 85.81 85.07 2.20×\times
COIECD 13.65 91.04 73.81 60.86 2.45×\times
Frontier Model
Llama-3.1-70B 16.35 94.12 87.48 87.31 5.10×\times
SFAD (Ours) 16.32 93.97 87.37 87.53 0.85×\times
Table 3: Long-Form QA results. SFAD significantly enhances faithfulness with higher efficiency than the vanilla model.
CLAPNQ ExpertQA HAGRID
Method R-L Faith R-L Faith R-L Faith Rel. Latency
Vanilla Baseline
Qwen3-14B 17.12 59.73 31.56 51.29 16.96 57.63 1.00×\times
Decoding-level Baselines (Qwen3-14B)
CAD 18.23 60.24 33.58 53.47 17.89 58.20 2.00×\times
AdaCAD 18.43 62.37 32.87 54.76 17.12 59.90 2.18×\times
COIECD 18.56 61.96 33.14 56.32 17.34 59.76 2.42×\times
Frontier Model
Llama-3.1-70B 42.15 92.45 46.10 72.40 52.07 82.20 5.30×\times
SFAD (Ours) 41.34 90.93 43.79 71.13 51.97 81.99 0.78×\times

6 Analysis

6.1 Effectiveness of ConFide and DPO

To isolate the contributions of our data construction pipeline and alignment strategy, we evaluate three draft model variants on 200 held-out knowledge-conflict instances from LLM-AggreFact. As illustrated in Figure 3, the transition from supervised fine-tuning (SFT) to preference optimization (DPO) triggers a fundamental shift in model behavior: while ConFide+SFT remains heavily anchored to internal priors (high MRM_{R}), DPO-trained variants demonstrate a superior ability to suppress parametric bias in favor of contextual evidence.

Notably, ConFide+DPO consistently outperforms ConFiQA+DPO across all metrics, validating the efficacy of our atomic perturbation mechanism. By exposing the model to hard negatives—such as entity swaps and relation inversions—ConFide provides a more granular discriminative signal that sharpens the model’s sensitivity to factual nuances. This contrastive optimization effectively penalizes hallucinated completions that appear superficially plausible but lack contextual grounding. These results confirm that our data construction pipeline combined with preference optimization is crucial for training context-faithful draft models capable of effective speculative decoding.

Figure 3: Impact of training strategies on knowledge conflict resolution. ConFide+DPO consistently outperforms baselines across context faithfulness (PcP_{c}, E​MEM) and memory reliance (MRM_{R}, PoP_{o}) metrics.
Figure 4: Latency-faithfulness trade-off analysis. SFAD achieves Pareto-optimal performance, simultaneously delivering 2.48×\times speedup and 85.2 faithfulness score. Decoding-level baselines suffer from severe latency penalties (<<0.85×\times) despite moderate faithfulness, while standard SD prioritizes speed (2.18×\times) at the cost of poor faithfulness (38.5). SFAD outperforms the best baseline by 21.7 points while maintaining 2.9×\times faster inference.

6.2 Latency-Faithfulness Trade-off Analysis

A critical question in deploying context-faithful LLMs is whether achieving high factuality necessitates sacrificing inference speed. To address this, we analyze the trade-off between inference acceleration and contextual faithfulness across different decoding strategies. We evaluate all methods on the same 200-instance LLM-AggreFact test set used in Section 6.1, measuring both Average Token Generation Acceleration (ATGA, Eq. (9)) and Faithfulness Score (MiniCheck FaithScore). Figure 4 visualizes this relationship, revealing three distinct performance regimes.

Decoding-level baselines exhibit poor speed-faithfulness balance. CAD, AdaCAD, and COIECD achieve moderate faithfulness scores but suffer severe latency penalties (ATGA 0.52×\times–0.82×\times). They need to perform two forward passes for both contextualized and baseline inputs. This introduces additional token overhead, which slows down the model’s inference speed.

Standard speculative decoding prioritizes speed over faithfulness. While achieving 2.18×\times speedup, vanilla SD with an unaligned draft model produces drastically poor faithfulness, falling below even the decoding baselines. The draft model’s lack of contextual grounding causes it to propose tokens based on parametric priors rather than provided evidence, leading to significant hallucinations in knowledge-intensive scenarios.

SFAD achieves Pareto-optimal performance. SFAD strikes an optimal balance between contextual faithfulness and inference speed, simultaneously delivering 2.48×\times speedup and 85.2 faithfulness score.

6.3 Epistemic Friction Analysis: Detecting Confident Hallucinations

To validate our Epistemic Friction mechanism, we analyze 1,200 samples from MQUAKE Zhong et al. (2023), tracking DJ​SD_{JS}, κt\kappa_{t}, and ℱt\mathcal{F}_{t} across generation. Figure 5 illustrates metric dynamics across a representative generation sequence. The raw divergence DJ​SD_{JS} fluctuates frequently, with peaks exceeding τ=0.5\tau=0.5, but relying solely on DJ​SD_{JS} would trigger excessive false positives at stylistic variations.

Figure 5: Epistemic Friction dynamics. Purple circles mark positions where DJ​S>τD_{JS}>\tau is suppressed by low κt\kappa_{t}. The red square marks a confident hallucination that triggers steering.

The specialist certainty κt\kappa_{t} provides crucial filtering. At positions marked with purple circles, DJ​SD_{JS} slightly exceeds τ\tau (0.52–0.58) but κt\kappa_{t} remains low (0.28–0.32), indicating uncertainty at linguistic alternatives. Consequently, ℱt=DJ​S×κt\mathcal{F}_{t}=D_{JS}\times\kappa_{t} is suppressed below threshold (0.15–0.17), correctly avoiding intervention.

In contrast, the red-square position exhibits a confident hallucination: DJ​S≈0.89D_{JS}\approx 0.89 combined with κt≈0.97\kappa_{t}\approx 0.97 produces a sharp spike in ℱt\mathcal{F}_{t} (∼\sim0.86), triggering full logit steering. By conditioning on both distributional conflict and specialist confidence, Epistemic Friction provides a principled trigger for factuality correction, avoiding the over-correction pitfalls of divergence-only metrics.

6.4 Quantifying Logit Steering Impact

To evaluate our asymmetric logit steering, we measure the average probability of faithful tokens across three decoding strategies on the MQUAKE test set, where faithful tokens are context-grounded entities that correctly answer the query.

Table 4 reveals the fundamental difference between verification-based and steering-based approaches. The original target model assigns 18.73% probability to faithful tokens, reflecting parametric bias where internal knowledge conflicts with contextual evidence. While the target model considers the faithful token, it remains overshadowed by parametric priors.

Standard speculative decoding provides virtually no correction, with faithful token probability remaining at 18.91%. This stagnation is inherent to SD’s design: the verification mechanism determines whether to accept or reject draft tokens based on P⁡(accept)∝PM​(x)Pm​(x)P(\text{accept})\propto\frac{P_{M}(x)}{P_{m}(x)}, but it does not modify the target model’s output distribution. When the draft proposes a faithful token that the target assigns low probability, SD simply rejects it and resamples from the original PMP_{M}, perpetuating the same parametric bias. This reveals a fundamental limitation: verification-only approaches operate through accept/reject decisions without modifying the target distribution, thus they cannot correct distributional biases but only accelerate generation.

Table 4: Average probability of faithful tokens. Standard SD provides negligible improvement, while SFAD achieves substantial gains through logit steering.
Strategy Prob (%) Abs. Gain Rel. Gain
Original Target 18.73 – 1.0×
Standard SD 18.91 +0.18 1.01×
SFAD (Ours) 62.45 +43.72 3.33×
Table 5: General utility evaluation.
GSM8K Just-Eval (1–5)
Method Acc Help. Clar. Fact. Depth Eng.
Target (Greedy) 91.35 4.15 4.90 4.30 4.55 4.75
Standard SD 91.32 4.14 4.89 4.28 4.54 4.74
SFAD (Ours) 91.27 4.11 4.88 4.33 4.53 4.73

In contrast, SFAD elevates faithful token probability to 62.45%, achieving a 3.33× relative improvement and a 243× larger absolute gain compared to standard SD. This dramatic boost results from our asymmetric steering mechanism (Eq. (7)): when epistemic friction ℱt\mathcal{F}_{t} exceeds threshold τ\tau, the system triggers ReLU-based logit injection zt∗=zM,t+λt⋅ReLU​(zm,t−zM,t)z^{*}_{t}=z_{M,t}+\lambda_{t}\cdot\text{ReLU}(z_{m,t}-z_{M,t}), directly amplifying the draft model’s high logits for contextually faithful tokens. Unlike verification that merely accepts or rejects, steering fundamentally reshapes the target’s output distribution, transforming faithful tokens from minority candidates (18.73%) to dominant choices (62.45%). Critically, this intervention occurs selectively as it only activates when both distributional conflict and specialist conviction are present, maintaining the 2.48× speedup while achieving 85.2 faithfulness score. This analysis demonstrates that logit-level steering provides a principled mechanism for resolving knowledge conflicts, achieving context faithfulness unattainable through verification alone.

6.5 SFAD Preserves General Utility

We evaluate on GSM8K Cobbe et al. (2021) and Just-Eval Lin et al. (2024) to verify SFAD maintains general capabilities. As shown in Table 5, SFAD achieves 91.27% on GSM8K with negligible Just-Eval degradation. This stems from the selective nature of ℱt\mathcal{F}_{t}: when no knowledge conflict exists, friction remains below τ\tau and the system defaults to standard speculative decoding, confirming SFAD’s factuality enhancement does not sacrifice general utility or reasoning capability.

6.6 Generalization and Parameter Analysis

To evaluate generalization across model families, we apply SFAD with Llama-3.1-8B as the target model. As shown in Table 6, SFAD consistently improves faithfulness while maintaining inference efficiency. The coefficient γ\gamma modulates specialist certainty κt\kappa_{t} to regulate steering sensitivity. As shown in Table 7, lower γ\gamma increases intervention frequency but risks incorporating unreliable signals, while higher γ\gamma improves speedup at the cost of faithfulness.

Table 6: Llama-3.1-8B results on Foundation QA.
Method TriviaQA HotpotQA PopQA Rel. Latency
Vanilla Baseline
Llama-3.1-8B 49.23 38.45 75.18 1.00×\times
Decoding-level Baselines (Llama-3.1-8B)
CAD 37.82 36.21 68.45 2.00×\times
AdaCAD 78.56 42.31 74.28 2.18×\times
COIECD 79.43 42.88 75.56 2.35×\times
SFAD (Ours) 81.67 48.73 83.21 0.85×\times
Table 7: Ablation of sharpening coefficient γ\gamma.
𝜸\boldsymbol{\gamma} Steer (%) Faith. (↑\uparrow) Speedup (↑\uparrow) R-L (↑\uparrow)
1 31.2 83.7 2.31×\times 15.89
2 (Default) 22.4 85.2 2.48×\times 16.32
3 15.8 84.1 2.61×\times 16.18
4 9.3 78.5 2.74×\times 15.94

7 Conclusion

In this work, we propose SFAD, the first speculative decoding framework unifying contextual faithfulness with acceleration. Leveraging a context-faithful drafter as factuality sentinel, SFAD detects hallucinations via epistemic friction. Drafters are trained on ConFide, a dataset with atomic-level hallucination perturbations. At inference, SFAD selectively steers target logits upon conflict detection while preserving efficiency, achieving simultaneous speedup and faithfulness gains.

Limitations

SFAD requires a domain-aligned draft model trained via DPO, which introduces additional data construction overhead compared to standard speculative decoding. Furthermore, the friction threshold τ\tau may require tuning when applied to out-of-distribution domains.

Acknowledgment

Di Wang is supported in part by the funding BAS/1/1689-01-01, RGC/3/7125-01-01, FCC/1/5940-20-05, FCC/1/5940-06-02, and King Abdullah University of Science and Technology (KAUST) – Center of Excellence for Generative AI, under award number 5940 and a gift from Google.

Lijie Hu is supported in part by the funding BF0100.

References

  • Achiam et al. (2023) Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. 2023. Gpt-4 technical report. arXiv preprint arXiv:2303.08774.
  • Ali et al. (2025) Muhammad Asif Ali, Nawal Daftardar, Mutayyba Waheed, Jianbin Qin, and Di Wang. 2025. MQA-KEAL: Multi-hop question answering under knowledge editing for Arabic language. In Proceedings of the 31st International Conference on Computational Linguistics, pages 5629–5644, Abu Dhabi, UAE. Association for Computational Linguistics.
  • Bachmann et al. (2025) Gregor Bachmann, Sotiris Anagnostidis, Albert Pumarola, Markos Georgopoulos, Artsiom Sanakoyeu, Yuming Du, Edgar Schönfeld, Ali Thabet, and Jonas Kohler. 2025. Judge decoding: Faster speculative sampling requires going beyond model alignment. In The Thirteenth International Conference on Learning Representations.
  • Bi et al. (2025) Baolong Bi, Shaohan Huang, Yiwei Wang, Tianchi Yang, Zihan Zhang, Haizhen Huang, Lingrui Mei, Junfeng Fang, Zehao Li, Furu Wei, Weiwei Deng, Feng Sun, Qi Zhang, and Shenghua Liu. 2025. Context-DPO: Aligning language models for context-faithfulness. In Findings of the Association for Computational Linguistics: ACL 2025, pages 10280–10300, Vienna, Austria. Association for Computational Linguistics.
  • Chen et al. (2023) Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper. 2023. Accelerating large language model decoding with speculative sampling. arXiv preprint arXiv:2302.01318.
  • Chen et al. (2025) Guanzheng Chen, Qilong Feng, Jinjie Ni, Xin Li, and Michael Qizhe Shieh. 2025. RAPID: Long-context inference with retrieval-augmented speculative decoding. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 8093–8107. PMLR.
  • Chen et al. (2022) Hung-Ting Chen, Michael Zhang, and Eunsol Choi. 2022. Rich knowledge sources bring complex knowledge conflicts: Recalibrating models to reflect conflicting evidence. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 2292–2307, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.
  • Chen et al. (2026) Qiaoling Chen, Zijun Liu, Peng Sun, Shenggui Li, Guoteng Wang, Ziming Liu, Yonggang Wen, Siyuan Feng, and Tianwei Zhang. 2026. Respec: Towards optimizing speculative decoding in reinforcement learning systems. In Proceedings of Machine Learning and Systems, volume 8. MLSys.
  • Cheng et al. (2024a) Keyuan Cheng, Muhammad Asif Ali, Shu Yang, Gang Lin, Yuxuan Zhai, Haoyang Fei, Ke Xu, Lu Yu, Lijie Hu, and Di Wang. 2024a. Leveraging logical rules in knowledge editing: A cherry on the top. arXiv preprint arXiv:2405.15452.
  • Cheng et al. (2025a) Keyuan Cheng, Zijian Kan, Zhuoran Zhang, Muhammad Asif Ali, Lijie Hu, and Di Wang. 2025a. COMPKE: Complex question answering under knowledge editing. In Findings of the Association for Computational Linguistics: ACL 2025, pages 2557–2576, Vienna, Austria. Association for Computational Linguistics.
  • Cheng et al. (2024b) Keyuan Cheng, Gang Lin, Haoyang Fei, Yuxuan Zhai, Lu Yu, Muhammad Asif Ali, Lijie Hu, and Di Wang. 2024b. Multi-hop question answering under temporal knowledge editing. In Proceedings of the 1st Conference on Language Modeling.
  • Cheng et al. (2025b) Keyuan Cheng, Xudong Shen, Yihao Yang, Tengyue Wang, Yang Cao, Muhammad Asif Ali, Hanbin Wang, Lijie Hu, and Di Wang. 2025b. CODEMENV: Benchmarking large language models on code migration. In Findings of the Association for Computational Linguistics: ACL 2025, pages 2719–2744, Vienna, Austria. Association for Computational Linguistics.
  • Cobbe et al. (2021) Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168.
  • Feng et al. (2024) Zhangyin Feng, Xiaocheng Feng, Dezhi Zhao, Maojin Yang, and Bing Qin. 2024. Retrieval-generation synergy augmented large language models. In 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 11661–11665. IEEE.
  • Guo et al. (2025a) Daya Guo et al. 2025a. DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning. arXiv preprint arXiv:2501.12948.
  • Guo et al. (2025b) Zikun Guo, Xinyue Xu, Pei Xiang, Shu Yang, Xin Han, Di Wang, and Lijie Hu. 2025b. Benchmarking and mitigate psychological sycophancy in medical vision-language models. arXiv e-prints, pages arXiv–2509.
  • Guu et al. (2020) Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Mingwei Chang. 2020. Retrieval augmented language model pre-training. In Proceedings of the 37th International Conference on Machine Learning, volume 119 of Proceedings of Machine Learning Research, pages 3929–3938. PMLR.
  • Hu et al. (2025) Jingyu Hu, Shu Yang, Xilin Gong, Hongming Wang, Weiru Liu, and Di Wang. 2025. Monica: Real-time monitoring and calibration of chain-of-thought sycophancy in large reasoning models. arXiv preprint arXiv:2511.06419.
  • Huang et al. (2025a) Lei Huang, Xiaocheng Feng, Weitao Ma, Yuchun Fan, Xiachong Feng, Yangfan Ye, Weihong Zhong, Yuxuan Gu, Baoxin Wang, Dayong Wu, Guoping Hu, and Bing Qin. 2025a. Improving contextual faithfulness of large language models via retrieval heads-induced optimization. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 16896–16913, Vienna, Austria. Association for Computational Linguistics.
  • Huang et al. (2025b) Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. 2025b. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2):1–55.
  • Huang et al. (2025c) Yuheng Huang, Jiayang Song, Zhijie Wang, Shengming Zhao, Huaming Chen, Felix Juefei-Xu, and Lei Ma. 2025c. Look before you leap: An exploratory study of uncertainty analysis for large language models. IEEE Transactions on Software Engineering, 51(2):413–429.
  • Huang et al. (2025d) Zongle Huang, Lei Zhu, Zongyuan Zhan, Ting Hu, Weikai Mao, Xianzhi Yu, Yongpan Liu, and Tianyu Zhang. 2025d. MoESD: Unveil speculative decoding’s potential for accelerating sparse MoE. In Advances in Neural Information Processing Systems, volume 38. Neural Information Processing Systems Foundation.
  • Hurst et al. (2024) Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276.
  • Ji et al. (2023) Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Delong Chen, Wenliang Dai, Ho Shu Chan, Andrea Madotto, and Pascale Fung. 2023. Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12):1–38.
  • Joshi et al. (2017) Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017. TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611, Vancouver, Canada. Association for Computational Linguistics.
  • Kaddour et al. (2023) Jean Kaddour, Joshua Harris, Maximilian Mozes, Herbie Bradley, Roberta Raileanu, and Robert McHardy. 2023. Challenges and applications of large language models. arXiv preprint arXiv:2307.10169.
  • Kamalloo et al. (2023) Ehsan Kamalloo, Aref Jafari, Xinyu Zhang, Nandan Thakur, and Jimmy Lin. 2023. HAGRID: A human-LLM collaborative dataset for generative information-seeking with attribution. arXiv preprint arXiv:2307.16883.
  • Khandelwal et al. (2025) Anant Khandelwal, Manish Gupta, and Puneet Agrawal. 2025. CoCoA: Confidence- and context-aware adaptive decoding for resolving knowledge conflicts in large language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 6835–6855, Suzhou, China. Association for Computational Linguistics.
  • Lei et al. (2025) Deren Lei, Yaxi Li, Siyao Li, Mengya Hu, Rui Xu, Ken Archer, Mingyu Wang, Emily Ching, and Alex Deng. 2025. FactCG: Enhancing fact checkers with graph-based multi-hop data. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 5002–5020, Albuquerque, New Mexico. Association for Computational Linguistics.
  • Leviathan et al. (2023) Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023. Fast inference from transformers via speculative decoding. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 19274–19286. PMLR.
  • Li and Yu (2025) Anguo Li and Lei Yu. 2025. Summary factual inconsistency detection based on LLMs enhanced by universal information extraction. In Findings of the Association for Computational Linguistics: ACL 2025, pages 25450–25465, Vienna, Austria. Association for Computational Linguistics.
  • Li et al. (2026a) Guanghao Li, Zhihui Fu, Min Fang, Qibin Zhao, Ming Tang, Chun Yuan, and Jun Wang. 2026a. DiffuSpec: Unlocking diffusion language models for speculative decoding. In Findings of the Association for Computational Linguistics: ACL 2026, pages 20896–20910, San Diego, California, United States. Association for Computational Linguistics.
  • Li et al. (2026b) Hongji Li, Manjiang Yu, Junchi Yao, Priyanka Singh, Xue Li, Di Wang, and Lijie Hu. 2026b. Towards reasoning-preserving unlearning in multimodal large language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10251–10261.
  • Li et al. (2025) Yiwei Li, Jiayi Shi, Shaoxiong Feng, Peiwen Yuan, Xinglin Wang, Yueqi Zhang, Ji Zhang, Chuyi Tan, Boyuan Pan, Yao Hu, and Kan Li. 2025. Speculative decoding for multi-sample inference. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 12523–12533, Suzhou, China. Association for Computational Linguistics.
  • Li et al. (2024) Yuhui Li, Fangyun Wei, Chao Zhang, and Hongyang Zhang. 2024. EAGLE: Speculative sampling requires rethinking feature uncertainty. In Proceedings of the 41st International Conference on Machine Learning, volume 235 of Proceedings of Machine Learning Research, pages 28935–28948. PMLR.
  • Lin et al. (2024) Bill Yuchen Lin, Abhilasha Ravichander, Ximing Lu, Nouha Dziri, Melanie Sclar, Khyathi Chandu, Chandra Bhagavatula, and Yejin Choi. 2024. The unlocking spell on base llms: Rethinking alignment via in-context learning. In The Twelfth International Conference on Learning Representations.
  • Malaviya et al. (2024) Chaitanya Malaviya, Subin Lee, Sihao Chen, Elizabeth Sieber, Mark Yatskar, and Dan Roth. 2024. ExpertQA: Expert-curated questions and attributed answers. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 3025–3045, Mexico City, Mexico. Association for Computational Linguistics.
  • Mallen et al. (2023) Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. When not to trust language models: Investigating effectiveness of parametric and non-parametric memories. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 9802–9822, Toronto, Canada. Association for Computational Linguistics.
  • MetaAI (2024) MetaAI. 2024. Introducing llama 3.1: Our most capable models to date. https://ai.meta.com/blog/meta-llama-3-1/.
  • Miao et al. (2024) Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, Chunan Shi, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia. 2024. Specinfer: Accelerating large language model serving with tree-based speculative inference and verification. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3, pages 932–949, La Jolla, CA, USA. Association for Computing Machinery.
  • Min et al. (2023) Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, and Hannaneh Hajishirzi. 2023. FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 12076–12100, Singapore. Association for Computational Linguistics.
  • Narayan et al. (2018) Shashi Narayan, Shay B. Cohen, and Mirella Lapata. 2018. Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1797–1807, Brussels, Belgium. Association for Computational Linguistics.
  • Pan et al. (2024) Shirui Pan, Linhao Luo, Yufei Wang, Chen Chen, Jiapu Wang, and Xindong Wu. 2024. Unifying large language models and knowledge graphs: A roadmap. IEEE Transactions on Knowledge and Data Engineering, 36(7):3580–3599.
  • Petroni et al. (2019) Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019. Language models as knowledge bases? In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2463–2473, Hong Kong, China. Association for Computational Linguistics.
  • Qin et al. (2025) Yujia Qin, Shengding Hu, Yankai Lin, Weize Chen, Ning Ding, Ganqu Cui, Zheni Zeng, Xuanhe Zhou, Yufei Huang, Chaojun Xiao, Chi Han, Yi Ren Fung, Yusheng Su, Huadong Wang, Cheng Qian, Runchu Tian, Kunlun Zhu, Shihao Liang, Xingyu Shen, and 23 others. 2025. Tool learning with foundation models. ACM Computing Surveys, 57(4).
  • Rafailov et al. (2023) Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning, Stefano Ermon, and Chelsea Finn. 2023. Direct preference optimization: Your language model is secretly a reward model. In Advances in Neural Information Processing Systems, volume 36, pages 53728–53741. Curran Associates, Inc.
  • Roberts et al. (2020) Adam Roberts, Colin Raffel, and Noam Shazeer. 2020. How much knowledge can you pack into the parameters of a language model? In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 5418–5426, Online. Association for Computational Linguistics.
  • Rosenthal et al. (2025) Sara Rosenthal, Avirup Sil, Radu Florian, and Salim Roukos. 2025. CLAPnq: Cohesive long-form answers from passages in natural questions for RAG systems. Transactions of the Association for Computational Linguistics, 13:53–72.
  • Sadhukhan et al. (2025) Ranajoy Sadhukhan, Jian Chen, Zhuoming Chen, Vashisth Tiwari, Ruihang Lai, Jinyuan Shi, Ian En-Hsu Yen, Avner May, Tianqi Chen, and Beidi Chen. 2025. Magicdec: Breaking the latency-throughput tradeoff for long context generation with speculative decoding. In The Thirteenth International Conference on Learning Representations.
  • Shi et al. (2024) Weijia Shi, Xiaochuang Han, Mike Lewis, Yulia Tsvetkov, Luke Zettlemoyer, and Wen-tau Yih. 2024. Trusting your evidence: Hallucinate less with context-aware decoding. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), pages 783–791, Mexico City, Mexico. Association for Computational Linguistics.
  • Sun et al. (2024) Hanshi Sun, Zhuoming Chen, Xinyu Yang, Yuandong Tian, and Beidi Chen. 2024. Triforce: Lossless acceleration of long sequence generation with hierarchical speculative decoding. In Proceedings of the Conference on Language Modeling.
  • Tang et al. (2024a) Liyan Tang, Philippe Laban, and Greg Durrett. 2024a. MiniCheck: Efficient fact-checking of LLMs on grounding documents. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 8818–8847, Miami, Florida, USA. Association for Computational Linguistics.
  • Tang et al. (2024b) Liyan Tang, Igor Shalyminov, Amy Wong, Jon Burnsky, Jake Vincent, Yu’an Yang, Siffi Singh, Song Feng, Hwanjun Song, Hang Su, Lijia Sun, Yi Zhang, Saab Mansour, and Kathleen McKeown. 2024b. TofuEval: Evaluating hallucinations of LLMs on topic-focused dialogue summarization. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 4455–4480, Mexico City, Mexico. Association for Computational Linguistics.
  • Tiwari et al. (2025) Rishabh Tiwari, Haocheng Xi, Aditya Tomar, Coleman Richard Charles Hooper, Sehoon Kim, Maxwell Horton, Mahyar Najibi, Michael W. Mahoney, Kurt Keutzer, and Amir Gholami. 2025. QuantSpec: Self-speculative decoding with hierarchical quantized KV cache. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 59668–59686. PMLR.
  • Wan et al. (2023) David Wan, Mengwen Liu, Kathleen McKeown, Markus Dreyer, and Mohit Bansal. 2023. Faithfulness-aware decoding strategies for abstractive summarization. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 2864–2880, Dubrovnik, Croatia. Association for Computational Linguistics.
  • Wang et al. (2025a) Cheng-Long Wang, Qi Li, Zihang Xiang, Yinzhi Cao, and Di Wang. 2025a. Towards lifecycle unlearning commitment management: Measuring sample-level unlearning completeness. In 34th USENIX Security Symposium (USENIX Security 25), pages 6481–6500, Seattle, WA. USENIX Association.
  • Wang et al. (2025b) Han Wang, Archiki Prasad, Elias Stengel-Eskin, and Mohit Bansal. 2025b. AdaCAD: Adaptively decoding to balance conflicts between contextual and parametric knowledge. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pages 11636–11652, Albuquerque, New Mexico. Association for Computational Linguistics.
  • Wang et al. (2026) Keyu Wang, Jin Li, Shu Yang, Zhuoran Zhang, and Di Wang. 2026. When truth is overridden: Uncovering the internal origins of sycophancy in large language models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pages 33566–33574.
  • Wang et al. (2025c) Liangyu Wang, Huanyi Xie, Xinhai Wang, Tianjin Huang, Mengdi Li, and Di Wang. 2025c. Infinite sampling: Efficient and stable grouped rl training for large language models. arXiv preprint arXiv:2506.22950.
  • Wang et al. (2025d) Xuekang Wang, Shengyu Zhu, and Xueqi Cheng. 2025d. Speculative safety-aware decoding. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 12827–12841, Suzhou, China. Association for Computational Linguistics.
  • Wei et al. (2024) Jason Wei, Nguyen Karina, Hyung Won Chung, Yunxin Joy Jiao, Spencer Papay, Amelia Glaese, John Schulman, and William Fedus. 2024. Measuring short-form factuality in large language models. arXiv preprint arXiv:2411.04368.
  • Xia et al. (2023) Heming Xia, Tao Ge, Peiyi Wang, Si-Qing Chen, Furu Wei, and Zhifang Sui. 2023. Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation. In Findings of the Association for Computational Linguistics: EMNLP 2023, pages 3909–3925, Singapore. Association for Computational Linguistics.
  • Xie et al. (2024) Jian Xie, Kai Zhang, Jiangjie Chen, Renze Lou, and Yu Su. 2024. Adaptive chameleon or stubborn sloth: Revealing the behavior of large language models in knowledge conflicts. In The Twelfth International Conference on Learning Representations.
  • Xu et al. (2024) Rongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang, Hongru Wang, Yue Zhang, and Wei Xu. 2024. Knowledge conflicts for LLMs: A survey. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 8541–8565, Miami, Florida, USA. Association for Computational Linguistics.
  • Xu (2023) Zhichao Xu. 2023. Context-aware decoding reduces hallucination in query-focused summarization. arXiv preprint arXiv:2312.14335.
  • Yan et al. (2025) Shi-Qi Yan, Quan Liu, and Zhen-Hua Ling. 2025. RPO: Retrieval preference optimization for robust retrieval-augmented generation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 5228–5240, Vienna, Austria. Association for Computational Linguistics.
  • Yang et al. (2025a) An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. 2025a. Qwen3 technical report. arXiv preprint arXiv:2505.09388.
  • Yang et al. (2026) Penghui Yang, Cunxiao Du, Fengzhuo Zhang, Haonan Wang, Tianyu Pang, Chao Du, and Bo An. 2026. LongSpec: Long-context lossless speculative decoding with efficient drafting and verification. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1826–1844, San Diego, California, United States. Association for Computational Linguistics.
  • Yang et al. (2025b) Shu Yang, Junchao Wu, Xilin Gong, Xuansheng Wu, Derek Wong, Ninhao Liu, and Di Wang. 2025b. Investigating cot monitorability in large reasoning models. arXiv preprint arXiv:2511.08525.
  • Yang et al. (2025c) Tiancheng Yang, Lin Zhang, Jiaye Lin, Guimin Hu, Di Wang, and Lijie Hu. 2025c. Tracing and mitigating hallucinations in multimodal llms via dynamic attention localization. arXiv preprint arXiv:2509.07864.
  • Yang et al. (2018) Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2369–2380, Brussels, Belgium. Association for Computational Linguistics.
  • Yuan et al. (2024) Xiaowei Yuan, Zhao Yang, Yequan Wang, Shengping Liu, Jun Zhao, and Kang Liu. 2024. Discerning and resolving knowledge conflicts through adaptive decoding with contextual information-entropy constraint. In Findings of the Association for Computational Linguistics: ACL 2024, pages 3903–3922, Bangkok, Thailand. Association for Computational Linguistics.
  • Zhang et al. (2025a) Zhuoran Zhang, Yongxiang Li, Zijian Kan, Keyuan Cheng, Lijie Hu, and Di Wang. 2025a. Locate-then-edit for multi-hop factual recall under knowledge editing. In Proceedings of the 42nd International Conference on Machine Learning, volume 267 of Proceedings of Machine Learning Research, pages 75369–75391. PMLR.
  • Zhang et al. (2025b) Zhuoran Zhang, Tengyue Wang, Xilin Gong, Yang Shi, Haotian Wang, Di Wang, and Lijie Hu. 2025b. When modalities conflict: How unimodal reasoning uncertainty governs preference dynamics in mllms. arXiv preprint arXiv:2511.02243.
  • Zhong et al. (2023) Zexuan Zhong, Zhengxuan Wu, Christopher Manning, Christopher Potts, and Danqi Chen. 2023. MQuAKE: Assessing knowledge editing in language models via multi-hop questions. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 15686–15702, Singapore. Association for Computational Linguistics.
  • Zhou et al. (2026) Wenrui Zhou, Mohamed Hendy, Shu Yang, Qingsong Yang, Zikun Guo, Yuyu Luo, Lijie Hu, and Di Wang. 2026. Flattery in motion: Benchmarking and analyzing sycophancy in video-LLMs. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8146–8172, San Diego, California, United States. Association for Computational Linguistics.

Appendix A Parameter Analysis

A.1 Role of Contextual Plausibility Mask (η\eta)

A pivotal challenge in logit-level steering is the potential for linguistic degradation: a specialist may propose a factually accurate token that is syntactically incompatible with the target model’s prefix. To evaluate how our Contextual Plausibility Mask (CPM) mitigates this risk, we conduct an ablation study on the threshold η\eta using the XSum dataset.

Table 8: Ablation of the Plausibility Threshold η\eta on XSum. η=0.1\eta=0.1 is the default setting used in Table 2; η=0\eta=0 denotes steering without the linguistic safety guard.
Configuration ROUGE-L (↑\uparrow) BERT-P (↑\uparrow) AlignScore (↑\uparrow)
SFAD (Standard, η=0.1\eta=0.1) 16.32 93.97 87.37
SFAD w/o CPM (η=0\eta=0) 14.78 91.25 87.49
Strict CPM (η=0.5\eta=0.5) 15.42 92.84 81.12
Vanilla Qwen3-14B 13.67 91.67 72.86

As illustrated in Table 8, removing the mask (η=0\eta=0) yields the highest AlignScore (87.49), confirming that the specialist model is indeed capable of forcing factual corrections. However, this comes at a substantial cost to linguistic integrity: ROUGE-L drops by 1.54 points and BERT-P falls by 2.72 points compared to the standard SFAD. Qualitative analysis reveals that without CPM, the model often injects correct entities with mismatched syntax (e.g., incorrect prepositional usage or broken noun phrases).

By introducing η=0.1\eta=0.1, SFAD successfully filters out these "linguistically dissonant" corrections. This explains the competitive results in Table 2, where SFAD nearly matches the 70B Frontier Model’s performance in both fluency and factuality. The CPM acts as a "safety guard," ensuring that steering only occurs within the target model’s plausible manifold, thus achieving a superior Pareto-optimal balance between knowledge correction and natural language generation.

A.2 Sensitivity Analysis of Friction Threshold τ\tau

The friction threshold τ\tau serves as the primary hyperparameter for calibrating the sensitivity of SFAD’s hallucination detector. It determines the operating point on the factuality-efficiency Pareto frontier by controlling the ratio between the Steering Path and the Fast Path. We evaluate the impact of τ\tau on the LLM-AggreFact test set, reporting the Steering Ratio (percentage of tokens where Ft≥τF_{t}\geq\tau), Faithfulness Score (MiniCheck), and inference speedup (ATGA).

Table 9: Sensitivity analysis of τ\tau on LLM-AggreFact. Steering Ratio denotes the percentage of generated tokens processed via the logit steering path. Speedup is relative to greedy decoding without speculation.
Threshold τ\tau Steering Ratio (%) Faithfulness Score (↑\uparrow) Speedup (ATGA) (↑\uparrow)
τ=0.1\tau=0.1 (Aggressive) 48.2% 86.7 2.12×\times
τ=0.3\tau=0.3 35.6% 85.9 2.31×\times
τ=0.5\tau=0.5 (Default) 22.4% 85.2 2.48×\times
τ=0.7\tau=0.7 12.1% 62.4 2.65×\times
τ=0.9\tau=0.9 (Conservative) 4.5% 41.8 2.82×\times
Standard SD 0.0% 38.5 2.18×\times

As shown in Table 9, the Steering Ratio decreases monotonically as τ\tau increases, reflecting a more selective intervention strategy. When τ=0.1\tau=0.1, the model intervenes frequently, achieving the highest faithfulness but incurring a latency penalty due to frequent target model logit computations. Conversely, at τ=0.9\tau=0.9, the system defaults to standard speculative decoding (Fast Path) most of the time, maximizing speed but failing to resolve knowledge conflicts.

Crucially, our default setting of τ=0.5\tau=0.5 achieves a "sweet spot": it resolves the majority of hallucinations (surpassing standard SD by 46.7 points) while maintaining a high speedup of 2.48×\times. This demonstrates that SFAD performs surgical interventions only when necessary.

A.3 Effectiveness of Asymmetric Logit Fusion Operators

To empirically validate the theoretical advantage of our asymmetric steering law (Theorem 3), we conduct an ablation study comparing different logit fusion operators. In this experiment, we keep the adaptive gating mechanism (λt\lambda_{t}) and the Contextual Plausibility Mask (CPM) constant, only varying how the specialist’s logits zmz_{m} are integrated with the generalist’s logits zMz_{M} when steering is triggered. We evaluate four strategies:

  1. (i)

    Linear Sum: z∗=zM+λt​zmz^{*}=z_{M}+\lambda_{t}z_{m}, which directly adds the specialist’s signals.

  2. (ii)

    Linear Interpolation: z∗=(1−λt)​zM+λt​zmz^{*}=(1-\lambda_{t})z_{M}+\lambda_{t}z_{m}, a standard weighted average used in model blending.

  3. (iii)

    Subtractive Contrast: z∗=zM+λt​(zM−zb​a​s​e)z^{*}=z_{M}+\lambda_{t}(z_{M}-z_{base}), similar to contrastive decoding but using the specialist as the positive signal.

  4. (iv)

    Asymmetric Steering (SFAD): z∗=zM+λt​ReLU​(zm−zM)z^{*}=z_{M}+\lambda_{t}\text{ReLU}(z_{m}-z_{M}), our proposed unidirectional injection.

The results on PopQA and HotpotQA are summarized in Table 10.

Table 10: Ablation of Logit Fusion Operators. Faithfulness is measured by Exact Match (EM) for PopQA and MiniCheck for HotpotQA. Fluency is represented by ROUGE-L (R-L).
Fusion Strategy PopQA HotpotQA
EM (↑\uparrow) R-L (↑\uparrow) Faith (↑\uparrow) R-L (↑\uparrow)
Linear Sum 78.42 14.12 46.21 35.89
Linear Interpolation 81.35 15.67 48.55 39.42
Subtractive Contrast 83.12 13.05 49.32 34.11
SFAD (Ours) 86.39 16.32 52.19 41.34

As shown in Table 10, the Asymmetric Steering operator consistently outperforms other fusion methods. Linear Sum and Interpolation tend to dilute the generalist’s linguistic priors, leading to a drop in ROUGE-L (fluency). Conversely, while Subtractive Contrast can highlight factual differences, it often results in "negative constraints" that suppress even valid tokens, harming coherence. SFAD’s ReLU-based injection ensures that the specialist only intervenes to boost tokens where it has higher confidence than the generalist, preserving the natural language manifold of the target model while effectively correcting factual errors.

Appendix B Generalization Across Model Families

We provide more results on Llama-3.1-8B in Table 11 and Table 12.

Table 11: Llama-3.1-8B results on Summarization.
XSum TofuEval
Method R-L BERT-P AlignScore AlignScore Rel. Latency
Vanilla Baseline
Llama-3.1-8B 12.89 90.83 71.24 58.12 1.00×\times
Decoding-level Baselines (Llama-3.1-8B)
CAD 13.76 92.87 82.56 81.45 2.00×\times
AdaCAD 14.08 93.41 84.03 83.29 2.22×\times
COIECD 12.91 90.22 72.15 59.34 2.48×\times
SFAD (Ours) 15.47 93.12 85.68 85.87 0.87×\times
Table 12: Llama-3.1-8B results on Long-Form QA.
CLAPNQ ExpertQA HAGRID
Method R-L Faith R-L Faith R-L Faith Rel. Latency
Vanilla Baseline
Llama-3.1-8B 16.34 57.89 30.21 49.56 16.12 55.87 1.00×\times
Decoding-level Baselines (Llama-3.1-8B)
CAD 17.41 58.43 32.15 51.72 17.03 56.45 2.00×\times
AdaCAD 17.58 60.51 31.49 52.93 16.31 58.12 2.21×\times
COIECD 17.72 60.14 32.78 54.48 16.52 57.94 2.45×\times
SFAD (Ours) 19.87 68.45 34.56 62.37 18.93 66.23 0.81×\times

These cross-family results indicate that SFAD’s gains are not tied to a single target backbone. Across summarization and long-form QA, the same decoding mechanism improves faithfulness while preserving the intended latency advantage, suggesting that the learned factuality-aware draft model and friction-triggered steering provide a portable correction signal across model families.

Appendix C Examples of Factual Perturbation for ConFide

We present two representative cases of how a faithful response is transformed into a hallucinated negative sample (yly_{l}) through our atomic-level perturbation mechanism.

Case 1: Entity Swap (τe​n​t\tau_{ent}) Source Context:
“…Muto graduated from Keio University in Tokyo with an economics degree two weeks ago. He will join Chelsea’s partner club Vitesse Arnhem on loan…”
Original Faithful Response (ywy_{w}):
“Yoshinori Muto, who recently completed his economics degree at Keio University, is set to join Chelsea.”
Atomic Fact Decomposition:
aj=(Yoshinori Muto,graduated from,Keio University)a_{j}=(\text{Yoshinori Muto},\text{graduated from},\text{{Keio University}})
Perturbation Operation:
Swap target entity Keio University with intra-context distractor
Vitesse Arnhem.
Generated Hallucinated Response (yly_{l}):
“Yoshinori Muto, who recently completed his economics degree at Vitesse Arnhem, is set to join Chelsea.”
Case 2: Relation Inversion (τr​e​l\tau_{rel}) Source Context:
“…Ogane claims that Chelsea’s interest in Muto is not connected to the £200million sponsorship deal they signed with Yokohama Rubber…”
Original Faithful Response (ywy_{w}):
“The club president stated that the move for Muto is independent of the sponsorship deal with Yokohama Rubber.”
Atomic Fact Decomposition:
aj=(CLOSE\displaystyle a_{j}=( Chelsea’s interest,is not connected to,\displaystyle\text{Chelsea's interest},\ \text{{is not connected to}}, OPENSponsorship Deal)\displaystyle\text{Sponsorship Deal}) Perturbation Operation:
Apply logical inversion (¬rj\neg r_{j}) to the relation is not connected to.
Generated Hallucinated Response (yly_{l}):
“The club president stated that the move for Muto is a direct result of the sponsorship deal with Yokohama Rubber.”

Appendix D Algorithms

This section provides the algorithmic details of the proposed framework. Algorithm 1 summarizes the inference-time procedure of SFAD, including friction-based conflict detection and asymmetric logit steering. Algorithm 2 describes the construction process of the ConFide preference dataset used for training the factuality-aware draft model.

Algorithm 1 SFAD: Speculative Factuality-Aware Decoding
1:  Input: Target model MM, DPO-aligned draft model mm, prefix context x<tx_{<t}, friction threshold τ\tau, sharpening coefficient γ\gamma, plausibility threshold η\eta, sigmoid scale β\beta.
2:  Output: Context-faithful generated token xtx_{t}.
3:  while xt≠EOSx_{t}\neq\text{EOS} do
4:   // Step 1: Distribution Computation
5:   Compute logit vectors 𝐳M,t\mathbf{z}_{M,t} from M⁡(x<t)M(x_{<t}) and 𝐳m,t\mathbf{z}_{m,t} from m⁡(x<t)m(x_{<t})
6:   Compute probabilities ℙM=Softmax​(𝐳M,t)\mathbb{P}_{M}=\text{Softmax}(\mathbf{z}_{M,t}) and ℙm=Softmax​(𝐳m,t)\mathbb{P}_{m}=\text{Softmax}(\mathbf{z}_{m,t})
7:   // Step 2: Factuality Conflict Detection
8:   Calculate Specialist Certainty: κt=(1−−∑v∈𝒱ℙm(v)logℙm(v)log⁡|𝒱|)γ\kappa_{t}=\left(1-\frac{-\sum_{v\in\mathcal{V}}\mathbb{P}_{m}(v)\log\mathbb{P}_{m}(v)}{\log|\mathcal{V}|}\right)^{\gamma}
9:   Calculate Epistemic Friction: ℱt=DJ​S(ℙM∥ℙm)⋅κt\mathcal{F}_{t}=D_{JS}(\mathbb{P}_{M}\parallel\mathbb{P}_{m})\cdot\kappa_{t}
10:   if ℱt≥τ\mathcal{F}_{t}\geq\tau then
11:   // Steering Path: Distribution Correction
12:   Identify plausible set: 𝒱C​P​C={x∈𝒱:ℙm​(x|x<t)≥η⋅maxw∈𝒱⁡ℙm​(w|x<t)}\mathcal{V}_{CPC}=\{x\in\mathcal{V}:\mathbb{P}_{m}(x|x_{<t})\geq\eta\cdot\max_{w\in\mathcal{V}}\mathbb{P}_{m}(w|x_{<t})\}
13:   Compute gating scalar: λt=11+exp⁡(−β⁡(ℱt−τ))\lambda_{t}=\frac{1}{1+\exp(-\beta(\mathcal{F}_{t}-\tau))}
14:   Apply Asymmetric Logit Steering:
15:   𝐳t∗=𝐳M,t+λt⋅max⁡(0,𝐳m,t−𝐳M,t)⋅𝕀⁡(x∈𝒱C​P​C)\mathbf{z}^{*}_{t}=\mathbf{z}_{M,t}+\lambda_{t}\cdot\max(0,\mathbf{z}_{m,t}-\mathbf{z}_{M,t})\cdot\mathbb{I}(x\in\mathcal{V}_{CPC})
16:   Sample next token: xt∼Softmax​(𝐳t∗)x_{t}\sim\text{Softmax}(\mathbf{z}^{*}_{t})
17:   else
18:   // Fast Path: Standard Speculative Decoding
19:   Sample draft token x~t\tilde{x}_{t} from ℙm\mathbb{P}_{m}
20:   xt←Verify​(x~t,ℙM)x_{t}\leftarrow\text{Verify}(\tilde{x}_{t},\mathbb{P}_{M}) {Standard rejection sampling}
21:   end if
22:   Update sequence: x<t+1←[x<t;xt]x_{<t+1}\leftarrow[x_{<t};x_{t}]
23:  end while
Algorithm 2 ConFide: Factuality-Aware Preference Data Construction
0:  𝒟={(x,yr​e​f,l)}i=1N\mathcal{D}=\{(x,y_{ref},l)\}_{i=1}^{N}, Teacher ℳT\mathcal{M}_{T}, Paraphraser πp​a​r​a\pi_{para}, Generator fr​e​cf_{rec}
1:  𝒫←∅\mathcal{P}\leftarrow\emptyset
2:  for each (x,yr​e​f,l)∈𝒟(x,y_{ref},l)\in\mathcal{D} do
3:   𝒜={a1,…,am}←fd​e​c​(yr​e​f)\mathcal{A}=\{a_{1},\dots,a_{m}\}\leftarrow f_{dec}(y_{ref}) {Atomic Fact Decomposition}
4:   // Negative Generation: Perturb aj∈𝒜a_{j}\in\mathcal{A} via Entity Swap, Numerical, or Relation Inversion
5:   aj′←ϕ⁡(aj),yl←fr​e​c​(𝒜∖{aj}∪{aj′})a^{\prime}_{j}\leftarrow\phi(a_{j}),\quad y_{l}\leftarrow f_{rec}(\mathcal{A}\setminus\{a_{j}\}\cup\{a^{\prime}_{j}\}) {Corrupt and Reconstruct}
6:   // Positive Refinement: Paraphrase faithful samples or utilize Teacher for corrections
7:   yw←(l=1)​?​πp​a​r​a​(yr​e​f):ℳT​(correct ​yr​e​f​ based on ​x)y_{w}\leftarrow(l=1)\ ?\ \pi_{para}(y_{ref})\ :\ \mathcal{M}_{T}(\text{correct }y_{ref}\text{ based on }x)
8:   𝒫←𝒫∪{(x,yw,yl)}\mathcal{P}\leftarrow\mathcal{P}\cup\{(x,y_{w},y_{l})\} {Assemble Preference Pair}
9:  end for
10:  return 𝒫\mathcal{P}

Appendix E Theoretical Analysis of SFAD

In this section, we provide a rigorous formal justification for the Speculative Factuality-Aware Decoding (SFAD) framework. Our analysis bridges the gap between empirical performance and distributional theory by focusing on three core pillars: (1) Semantic Integrity, establishing that our steering mechanism preserves the linguistic manifold; (2) Factuality Amplification, proving how DPO-induced margins yield exponential gains in faithful tokens; and (3) Dynamic Stability, demonstrating the robustness of asymmetric steering against the "zero-probability trap" prevalent in contrastive methods.

E.1 Semantic Consistency and the Linguistic Manifold Bound

A fundamental challenge in logit steering is the alignment-fidelity trade-off: factual corrections must not drive the model’s output distribution off the natural language manifold ℳ\mathcal{M}. We formalize the role of the Contextual Plausibility Mask (CPM) as a projection operator that constrains the steering signal to the support of the generalist’s distribution.

Theorem 1 (Manifold Projection Bound).

Let 𝒱C​P​C={v∈𝒱:Pm​(v|x<t)≥η⋅maxw⁡Pm​(w|x<t)}\mathcal{V}_{CPC}=\{v\in\mathcal{V}:P_{m}(v|x_{<t})\geq\eta\cdot\max_{w}P_{m}(w|x_{<t})\} be the set of plausible candidates. For any steering intensity λt∈[0,1]\lambda_{t}\in[0,1], the Total Variation (TV) distance between the steered distribution P∗P^{*} and the original generalist distribution PMP_{M} is bounded by the probability mass concentrated on the plausible manifold:

dT​V​(P∗,PM)\displaystyle d_{TV}(P^{*},P_{M}) ≤12​∑v∈𝒱C​P​CPM​(v)​|eλt​Δ​zv𝒵t∗−1|\displaystyle\leq\frac{1}{2}\sum_{v\in\mathcal{V}_{CPC}}P_{M}(v)\left|\frac{e^{\lambda_{t}\Delta z_{v}}}{\mathcal{Z}_{t}^{*}}-1\right| (10)
+12​|1−1𝒵t∗|​(1−PM​(𝒱C​P​C)).\displaystyle+\frac{1}{2}\left|1-\frac{1}{\mathcal{Z}_{t}^{*}}\right|(1-P_{M}(\mathcal{V}_{CPC})).

where Δ​zv=(zm,v−zM,v)+\Delta z_{v}=(z_{m,v}-z_{M,v})_{+} and 𝒵t∗=𝔼v∼PM​[eλt​Δ​zv]\mathcal{Z}_{t}^{*}=\mathbb{E}_{v\sim P_{M}}[e^{\lambda_{t}\Delta z_{v}}] is the partition normalization factor.

Proof.

By the definition of the CPM-based steering in Eq. (7), for any token v∉𝒱C​P​Cv\notin\mathcal{V}_{CPC}, the steering signal Δ​zv\Delta z_{v} is nullified by the indicator function 𝕀⁡(x∈𝒱C​P​C)\mathbb{I}(x\in\mathcal{V}_{CPC}). Consequently, the steered logit zv∗z^{*}_{v} equals the original logit zM,vz_{M,v}, and the probability simplifies to P∗​(v)=PM​(v)/𝒵t∗P^{*}(v)=P_{M}(v)/\mathcal{Z}_{t}^{*}.

The TV distance is defined as dT​V​(P∗,PM)=12​∑v∈𝒱|P∗​(v)−PM​(v)|d_{TV}(P^{*},P_{M})=\frac{1}{2}\sum_{v\in\mathcal{V}}|P^{*}(v)-P_{M}(v)|. Partitioning the vocabulary 𝒱\mathcal{V} into the plausible set 𝒱C​P​C\mathcal{V}_{CPC} and its complement 𝒱c\mathcal{V}^{c}, we obtain:

2​dT​V\displaystyle 2d_{TV} =∑v∈𝒱C​P​C|PM​(v)​eλt​Δ​zv𝒵t∗−PM​(v)|\displaystyle=\sum_{v\in\mathcal{V}_{CPC}}\left|\frac{P_{M}(v)e^{\lambda_{t}\Delta z_{v}}}{\mathcal{Z}_{t}^{*}}-P_{M}(v)\right|
+∑v∈𝒱c|PM​(v)𝒵t∗−PM(v)|\displaystyle\quad+\sum_{v\in\mathcal{V}^{c}}\left|\frac{P_{M}(v)}{\mathcal{Z}_{t}^{*}}-P_{M}(v)\right|
=∑v∈𝒱C​P​CPM​(v)​|eλt​Δ​zv𝒵t∗−1|\displaystyle=\sum_{v\in\mathcal{V}_{CPC}}P_{M}(v)\left|\frac{e^{\lambda_{t}\Delta z_{v}}}{\mathcal{Z}_{t}^{*}}-1\right|
+|1𝒵t∗−1|∑v∈𝒱cPM(v).\displaystyle\quad+\left|\frac{1}{\mathcal{Z}_{t}^{*}}-1\right|\sum_{v\in\mathcal{V}^{c}}P_{M}(v). (11)

Substituting PM​(𝒱c)=1−PM​(𝒱C​P​C)P_{M}(\mathcal{V}^{c})=1-P_{M}(\mathcal{V}_{CPC}) completes the proof. This bound illustrates that the divergence is strictly governed by the specialist’s corrective signal magnitude within the linguistic manifold. Since mm and MM share a common pre-training ancestry, their high-density regions in ℳ\mathcal{M} are naturally aligned, ensuring that steering only re-allocates mass within valid semantic clusters. ∎

Remark 1.

The Anchor Property. Theorem 1 reveals that SFAD acts as a "factual re-ranker" rather than a "stochastic generator." By anchoring the correction to the generalist’s manifold ℳ\mathcal{M}, the framework prevents out-of-distribution (OOD) artifacts. Unlike additive noise or unconstrained logit shifts, the CPM ensures that the model never samples tokens that are linguistically nonsensical, explaining the high fluency scores observed in our experiments.

E.2 Factuality Amplification via DPO-Induced Margins

While the previous section established safety, we now prove that SFAD effectively resolves knowledge conflicts by leveraging the preference gap instilled during the Direct Preference Optimization (DPO) phase.

Proposition 1 (Exponential Posterior Gain).

Let vc​t​x​tv_{ctxt} be a faithful token and vm​e​mv_{mem} be a hallucinated token (favored by the target model’s parametric prior). If the specialist mm satisfies the DPO optimality condition with a learned margin γm=zm​(vc​t​x​t)−zm​(vm​e​m)\gamma_{m}=z_{m}(v_{ctxt})-z_{m}(v_{mem}), the posterior odds ratio under SFAD steering satisfies:

P∗​(vc​t​x​t)P∗​(vm​e​m)=PM​(vc​t​x​t)PM​(vm​e​m)⋅exp⁡(λt⋅γe​f​f)\frac{P^{*}(v_{ctxt})}{P^{*}(v_{mem})}=\frac{P_{M}(v_{ctxt})}{P_{M}(v_{mem})}\cdot\exp\left(\lambda_{t}\cdot\gamma_{eff}\right) (12)

where γe​f​f=max⁡(0,zm,vc​t​x​t−zM,vc​t​x​t)\gamma_{eff}=\max(0,z_{m,v_{ctxt}}-z_{M,v_{ctxt}}) is the effective corrective margin.

Proof.

Based on the Bradley-Terry model utilized in DPO, the specialist mm is trained to maximize the log-odds of contextually loyal pairs. Following the asymmetric steering law zt∗​(v)=zM,t​(v)+λt​Δ​zvz_{t}^{*}(v)=z_{M,t}(v)+\lambda_{t}\Delta z_{v}, we examine the log-odds ratio:

log⁡P∗​(vc​t​x​t)P∗​(vm​e​m)\displaystyle\log\frac{P^{*}(v_{ctxt})}{P^{*}(v_{mem})} =(zM,vc​t​x​t+λt​Δ​zvc​t​x​t)\displaystyle=\bigl(z_{M,v_{ctxt}}+\lambda_{t}\Delta z_{v_{ctxt}}\bigr)
−(zM,vm​e​m+λt​Δ​zvm​e​m).\displaystyle\quad-\bigl(z_{M,v_{mem}}+\lambda_{t}\Delta z_{v_{mem}}\bigr). (13)

In a factuality conflict, vm​e​mv_{mem} is the "surface pattern" favored by MM, thus zm,vm​e​m≤zM,vm​e​mz_{m,v_{mem}}\leq z_{M,v_{mem}}, which implies Δ​zvm​e​m=0\Delta z_{v_{mem}}=0 due to the ReLU activation. Conversely, for the faithful token vc​t​x​tv_{ctxt}, the specialist (having been optimized via DPO) yields zm,vc​t​x​t>zM,vc​t​x​tz_{m,v_{ctxt}}>z_{M,v_{ctxt}}. Thus:

log⁡P∗​(vc​t​x​t)P∗​(vm​e​m)\displaystyle\log\frac{P^{*}(v_{ctxt})}{P^{*}(v_{mem})} =log⁡PM​(vc​t​x​t)PM​(vm​e​m)\displaystyle=\log\frac{P_{M}(v_{ctxt})}{P_{M}(v_{mem})}
+λt​(zm,vc​t​x​t−zM,vc​t​x​t).\displaystyle\quad+\lambda_{t}\bigl(z_{m,v_{ctxt}}-z_{M,v_{ctxt}}\bigr). (14)

Exponentiating both sides demonstrates that the faithful candidate’s probability mass is amplified exponentially relative to the hallucination, controlled by the steering intensity λt\lambda_{t}. ∎

Remark 2.

Selective Pressure. This result highlights the advantage of SFAD over standard speculative decoding. While the latter performs binary rejection (0 or 1), SFAD applies continuous selective pressure. Even if the generalist MM is initially biased toward a hallucination, a sufficiently certain specialist can shift the distribution’s mode toward the truth without requiring a complete rejection of the sequence.

E.3 Information-Theoretic Validity of Epistemic Friction

We now justify the use of Epistemic Friction FtF_{t} as a dynamic trigger. It must distinguish between benign diversity and factual conflict.

Lemma 1 (Friction as a High-Precision Trigger).

The Epistemic Friction Ft=𝒟J​S(PM∥Pm)⋅κtF_{t}=\mathcal{D}_{JS}(P_{M}\|P_{m})\cdot\kappa_{t} is a lower bound on the expected reduction in epistemic uncertainty when the generalist conditions its output on the specialist’s contextual evidence.

Proof.

The Jensen-Shannon divergence 𝒟J​S\mathcal{D}_{JS} is a symmetric and bounded metric of distributional tension. In the context of LLMs, high divergence can arise from two sources: (i) factual disagreement or (ii) high entropy (uncertainty) in one model. Standard entropy-based triggers fail because they cannot distinguish between these cases. By weighting 𝒟J​S\mathcal{D}_{JS} with the specialist certainty κt=(1−H⁡(Pm)/log⁡|𝒱|)γ\kappa_{t}=(1-H(P_{m})/\log|\mathcal{V}|)^{\gamma}, FtF_{t} acts as a filter. As κt→1\kappa_{t}\to 1 (certainty), Ft→𝒟J​SF_{t}\to\mathcal{D}_{JS}. As κt→0\kappa_{t}\to 0 (uncertainty), Ft→0F_{t}\to 0, suppressing the trigger. Thus, FtF_{t} only activates steering when the specialist’s corrective signal is both strong and high-confidence, minimizing noise injection. ∎

E.4 Stability and Support Preservation

A critical failure in subtractive contrastive methods (e.g., CAD) is the "Zero-Probability Trap," where valid tokens are suppressed to numerical underflow.

Theorem 2 (Numerical Stability and Support Preservation).

For any finite steering scale λt∈[0,∞)\lambda_{t}\in[0,\infty) and any token v∈𝒱v\in\mathcal{V}, the SFAD framework satisfies the support preservation property: if PM​(v)>0P_{M}(v)>0, then P∗​(v)>0P^{*}(v)>0.

Proof.

The steered logit is zv∗=zM,v+λt⋅ReLU​(zm,v−zM,v)z^{*}_{v}=z_{M,v}+\lambda_{t}\cdot\text{ReLU}(z_{m,v}-z_{M,v}). Since ReLU​(⋅)≥0\text{ReLU}(\cdot)\geq 0 and λt≥0\lambda_{t}\geq 0, it holds that zv∗≥zM,vz^{*}_{v}\geq z_{M,v} for all vv. The steered probability is P∗​(v)=exp⁡(zv∗)/∑wexp⁡(zw∗)P^{*}(v)=\exp(z^{*}_{v})/\sum_{w}\exp(z^{*}_{w}). Since the exponential function is strictly positive, and zv∗z^{*}_{v} is bounded from below by the original logit zM,vz_{M,v}, the numerator remains positive. Unlike subtractive methods where z∗=zM−α​zn​e​gz^{*}=z_{M}-\alpha z_{neg} can lead to −∞-\infty and numerical collapse, our asymmetric law ensures that no token is ever strictly "killed." They are only relatively de-emphasized by the factual amplification of superior candidates. ∎

Remark 3.

Robustness by Design. This theorem explains why SFAD is robust to the hyperparameter λt\lambda_{t}. The additive, ReLU-based nature of the steering ensures that even with high steering intensity, the model maintains its linguistic foundation, avoiding the "repetitive gibberish" or "empty strings" common in negative-constraint decoding.

E.5 Mode-Switching Optimality under Knowledge Conflicts

Finally, we demonstrate why asymmetric steering is superior to linear logit interpolation (averaging).

Theorem 3 (Optimality of Asymmetric Mode-Switching).

In a bimodal conflict between a factual token vfv_{f} and a hallucinated token vhv_{h}, the asymmetric law z∗=zM+λt​(zm−zM)+z^{*}=z_{M}+\lambda_{t}(z_{m}-z_{M})_{+} is a rank-preserving transformation for all neutral tokens vnv_{n} where zm,vn≤zM,vnz_{m,v_{n}}\leq z_{M,v_{n}}.

Proof.

Consider two neutral tokens v1,v2v_{1},v_{2} where the generalist and specialist are in relative agreement (i.e., zm≤zMz_{m}\leq z_{M} for both). Under SFAD, Δ​zv1=Δ​zv2=0\Delta z_{v_{1}}=\Delta z_{v_{2}}=0. Therefore, their steered logits remain zv1∗=zM,v1z^{*}_{v_{1}}=z_{M,v_{1}} and zv2∗=zM,v2z^{*}_{v_{2}}=z_{M,v_{2}}. Their relative rank zv1∗−zv2∗=zM,v1−zM,v2z^{*}_{v_{1}}-z^{*}_{v_{2}}=z_{M,v_{1}}-z_{M,v_{2}} is perfectly preserved. In contrast, a linear interpolation zi​n​t​e​r​p=(1−α)​zM+α​zmz_{interp}=(1-\alpha)z_{M}+\alpha z_{m} would modify both logits, potentially flipping their rank if the specialist has slight fluctuations in its low-probability tail. SFAD thus provides a "surgical" correction: it only perturbs the distribution where the specialist has a strictly better (more factual) proposal. ∎

Remark 4.

Theoretical Scalability. This result suggests that as the Specialist model mm improves in factuality (e.g., through more rigorous DPO), SFAD’s performance will scale without requiring a corresponding increase in the Generalist’s capacity, making it a sustainable architecture for deploying large-scale factual models.

E.6 Optimal Risk-Aware Switching and the Factuality-Efficiency Frontier

A distinctive feature of SFAD is the hybrid policy (Eq. (8)) that dynamically chooses between the Steering Path and the Fast Path. We now provide a decision-theoretic justification for this switching logic, demonstrating that the Epistemic Friction FtF_{t} serves as an optimal proxy for balancing factual integrity and computational latency.

Theorem 4 (Factuality-Risk Minimization).

Let ℛ⁡(π)\mathcal{R}(\pi) be the expected risk of generating a hallucinated token under policy π\pi, and let 𝒞⁡(π)\mathcal{C}(\pi) be the computational cost (latency). Define the factual risk rtr_{t} at step tt as the probability that the draft token x~t\tilde{x}_{t} deviates from the specialist’s high-confidence manifold. The SFAD hybrid policy πS​F​A​D\pi_{SFAD} is a solution to the constrained optimization problem:

minπ⁡𝔼⁡[ℛ⁡(π)]subject to𝔼⁡[𝒞⁡(π)]≤ℬ\min_{\pi}\mathbb{E}[\mathcal{R}(\pi)]\quad\text{subject to}\quad\mathbb{E}[\mathcal{C}(\pi)]\leq\mathcal{B} (15)

where the switching threshold τ\tau acts as the Lagrange multiplier β−1\beta^{-1} that determines the operating point on the Factuality-Efficiency Pareto frontier.

Proof.

Consider a binary decision at each step tt: either accept the standard speculative verification (Fast Path, at=0a_{t}=0) or intervene with logit steering (Steering Path, at=1a_{t}=1). The total risk is:

𝔼⁡[ℛ]=∑t[P⁡(at=0)⋅rt+P⁡(at=1)⋅ϵ]\mathbb{E}[\mathcal{R}]=\sum_{t}\left[P(a_{t}=0)\cdot r_{t}+P(a_{t}=1)\cdot\epsilon\right] (16)

where ϵ\epsilon is the residual risk after steering (which is minimal per Prop. 1). The computational cost is 𝒞f​a​s​t\mathcal{C}_{fast} for the Fast Path and 𝒞s​t​e​e​r\mathcal{C}_{steer} for the Steering Path (𝒞s​t​e​e​r>𝒞f​a​s​t\mathcal{C}_{steer}>\mathcal{C}_{fast}).

By the Neyman-Pearson Lemma, the optimal decision rule ata_{t} that minimizes risk for a fixed cost is a likelihood-ratio test. In SFAD, the Epistemic Friction FtF_{t} quantifies the "hallucination likelihood" by measuring the distributional tension weighted by specialist certainty. When Ft≥τF_{t}\geq\tau, the potential risk reduction Δ​rt=rt−ϵ\Delta r_{t}=r_{t}-\epsilon outweighs the marginal cost Δ​𝒞\Delta\mathcal{C}, triggering the Steering Path. Thus, the switching threshold τ\tau effectively calibrates the sensitivity of the detector to maximize the Factuality Gain per Unit of Latency. ∎

Remark 5.

The Latency-Accuracy Frontier. Theorem 4 formalizes why SFAD maintains the speed of speculative decoding while approaching the factuality of much larger, slow-inference models. By only invoking the "Steering Path" when FtF_{t} signals high epistemic conflict, SFAD avoids the redundant computation of constant logit correction, allowing the system to stay on the optimal Pareto frontier.

E.7 Summary of Theoretical Guarantees

Synthesizing the above proofs, we establish that SFAD is not merely an empirical heuristic but a mathematically grounded framework for verifiable generation.

  • •

    Theorem 1 guarantees that our intervention is linguistically safe (Semantic Integrity).

  • •

    Proposition 1 ensures that we effectively resolve conflicts in favor of the truth (Factuality Amplification).

  • •

    Theorem 2 protects against numerical instability and the suppression of valid tokens (Support Preservation).

  • •

    Theorem 4 justifies the adaptive switching mechanism for real-world deployment (Efficiency Optimality).

Collectively, these theoretical foundations formalize SFAD as a principled safeguarding framework, providing rigorous guarantees for factual integrity through an inherent ’safe-by-design’ decoding paradigm.