arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2609.00543v1 [cs.AI] 01 Sep 2026

Feedback-Assisted Trust Propagation over Document Relation Graphs for Retrieval-Augmented Generation

Zhuoheng Li    Ying Chen Affiliation: College of Information Sciences and Technology, The Pennsylvania State University Email: {zml5515,yingchen}@psu.edu
Abstract

Retrieval-augmented generation (RAG) systems rely on external corpora that may contain outdated, contradictory, noisy, or unreliable documents, introducing reliability risks. Prior work has leveraged document relations to improve the answer reliability of RAG. To propagate reliability signals beyond directly compared document pairs, we propose TrustPropRAG, which structures document relations as a graph and estimates document reliability through multi-hop propagation across the graph. TrustPropRAG anchors this propagation with a limited set of human feedback on document reliability, extending these costly-to-collect feedback-based reliability signals across the whole corpus. Specifically, based on the constructed document relation graph, TrustPropRAG estimates a trust score for each document by formulating and solving an optimization problem that jointly captures pairwise document relations and user feedback. These scores are then used to improve the selection of reliable documents and support trust-aware answer generation. Evaluation results show that TrustPropRAG improves both retrieval quality and exact match over baselines, and remains robust under sparse and noisy feedback.

1 Introduction

Retrieval-augmented generation (RAG) has emerged as an effective approach for mitigating the limitations of large language models (LLMs), whose parametric knowledge may be incomplete or outdated Lewis et al. (2020); Borgeaud et al. (2022); Izacard et al. (2023). By retrieving knowledge from external corpora, LLMs can incorporate relevant and up-to-date information during answer generation. However, RAG’s reliance on external corpora can introduce reliability risks, as the corpora may contain noisy, outdated, or contaminated content (Zhong et al., 2023; Zou et al., 2025), as well as knowledge conflicts (Xie et al., 2024; Chen et al., 2022).

Prior work aims to improve RAG reliability in both the retrieval Ma et al. (2023) and answer generation Wang et al. (2025a) phases. Notably, a line of work leverages relationships among documents to improve the answer reliability of RAG, such as aggregating answers from different documents through majority agreement Xiang et al. (2024), or selecting a subset of documents that are mutually consistent before answer generation Shen et al. (2025). Document relations can be inferred by using natural language inference (NLI) models or leveraging source metadata.

We propose TrustPropRAG, a framework that structures document relations as a graph, as graph representations enable information to propagate through multi-hop relations. Recent studies have explored graph-enhanced RAG, but they mainly focus on organizing knowledge graphs Edge et al. (2024), or constructing index graphs for retrieval Guo et al. (2025). In contrast, we use the graph structure to estimate document reliability and propagate these estimates across the corpus through multi-hop relations. Graph propagation, however, requires anchoring, as document relations reveal consistency or conflict, but not document reliability itself. We anchor this propagation with a set of human feedback on document reliability. In practice, explicit signals (e.g., thumbs-up or thumbs-down ratings) and implicit signals (e.g., clicks or dwell time) can be attributed to the documents that contribute to the generated answer using traceback methods Wang et al. (2025b); Cohen-Wang et al. (2024), thereby providing positive or negative signals for those documents. Since user feedback is costly to collect, TrustPropRAG leverages document relation graphs to propagate sparse feedback across the corpus, similar to how label propagation methods in semi-supervised learning (Zhu et al., 2003; Zhou et al., 2003) spread a few labeled examples across many unlabeled ones. TrustPropRAG then turns the propagated document reliability scores into performance improvements in retrieval and answer generation.

To estimate document reliability and propagate estimates across the corpus, TrustPropRAG first constructs a document relation graph, where nodes represent documents and edges capture document relations (e.g., shared source, mutual support, or contradiction). Based on this graph, TrustPropRAG estimates a trust score for each document by formulating and solving an optimization problem. The problem includes two objective function terms to jointly capture pairwise document relations and incorporate user feedback. At query time, these trust scores are used for trust-aware rescoring, which helps select more reliable documents, and for trust-aware prompting, where the LLM is informed of document reliability during answer generation. Trust score optimization is performed offline over the corpus, amortizing its computational cost across future queries.

Our contributions are threefold. First, we propose a framework that estimates document trust scores by propagating human feedback over a document relation graph, and uses these scores to improve both retrieval and answer generation. Second, we formulate and solve an optimization problem for trust score estimation. We also provide a theoretical analysis of how feedback influences trust score propagation, as well as the convergence rate of the propagation. Third, we evaluate TrustPropRAG with three open-domain question-answering (QA) benchmarks, three retrievers, and six LLMs, and further test its robustness under sparse and noisy feedback. Our code is publicly available at https://github.com/zhliOvO/TrustPropRAG.

2 Related Work

RAG reliability. A growing body of work has studied how to improve RAG reliability. For example, InstructRAG guides LLMs to denoise retrieved contexts using rationales Wei et al. (2025). AstuteRAG compares retrieved evidence with the LLM’s parametric knowledge to handle knowledge conflicts (Wang et al., 2025a). A notable line of work uses relationships (agreements or contradictions) among documents to improve reliability. For instance, RobustRAG isolates retrieved documents, generates per-document answers independently, and aggregates them through majority agreement Xiang et al. (2024). TrustRAG leverages similarity relations among retrieved documents by clustering them and then filtering out contaminated documents before generation Zhou et al. (2025). ReliabilityRAG detects contradictions between documents, selects a subset of documents that are mutually consistent, and produces the final answer via keyword aggregation (Shen et al., 2025). However, these relationships are considered within a subset of documents, and reliability is not propagated beyond the directly compared documents. To move beyond this one-hop relation, TrustPropRAG structures document relations as a graph and captures multi-hop relations among documents through the graph representation.

Graph-empowered RAG. Graphs have long been used to model relationships among textual evidence in tasks such as multi-document summarization (Christensen et al., 2013) and fact verification (Zhong et al., 2020). Recent work has also incorporated graph structures into RAG. For example, GraphRAG organizes corpora with entity knowledge graphs, and leverages community summaries for all groups of closely related entities Edge et al. (2024). LightRAG incorporates graph structures into text indexing and relevant information retrieval for efficient RAG (Guo et al., 2025). NodeRAG introduces heterogeneous graph structures to better leverage the structural nature of graphs Xu et al. (2025). However, these methods primarily use graphs to organize knowledge and guide retrieval. In contrast, we use graph structure to propagate document reliability estimates across the corpus that may contain noisy, outdated, or contradictory information.

Human feedback for enhancing LLMs. Human feedback has been used to align LLM systems with user preferences or intents. For example, InstructGPT trains language models with human feedback using reinforcement learning (Ouyang et al., 2022). RAG-Reward uses preference-based reward models to improve the quality of generated answers (Zhang et al., 2025a). Pistis-RAG aligns human feedback by training a ranking model to improve content ranking and retrieval mechanisms (Bai et al., 2024). Some methods further study feedback in interactive settings. For example, Liu et al. (2025b) leverages implicit feedback in human–LLM dialogues. Zhang et al. (2025b) leverages engagement and disengagement signals for improving LLM-based recommenders. These works mainly use feedback to fine-tune model behavior or improve outputs in human-LLM interactions. TrustPropRAG instead uses limited feedback to estimate document reliability across the large-scale corpus without model fine-tuning.

3 Methodology

Our objective is to enhance RAG by selecting and using more reliable documents from an external corpus for answer generation. The external document corpus may contain unreliable documents, such as those with misinformation, outdated information, or conflicting claims (Xu et al., 2024; Zou et al., 2025). User feedback can indicate whether certain documents are reliable or unreliable. For example, correct and incorrect RAG-generated answers, reported by users or flagged by fact-verification systems, can be traced back to their contributing documents (e.g., using traceback methods Wang et al. (2025b); Cohen-Wang et al. (2024)), providing positive or negative feedback for the corresponding documents. However, user feedback can cover only a small fraction of a large external corpus. To address this, TrustPropRAG takes advantage of document relations to propagate sparse feedback across the corpus.

As shown in Figure 1, TrustPropRAG consists of four main stages. First, we construct a graph where each node represents a document, and each edge encodes a relation between two documents, such as a same-source relation, mutual support, or contradiction (Section 3.1). Second, based on the document relation graph, we formulate and solve optimization problems for estimating trust scores (Section 3.2). We estimate a trust score τi∈[0,1]\tau_{i}\in[0,1] for each document did_{i} by jointly optimizing these scores over all documents in the corpus 𝒞={d1,d2,…,dN}\mathcal{C}=\{d_{1},d_{2},\ldots,d_{N}\} with NN documents. Third, we perform trust-aware rescoring (Section 3.3). Given a user query qq, the objective is to produce a ranked list ℛk​(q)\mathcal{R}_{k}(q) of kk documents that are provided to the LLM for answer generation, favoring query-relevant documents with higher trust scores. Finally, we perform trust-aware prompting (Section 3.4) by leveraging document trust scores during answer generation.

TrustPropRAG separates offline trust estimation from query-time answer generation. Graph construction and trust score optimization are performed offline over the corpus, amortizing their computational cost across future queries, while trust-aware rescoring and prompting are performed for each query with limited additional overhead.

Refer to caption
Figure 1: Overview of the TrustPropRAG pipeline.

3.1 Document Relation Graph Construction

Given the corpus 𝒞\mathcal{C}, we construct a document relation graph G=(𝒞,E)G=(\mathcal{C},E), where each node represents a document, and each edge ei​je_{ij} encodes a relation between documents did_{i} and djd_{j}. Each edge ei​je_{ij} is assigned a relation label ℓi​j\ell_{ij}. Specifically, ℓi​j=+1\ell_{ij}=+1 means that the two documents are expected to receive similar trust scores, ℓi​j=−1\ell_{ij}=-1 means that they contain conflicting claims and should not both be highly trusted, and ℓi​j=0\ell_{ij}=0 means that no reliable relation can be identified. Each edge also has a confidence weight wi​j∈[0,1]w_{ij}\in[0,1], which reflects the confidence associated with the relation label ℓi​j\ell_{ij}.

The relation label and confidence weight can be obtained in different ways. For example, we can use source metadata to connect two documents did_{i} and djd_{j} from the same source group with supportive edges (ℓi​j=+1\ell_{ij}=+1), because they are expected to have correlated trustworthiness and thus similar trust scores. Document relations can also be inferred using an NLI model, which derives whether documents did_{i} and djd_{j} are supportive (ℓi​j=+1\ell_{ij}=+1), contradictory (ℓi​j=−1\ell_{ij}=-1), or neutral (ℓi​j=0\ell_{ij}=0). For edges derived by NLI, the confidence weight wi​jw_{ij} is set to the softmax probability associated with the relation label.

3.2 Trust Score Optimization

Pairwise consistency loss. We encourage two documents with positive relation labels to have similar trust scores, while enforcing that contradictory documents have diverging trust scores. Specifically, if did_{i} and djd_{j} are mutually supportive (ℓi​j=+1\ell_{ij}=+1), their trust scores τi\tau_{i} and τj\tau_{j} should be close to each other, i.e., we aim to minimize (τi−τj)2(\tau_{i}-\tau_{j})^{2}; whereas if ℓi​j=−1\ell_{ij}=-1, meaning that did_{i} and djd_{j} are contradictory, their trust scores should be far apart so that at most one document can be trusted, i.e., we aim to minimize (τi+τj−1)2(\tau_{i}+\tau_{j}-1)^{2}. Hence, we define the pairwise consistency loss ℒpair\mathcal{L}_{\text{pair}}:

ℒpair=∑ei​j∈Eℓi​j=+1wi​j​(τi−τj)2+∑ei​j∈Eℓi​j=−1wi​j​(τi+τj−1)2.\mathcal{L}_{\mathrm{pair}}=\!\!\!\sum_{\begin{subarray}{c}e_{ij}\in E\\ \ell_{ij}=+1\end{subarray}}\!\!\!w_{ij}(\tau_{i}-\tau_{j})^{2}\;+\!\!\!\sum_{\begin{subarray}{c}e_{ij}\in E\\ \ell_{ij}=-1\end{subarray}}\!\!\!w_{ij}(\tau_{i}+\tau_{j}-1)^{2}.

Feedback loss. Over time, we maintain a feedback set ℱ=ℱ+∪ℱ−\mathcal{F}=\mathcal{F}^{+}\cup\mathcal{F}^{-}: ℱ+\mathcal{F}^{+} for documents marked as reliable through user feedback, and ℱ−\mathcal{F}^{-} for documents marked as unreliable. We incorporate this feedback into the optimization problem of trust score estimation. Nodes in ℱ+\mathcal{F}^{+} are encouraged to have high trust scores, e.g., by minimizing (τi−τ+)2(\tau_{i}-\tau^{+})^{2}, where τ+\tau^{+} is close to 11. In contrast, nodes in ℱ−\mathcal{F}^{-} are pushed toward lower trust scores, e.g., by minimizing (τi−τ−)2(\tau_{i}-\tau^{-})^{2}, where τ−\tau^{-} is close to 00, to limit their influence on answer generation. We implement this by adding a feedback term ℒfb\mathcal{L}_{\mathrm{fb}} to the optimization problem of trust score estimation. ℒfb\mathcal{L}_{\mathrm{fb}} is formulated as

ℒfb=∑di∈ℱ(τi−yi)2,yi={τ+,if ​di∈ℱ+τ−,if ​di∈ℱ−,\mathcal{L}_{\mathrm{fb}}=\sum_{d_{i}\in\mathcal{F}}(\tau_{i}-y_{i})^{2},\quad y_{i}=\begin{cases}\tau^{+},&\text{if }d_{i}\in\mathcal{F}^{+}\\ \tau^{-},&\text{if }d_{i}\in\mathcal{F}^{-}\end{cases},

where yiy_{i} is the feedback value.

Formulating trust score optimization problem. To capture pairwise consistency of documents and incorporate user feedback, we jointly optimize objective function terms ℒpair\mathcal{L}_{\mathrm{pair}} and ℒfb\mathcal{L}_{\mathrm{fb}}. We formulate the optimization problem to assign trust scores to all documents in the corpus:

𝝉∗=arg⁡min𝝉∈[0,1]N⁡ℒpair+λf​ℒfb,\boldsymbol{\tau}^{*}\;=\;\arg\min_{\boldsymbol{\tau}\in[0,1]^{N}}\;\;\mathcal{L}_{\mathrm{pair}}\;+\;\lambda_{f}\,\mathcal{L}_{\mathrm{fb}}, (1)

where 𝝉={τ1,τ2,⋯,τN}\boldsymbol{\tau}=\{\tau_{1},\tau_{2},\cdots,\tau_{N}\} is the set of trust scores for all documents, and λf\lambda_{f} controls the trade-off between the two terms.

Solving trust score optimization problem. Eq. (1) is a quadratic programming (QP) problem. We solve the problem with projected gradient descent (PGD). The document relation graph can contain many nodes but remains sparse (i.e., each node has supportive or contradictory relation labels with only a small number of other nodes). As a result, each PGD step only needs to compute the loss and gradients for document pairs connected by these relations, rather than all possible document pairs. Section 4.2 reports the measured cost of each stage in TrustPropRAG.

Before running PGD, we initialize each document’s trust score using its relations with other documents in the document relation graph. Specifically, for each document did_{i}, we compute an initial score by aggregating the signed weights of its connected edges: init⁡(di)=∑(i,j)∈Eℓi​j⋅wi​j\mathrm{init}(d_{i})=\sum_{(i,j)\in E}\ell_{ij}\cdot w_{ij}. Values are min-max normalized to [τmin,τmax][\tau_{\text{min}},\tau_{\text{max}}], with 0<τmin<0.5<τmax<10<\tau_{\text{min}}<0.5<\tau_{\text{max}}<1. Intuitively, documents connected to others through stronger supportive relations receive higher initial trust scores, while documents involved in stronger contradictory relations receive lower initial trust scores. Graph construction and trust score optimization are performed offline over the corpus, amortizing their computational cost across future queries.

Theoretical analysis. Trust scores exhibit multi-hop propagation. Feedback on a document shifts its trust score; through pairwise document relations, this effect then spreads to its neighbors, which further propagate the influence to their neighbors. We formalize this behavior in two lemmas below. Since ℒ=ℒpair+λf​ℒfb\mathcal{L}=\mathcal{L}_{\mathrm{pair}}\;+\;\lambda_{f}\,\mathcal{L}_{\mathrm{fb}} is quadratic in 𝝉\boldsymbol{\tau}, its Hessian A=∇2ℒA=\nabla^{2}\mathcal{L} is symmetric positive semidefinite. We assume that every connected component of GG contains at least one document with feedback, which is sufficient to make AA strictly positive definite (see proof in Appendix C). Let σmin≤⋯≤σmax\sigma_{\min}\leq\cdots\leq\sigma_{\max} denote its eigenvalues and κ=σmax/σmin\kappa=\sigma_{\max}/\sigma_{\min} its condition number. The unconstrained minimizer satisfies the linear system A​𝝉∗=𝐛A\boldsymbol{\tau}^{*}=\mathbf{b}. We use this unconstrained solution to analyze trust propagation and convergence rate.

Lemma 1 (Trust propagation)

Let f1,…,fmf_{1},\ldots,f_{m} be the feedback documents in ℱ\mathcal{F}. Suppose the feedback value yky_{k} at each fk∈ℱf_{k}\in\mathcal{F} is perturbed with a change of Δ​yfk\Delta y_{f_{k}}. Then the resulting change in the trust score of document did_{i} satisfies

|Δ​τi|≤∑k=1mγk⋅qdG​(i,fk),q=κ−1κ+1,\lvert\Delta\tau_{i}\rvert\leq\sum_{k=1}^{m}\gamma_{k}\cdot q^{\,d_{G}(i,\,f_{k})},\qquad q=\frac{\sqrt{\kappa}-1}{\sqrt{\kappa}+1},

where γk=2​λf​γ​|Δ​yfk|\gamma_{k}=2\lambda_{f}\gamma|\Delta y_{f_{k}}|, γ=(1+κ)2/(2​σmax)=O⁡(1/σmin)\gamma=(1+\sqrt{\kappa})^{2}/(2\sigma_{\max})=O(1/\sigma_{\min}), and dG​(i,j)d_{G}(i,j) is the shortest-path distance between nodes ii and jj in GG.

Proof. See Appendix C.1.

This lemma shows that trust propagation is distance-weighted and cumulative. Each document receives influence from all documents with feedback. Each influence term is weighted by qdG​(i,fk)q^{d_{G}(i,f_{k})}, which depends on its graph distance to the corresponding document with feedback. The result suggests that sparse feedback can still be effective when the documents with feedback are well positioned in the relation graph, allowing their influence to reach other documents through short relation paths.

Lemma 2 (Convergence)

With the fixed step size η⋆=2/(σmin+σmax)\eta^{\star}=2/(\sigma_{\min}+\sigma_{\max}), the projected gradient descent iterates satisfy

∥𝝉(t)−𝝉∗∥≤ρ⋅∥𝝉(t−1)−𝝉∗∥,ρ=κ−1κ+1.\lVert\boldsymbol{\tau}^{(t)}-\boldsymbol{\tau}^{*}\rVert\;\leq\;\rho\cdot\lVert\boldsymbol{\tau}^{(t-1)}-\boldsymbol{\tau}^{*}\rVert,\qquad\rho=\frac{\kappa-1}{\kappa+1}.

Proof. See Appendix C.2.

At each iteration, the error contracts by a factor of at most ρ\rho, so the number of iterations needed to reach an ϵ\epsilon-accurate solution (i.e. ∥𝝉(t)−𝝉∗∥≤ϵ\lVert\boldsymbol{\tau}^{(t)}-\boldsymbol{\tau}^{*}\rVert\leq\epsilon) scales as O⁡(κ​log⁡(1/ϵ))O(\kappa\log(1/\epsilon)). In our experiments, PGD converges within 10001000 iterations on tested datasets.

3.3 Trust-aware Rescoring

We select documents with both high trust scores and high query relevance, and provide them to the LLM as retrieved context for answer generation. To this end, we perform trust-aware rescoring for every document by using its estimated trust score and its similarity score. The similarity score is computed by the retriever, e.g., as a dot product score, to measure how relevant the document is to the query. Specifically, for document did_{i}, we combine its trust score τi\tau_{i} with its normalized similarity score si{s}_{i} and define the trust-aware ranking score as ri=α​τi+(1−α)​sir_{i}=\alpha\tau_{i}+(1-\alpha){s}_{i}, where α∈[0,1]\alpha\in[0,1] controls the trade-off between document trustworthiness and its relevance to the query. We select the top-kk documents with the largest rir_{i} and provide them to the LLM for answer generation.

3.4 Trust-aware Answer Generation

Apart from using trust-aware rescoring for document retrieval, we further leverage the estimated trust scores during answer generation. Specifically, for the top-kk documents with the largest trust-aware ranking scores, we format the prompt so that each document is accompanied by its trust score. We also include an instruction in the prompt to guide the LLM to place greater reliance on documents with higher trust scores during answer generation. This trust-aware prompting complements trust-aware rescoring: while rescoring filters out less reliable documents before answer generation, this trust-aware prompting helps the LLM account for document reliability during answer generation.

4 Evaluation

4.1 Evaluation Setup

Datasets. We evaluate TrustPropRAG with three open-domain QA datasets: MS MARCO (Bajaj et al., 2016), Natural Questions (NQ) (Kwiatkowski et al., 2019), and TriviaQA (Joshi et al., 2017). These datasets are widely adopted in retrieval-augmented and open-domain QA evaluation, and together they cover diverse question types. Following the evaluation settings commonly used in prior RAG reliability and robustness studies (Shen et al., 2025; Zhong et al., 2023; Wei et al., 2025; Wang et al., 2025a), we randomly sample 300 query–answer instances for each dataset. Half of the sampled queries are paired with synthetic contradictory documents generated by following the method in PoisonedRAG (Zou et al., 2025). These documents are written to remain fluent and relevant to the query while supporting answers that conflict with the ground-truth answers. The remaining queries are paired only with factual documents from the original datasets. In this way, we construct the evaluation corpus from the relevant documents associated with the sampled queries, resulting in 805, 893, and 1,037 documents for MS MARCO, NQ, and TriviaQA, respectively. This setup evaluates whether the end-to-end system can produce correct answers when the corpus contains both reliable evidence and fluent but misleading contradictory content.

Retrievers. We evaluate with both sparse and dense retrievers. We use BM25 (Robertson and Zaragoza, 2009), a sparse lexical retriever, and two dense bi-encoder retrievers: Contriever (Izacard et al., 2022) and MiniLM based on all-MiniLM-L6-v2 (Reimers and Gurevych, 2019). For each query, we retrieve k=5k{=}5 documents.

Graph construction. We construct a document relation graph using two types of information: document sources and semantic relations between documents. Documents ii and jj from the same source are connected with supportive edges (ℓi​j=+1\ell_{ij}=+1). In TriviaQA, we determine whether two documents share the same source based on their source Uniform Resource Locators (URLs). For MS MARCO and NQ, which do not provide source metadata, we use simulated source grouping: factual documents are partitioned into multiple source groups, and contradictory documents are partitioned into another set of source groups. Documents in the same group are treated as sharing the same source. We adopt this simulation as a proxy for real source metadata, following recent source-reliability-aware RAG work that simulates sources with varying reliability (Hwang et al., 2025). We also apply an NLI model (nli-deberta-v3-small) to each document pair. If the model predicts that two documents ii and jj support each other, we add a supportive edge (ℓi​j=+1\ell_{ij}=+1). If it predicts that they contradict each other, we add a contradictory edge (ℓi​j=−1\ell_{ij}=-1). Appendix F reports the precision and recall of the NLI-derived edges on all three datasets. Table 4 evaluates the impact of removing the source-group edges to isolate the contribution of the NLI-derived edges.

Feedback construction. We construct the feedback set by randomly sampling a fraction of factual and contradictory documents and simulating feedback labels for them. Sampled factual documents are assigned positive feedback labels, while sampled contradictory documents are assigned negative feedback labels. We refer to this sampling fraction as the feedback ratio (FB); for instance, FB = 30% means that feedback labels are provided for 30% of the documents. As real-world feedback may be imperfect, we evaluate TrustPropRAG under noisy feedback in Figure 2, where a portion of the feedback labels are incorrect.

Trust optimization. For PGD, we use a step size of η=0.01\eta=0.01, run for at most T=1000T=1000 iterations, and stop when the convergence tolerance reaches ϵ=10−6\epsilon=10^{-6}. We set feedback loss weight λf=1.0\lambda_{f}{=}1.0. Trust scores are optimized over all corpus documents, resulting in large-scale optimization problems over 805, 893, and 1,037 documents for MS MARCO, NQ, and TriviaQA, respectively.

Rescoring. We set the trust coefficient to α=0.6\alpha{=}0.6. We also test the impact of α\alpha in Section 4.2. The top-kk documents, where k=5k=5 in our experiments, are provided as retrieved context to the LLM. We also test the impact of kk in Appendix B.

Answer generation. We use proprietary and open-source LLMs, including GPT-4o-mini, GPT-5, Gemini 2.5 Flash, Gemini 2.5 Pro, Llama-3.1-8B-Instruct, and Mistral-7B-Instruct. TrustPropRAG leverages the trust-aware prompt that annotates each document with its trust score (as detailed in Section 3.4).

Evaluation metrics. We adopt exact match (EM) and fact precision@kk (FP@kk) as evaluation metrics. EM is a widely used metric in open-domain QA. EM checks whether the normalized generated answer contains the ground-truth answer, with normalization including lowercasing and the removal of articles and punctuation. FP@kk evaluates retrieval quality, which measures the reliability of the top-ranked documents used for answer generation. Specifically, let 𝒟fact\mathcal{D}_{\mathrm{fact}} and 𝒟contr\mathcal{D}_{\mathrm{contr}} denote the sets of factual and contradictory documents, respectively, and let ℛk​(q)\mathcal{R}_{k}(q) denote the top-kk documents for query qq. FP@kk is defined as the fraction of factual documents among these top-kk documents: FP@​k=|ℛk​(q)∩𝒟fact||ℛk​(q)∩(𝒟fact∪𝒟contr)|.\text{FP@}k=\frac{|\mathcal{R}_{k}(q)\cap\mathcal{D}_{\mathrm{fact}}|}{|\mathcal{R}_{k}(q)\cap(\mathcal{D}_{\mathrm{fact}}\cup\mathcal{D}_{\mathrm{contr}})|}.

Baselines. We compare TrustPropRAG against the following methods: (1) VanillaRAG ranks documents using the relevance scores produced by the retriever and feeds the top-ranked documents to the LLM. (2) InstructRAG (Wei et al., 2025) guides the LLM to better use retrieved evidence during generation. It decomposes retrieved documents into atomic claims, and prompts the LLM to generate a rationale connecting relevant claims to the final answer. (3) AstuteRAG (Wang et al., 2025a) combines the LLM’s internal knowledge with retrieved evidence. It first prompts the LLM to produce an answer-relevant document from its internal knowledge, and then compares this document with the retrieved documents and resolves conflicts before generating the answer. (4) ReliabilityRAG (Shen et al., 2025) uses NLI to identify contradictions among retrieved documents. It then solves a maximum independent set problem to select a subset of documents with no detected pairwise contradictions and uses keyword aggregation to generate the final answer. (5) TrustRAG (Zhou et al., 2025) includes a first-stage filtering procedure that clusters the retrieved documents by embedding similarity and filters out clusters identified as suspicious. We use this first-stage procedure as the baseline.

4.2 Evaluation Results

Dataset Method Contriever BM25 MiniLM
MS MARCO VanillaRAG 0.27 0.24 0.26
InstructRAG 0.35 0.40 0.37
AstuteRAG 0.38 0.39 0.37
ReliabilityRAG 0.42 0.40 0.39
TrustRAGstage 1 0.36 0.35 0.33
TrustPropRAG 0.45 0.47 0.44
NQ VanillaRAG 0.37 0.37 0.39
InstructRAG 0.46 0.48 0.43
AstuteRAG 0.44 0.50 0.42
ReliabilityRAG 0.45 0.50 0.48
TrustRAGstage 1 0.44 0.46 0.41
TrustPropRAG 0.55 0.59 0.59
TriviaQA VanillaRAG 0.57 0.62 0.49
InstructRAG 0.68 0.67 0.69
AstuteRAG 0.70 0.71 0.72
ReliabilityRAG 0.75 0.72 0.67
TrustRAGstage 1 0.66 0.65 0.63
TrustPropRAG 0.78 0.75 0.77
Table 1: EM averaged over GPT-4o-mini, Gemini 2.5 Flash, and GPT-5 (k=5k{=}5, FB=30%). Best results are bolded; second-best results are underlined.
Dataset Retriever Baselines TrustPropRAG
MS MARCO Contriever 44% 77%
BM25 57% 92%
MiniLM 54% 88%
NQ Contriever 53% 79%
BM25 46% 88%
MiniLM 48% 96%
TriviaQA Contriever 38% 77%
BM25 57% 90%
MiniLM 49% 92%
Table 2: Retrieval quality measured by FP@5 (k=5k{=}5, FB=30%).

TrustPropRAG outperforms baselines on EM and FP@5. Tables 1 and 2 report the EM and FP@5 averaged over GPT-4o-mini, Gemini 2.5 Flash, and GPT-5. Table 1 shows that TrustPropRAG achieves the best EM across all datasets and retrievers, with absolute EM gains of 0.03–0.11 compared with the strongest baseline. Table 2 shows FP@5 (the fraction of factual documents among top-kk documents) of TrustPropRAG and baselines. All baselines have the same FP@5 as VanillaRAG under this retrieval-level metric: InstructRAG, AstuteRAG, and TrustRAGstage 1{}_{\text{stage 1}} operate after the retriever selects the top-kk documents based on query relevance and thus do not change the ranked top-kk documents; although ReliabilityRAG filters documents before final answer generation, its filtering is driven by intermediate LLM-generated answers and their consistency, rather than by producing a new top-kk list. The baseline methods have an FP@5 of 38–57%, while TrustPropRAG raises it to 77–96%. The improvement is observed with all three retrievers, indicating that the method is not tied to a specific retriever. These improvements in FP@5 help explain the EM gains of TrustPropRAG. Specifically, a higher FP@5 indicates that the top-kk list contains a larger fraction of factual documents, which suggests that trust score optimization assigns higher trust to more reliable documents, and trust-aware rescoring promotes them into the top-kk list. As a result, the LLM inputs contain more documents that support the correct answer.

Ablation study. We conduct an ablation study to examine the contribution of each component in TrustPropRAG. Table 3 evaluates five variants on MS MARCO with FB=30% using Gemini 2.5 Flash. We observe that the combination of trust score optimization and trust-aware rescoring leads to a substantial improvement. Compared with VanillaRAG, adding these two components increases FP@5 from 54% to 87% and improves EM from 0.28 to 0.41. This suggests that the optimized trust scores are effective for identifying more trustworthy documents and that rescoring can promote them into the top-kk context used for answer generation. We also observe that trust-aware prompting provides an additional benefit when combined with trust score optimization and trust-aware rescoring, increasing EM from 0.41 to 0.46. This suggests that some low-trust documents may still remain in the top-kk list after rescoring, and trust-aware prompting helps the LLM reduce their influence during answer generation. Finally, we observe that using trust-aware prompting alone does not improve performance. Without trust score optimization, FP@5 remains at 54%, and EM slightly decreases from 0.28 to 0.27 compared with VanillaRAG. This suggests the trust-aware prompt becomes useful only when the trust scores have been optimized and can provide a meaningful signal for answer generation.

Variant Opt. Rescore Prompt EM FP@5
VanillaRAG – – – 0.28 54%
Opt.+Prompt ✓ – ✓ 0.35 54%
Opt.+Rescore ✓ ✓ – 0.41 87%
Prompt only – – ✓ 0.27 54%
Full TrustPropRAG ✓ ✓ ✓ 0.46 88%
Table 3: Effects of different TrustPropRAG components. Opt: trust score optimization, Rescore: trust-aware rescoring, and Prompt: trust-aware prompting.
Variant EM FP@5
VanillaRAG 0.31 51%
TrustPropRAG (trivial feedback) 0.36 62%
TrustPropRAG (NLI-only) 0.44 79%
TrustPropRAG 0.49 92%
Table 4: Effects of trust propagation and source-group edges (MS MARCO and NQ, MiniLM, GPT-4o-mini, k=5k{=}5, FB=30%).
Model MS MARCO NQ TriviaQA Δ¯\overline{\Delta}
VanillaRAG TrustPropRAG VanillaRAG TrustPropRAG VanillaRAG TrustPropRAG
Proprietary
GPT-4o-mini 0.21 0.40 0.41 0.58 0.50 0.76 +0.21
GPT-5 0.30 0.49 0.43 0.62 0.53 0.85 +0.23
Gemini 2.5 Flash 0.28 0.46 0.33 0.56 0.44 0.69 +0.22
Gemini 2.5 Pro 0.32 0.48 0.37 0.60 0.46 0.77 +0.23
Open-source
Llama-3.1-8B-Instruct 0.19 0.32 0.38 0.55 0.48 0.77 +0.20
Mistral-7B-Instruct 0.22 0.35 0.33 0.47 0.51 0.74 +0.17
Table 5: Performance with different LLMs across three datasets. Δ¯\overline{\Delta} denotes TrustPropRAG’s absolute EM improvement over VanillaRAG, averaged across the three datasets.
Figure 2: Impact of feedback ratio and feedback accuracy on EM (MS MARCO, MiniLM, Gemini 2.5 Flash).

The effects of trust propagation and source-group edges. To separate the benefit of trust propagation from simply having access to feedback, the trivial feedback variant uses the same feedback set as TrustPropRAG but does not propagate the feedback through the graph. Documents with positive feedback receive a boost to their retrieval scores, while documents with negative feedback are not included in the top-kk list. To isolate the contribution of the NLI-derived edges from source-group edges, the NLI-only variant removes all simulated source-group edges and retains only the NLI-derived edges.

From Table 4, the trivial feedback variant achieves an EM of 0.36 and an FP@5 of 62%. This confirms that directly using feedback is beneficial. However, its performance remains substantially below TrustPropRAG, which achieves an EM of 0.49 and an FP@5 of 92%, respectively. This gap shows that the gain of TrustPropRAG cannot be attributed only to access to the feedback set. Propagating feedback through document relation graphs provides trust estimates for documents without direct feedback. Table 4 also shows that the NLI-only variant achieves an EM of 0.44 and an FP@5 of 79%. Although removing the source-group edges leads to a performance drop, the NLI-only variant still outperforms the trivial feedback variant, showing that propagation remains beneficial even when the graph contains only NLI-derived relations.

Performance under noisy feedback. Figure 2 evaluates the robustness of TrustPropRAG by varying the accuracy of feedback labels from 60% to 100%. Feedback accuracy denotes the fraction of feedback labels that are correct. For example, 80% feedback accuracy means that 20% of the feedback labels are flipped: a factual document may be labeled as low-trust, or a contradictory document may be labeled as high-trust. Figure 2 reports EM on MS MARCO with Gemini 2.5 Flash. We observe that TrustPropRAG is robust to noisy feedback. Even when feedback accuracy is only 60% (i.e., 40% of feedback labels are incorrect), TrustPropRAG still improves EM over the setting without feedback across all feedback ratios. This is because individual feedback errors can be mitigated through trust propagation over the document relation graph, where relations among documents help mitigate the impact of noisy feedback signals.

Figure 3: Impact of trust coefficient α\alpha on EM and FP@5. (MS MARCO, MiniLM, Gemini 2.5 Flash, FB=30%)

The impact of feedback ratio. Figure 2 also shows EM as a function of the feedback ratio on MS MARCO with Gemini 2.5 Flash. Compared with the setting without feedback, TrustPropRAG improves sharply when only 10% of documents receive feedback, after which the improvement becomes more gradual. This indicates that a small portion of feedback is sufficient for trust propagation to produce reliable trust estimates across the document relation graph.

The impact of trust coefficient α\alpha. Figure 3 shows EM and FP@5 as the trust coefficient α\alpha varies from 0 to 1. When α=0\alpha=0, rescoring uses only the retrieval relevance score. As α\alpha increases from 00, both EM and FP@5 improve, showing that incorporating trust scores helps promote documents that support the correct answer into the top-kk context. EM reaches its best value at α=0.6\alpha=0.6, and remains stable between 0.50.5 and 0.80.8. However, when α\alpha is too large, the top-kk list tends to include documents with high trust scores but lower query relevance, leading to a slight decrease in EM. As TrustPropRAG is not sensitive to the exact choice of α\alpha within a range (e.g., between 0.5 and 0.8), we set α=0.6\alpha{=}0.6 for the rest of the experiments.

Performance with different LLMs. We evaluate the performance of TrustPropRAG with a diverse set of proprietary and open-source LLMs of different model scales in Table 5. Since FP@5 depends only on retrieval and rescoring, it is unchanged across LLMs; therefore, we report EM as the main metric. We use MiniLM as the retriever and set k=5k=5 and FB=30%. Table 5 shows that TrustPropRAG consistently outperforms VanillaRAG across all evaluated LLMs. The average EM improvement ranges from 0.17 to 0.23 across both proprietary models (GPT-4o-mini, GPT-5, Gemini 2.5 Flash, Gemini 2.5 Pro) and smaller open-source models (Llama-3.1-8B-Instruct, Mistral-7B-Instruct). Notably, TrustPropRAG still achieves an average EM gain of 0.23 with strong LLMs such as GPT-5 and Gemini 2.5 Pro. This suggests that improving the reliability of the retrieved context with trust propagation remains useful even for strong LLMs.

Computational cost. Table 6 reports the computational cost of TrustPropRAG, including offline graph construction and trust optimization, as well as query-time retrieval. The computational resources are detailed in Appendix A. Constructing the graph accounts for most of the offline cost, taking 578.3–1077.1 seconds. Given the constructed graph, trust score optimization is efficient, requiring only 0.3–0.4 seconds, indicating that trust propagation itself introduces negligible computational overhead. Both graph construction and trust optimization are performed offline and their costs are therefore amortized across future queries. At query time, retrieval takes 4.4–4.8 seconds and is shared by TrustPropRAG and the baselines. Beyond retrieval, TrustPropRAG only performs lightweight trust-aware rescoring using the precomputed trust scores, introducing limited additional query-time overhead.

Stage MS MARCO NQ TriviaQA
Offline
Graph construction 578.3 s 852.4 s 1077.1 s
Trust optimization (PGD) 0.3 s 0.3 s 0.4 s
Query time
Retrieval 4.4 s 4.8 s 4.8 s
Table 6: Computational cost breakdown.

5 Conclusions

We present TrustPropRAG, a framework that improves RAG reliability by propagating sparse human feedback over a document relation graph. TrustPropRAG estimates a trust score for each document by optimizing an objective that combines pairwise consistency loss and feedback loss. The resulting trust scores are used to select more reliable documents and guide trust-aware answer generation. TrustPropRAG achieves absolute EM gains of 0.03–0.11 over the strongest baseline, and demonstrates robustness to noisy feedback.

Acknowledgments

We thank the anonymous reviewers for insightful feedback. This work was supported by Seed Grant of IST, and the National Science Foundation under Grants 2550742, 2623125, 2555329, and 2549266.

Limitations

TrustPropRAG incorporates human feedback as binary labels, marking documents as either reliable or unreliable. In practice, feedback can be graded rather than binary. Studying TrustPropRAG under graded feedback is left to future work. Following common practice in RAG evaluation, we focus on unreliable content that is injected as synthetic contradictions rather than drawn from naturally occurring conflicts. Finally, we focus on open-domain QA with short factual answers. Applying trust propagation to long-form or multi-hop generation, where reliability interacts with compositional reasoning, is left to future work.

Ethical Considerations

This work addresses the problem of unreliable content in retrieval-augmented generation. Our experiments use synthetically generated contradictions. All datasets and models used in this work are publicly released and licensed for research use, and our usage is consistent with their intended research purposes. MS MARCO, NQ, and TriviaQA are publicly available open-domain QA benchmarks released for research. The retrievers (BM25, Contriever, and the all-MiniLM-L6-v2 encoder) and the NLI model (nli-deberta-v3-small) are also publicly released.

References

  • Bai et al. (2024) Y. Bai, Y. Miao, L. Chen, D. Wang, D. Li, Y. Ren, H. Xie, C. Yang, and X. Cai Pistis-RAG: enhancing retrieval-augmented generation with human feedback. arXiv preprint arXiv:2407.00072. Cited by: §2.
  • Bajaj et al. (2016) P. Bajaj, D. Campos, N. Craswell, L. Deng, J. Gao, X. Liu, R. Majumder, A. McNamara, B. Mitra, T. Nguyen, M. Rosenberg, X. Song, A. Stoica, S. Tiwary, and T. Wang MS MARCO: a human generated machine reading comprehension dataset. In Proceedings of the Workshop on Cognitive Computation: Integrating Neural and Symbolic Approaches (CoCo@NIPS), Cited by: §4.1.
  • Benzi and Razouk (2007) M. Benzi and N. Razouk Decay bounds and O(nn) algorithms for approximating functions of sparse matrices. Electronic Transactions on Numerical Analysis 28, pp. 16–39. Cited by: §C.1.
  • Borgeaud et al. (2022) S. Borgeaud, A. Mensch, J. Hoffmann, T. Cai, E. Rutherford, K. Millican, G. van den Driessche, J. Lespiau, B. Damoc, A. Clark, D. de Las Casas, A. Guy, J. Menick, R. Ring, T. Hennigan, S. Huang, L. Maggiore, C. Jones, A. Cassirer, A. Brock, M. Paganini, G. Irving, O. Vinyals, S. Osindero, K. Simonyan, J. W. Rae, E. Elsen, and L. Sifre Improving language models by retrieving from trillions of tokens. In Proceedings of the 39th International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, Vol. 162, pp. 2206–2240. Cited by: §1.
  • Chen et al. (2022) H. Chen, M. J.Q. Zhang, and E. Choi Rich knowledge sources bring complex knowledge conflicts: recalibrating models to reflect conflicting evidence. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, pp. 2292–2307. Cited by: §1.
  • Christensen et al. (2013) J. Christensen, Mausam, S. Soderland, and O. Etzioni Towards coherent multi-document summarization. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Atlanta, Georgia, pp. 1163–1173. Cited by: §2.
  • Cohen-Wang et al. (2024) B. Cohen-Wang, H. Shah, K. Georgiev, and A. Mądry ContextCite: attributing model generation to context. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 37. Cited by: §1, §3.
  • Demko et al. (1984) S. Demko, W. F. Moss, and P. W. Smith Decay rates for inverses of band matrices. Mathematics of Computation 43 (168), pp. 491–499. Cited by: §C.1, Appendix D.
  • Edge et al. (2024) D. Edge, H. Trinh, N. Cheng, J. Bradley, A. Chao, A. Mody, S. Truitt, D. Metropolitansky, R. O. Ness, and J. Larson From local to global: a graph RAG approach to query-focused summarization. arXiv preprint arXiv:2404.16130. Cited by: §1, §2.
  • Guo et al. (2025) Z. Guo, L. Xia, Y. Yu, T. Ao, and C. Huang LightRAG: simple and fast retrieval-augmented generation. In Findings of the Association for Computational Linguistics: EMNLP 2025, pp. 10746–10761. Cited by: §1, §2.
  • Hwang et al. (2025) J. Hwang, J. Park, H. Park, D. Kim, S. Park, and J. Ok Retrieval-augmented generation with estimation of source reliability. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 34279–34303. Cited by: §4.1.
  • Izacard et al. (2022) G. Izacard, M. Caron, L. Hosseini, S. Riedel, P. Bojanowski, A. Joulin, and E. Grave Unsupervised dense information retrieval with contrastive learning. Transactions on Machine Learning Research. Cited by: §4.1.
  • Izacard et al. (2023) G. Izacard, P. Lewis, M. Lomeli, L. Hosseini, F. Petroni, T. Schick, J. Dwivedi-Yu, A. Joulin, S. Riedel, and E. Grave Few-shot learning with retrieval augmented language models. Journal of Machine Learning Research 24 (251), pp. 1–43. Cited by: §1.
  • Joshi et al. (2017) M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer TriviaQA: a large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Cited by: §4.1.
  • Kwiatkowski et al. (2019) T. Kwiatkowski, J. Palomaki, O. Redfield, M. Collins, A. Parikh, C. Alberti, D. Epstein, I. Polosukhin, J. Devlin, K. Lee, K. Toutanova, L. Jones, M. Kelcey, M. Chang, A. M. Dai, J. Uszkoreit, Q. Le, and S. Petrov Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics 7, pp. 452–466. Cited by: §4.1.
  • Lewis et al. (2020) P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. Yih, T. Rocktäschel, S. Riedel, and D. Kiela Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems, Vol. 33. Cited by: §1.
  • Liu et al. (2025a) S. Liu, Q. Ning, K. Halder, Z. Qi, W. Xiao, P. M. Htut, Y. Zhang, N. A. John, B. Min, Y. Benajiba, et al. Open domain question answering with conflicting contexts. In Findings of the Association for Computational Linguistics: NAACL 2025, pp. 1838–1854. Cited by: Appendix G.
  • Liu et al. (2025b) Y. Liu, M. J. Zhang, and E. Choi User feedback in human-LLM dialogues: a lens to understand users but noisy as a learning signal. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 2666–2681. Cited by: §2.
  • Ma et al. (2023) X. Ma, Y. Gong, P. He, H. Zhao, and N. Duan Query rewriting in retrieval-augmented large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 5303–5315. Cited by: §1.
  • Ouyang et al. (2022) L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. L. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. Christiano, J. Leike, and R. Lowe Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, Vol. 35. Cited by: §2.
  • Reimers and Gurevych (2019) N. Reimers and I. Gurevych Sentence-BERT: sentence embeddings using siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, pp. 3982–3992. Cited by: §4.1.
  • Robertson and Zaragoza (2009) S. Robertson and H. Zaragoza The probabilistic relevance framework: BM25 and beyond. Foundations and Trends in Information Retrieval 3 (4), pp. 333–389. Cited by: §4.1.
  • Shen et al. (2025) Z. Shen, B. Imana, T. Wu, C. Xiang, P. Mittal, and A. Korolova ReliabilityRAG: effective and provably robust defense for RAG-based web-search. In Advances in Neural Information Processing Systems, Cited by: Appendix B, §1, §2, §4.1, §4.1.
  • Wang et al. (2025a) F. Wang, X. Wan, R. Sun, J. Chen, and S. Ö. Arik Astute RAG: overcoming imperfect retrieval augmentation and knowledge conflicts for large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria, pp. 30553–30571. Cited by: §1, §2, §4.1, §4.1.
  • Wang et al. (2025b) Y. Wang, W. Zou, R. Geng, and J. Jia TracLLM: a generic framework for attributing long context llms. In USENIX Security Symposium, Cited by: §1, §3.
  • Wei et al. (2025) Z. Wei, W. Chen, and Y. Meng InstructRAG: instructing retrieval-augmented generation via self-synthesized rationales. In Proceedings of the Thirteenth International Conference on Learning Representations (ICLR), Cited by: §2, §4.1, §4.1.
  • Xiang et al. (2024) C. Xiang, T. Wu, Z. Zhong, D. Wagner, D. Chen, and P. Mittal Certifiably robust RAG against retrieval corruption. arXiv preprint arXiv:2405.15556. Cited by: §1, §2.
  • Xie et al. (2024) J. Xie, K. Zhang, J. Chen, R. Lou, and Y. Su Adaptive chameleon or stubborn sloth: revealing the behavior of large language models in knowledge conflicts. In Proceedings of the Twelfth International Conference on Learning Representations (ICLR), Cited by: §1.
  • Xu et al. (2024) R. Xu, Z. Qi, Z. Guo, C. Wang, H. Wang, Y. Zhang, and W. Xu Knowledge conflicts for LLMs: a survey. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, Florida, USA, pp. 8541–8565. Cited by: §3.
  • Xu et al. (2025) T. Xu, H. Zheng, C. Li, H. Chen, Y. Liu, R. Chen, and L. Sun NodeRAG: structuring graph-based rag with heterogeneous nodes. arXiv preprint arXiv:2504.11544. Cited by: §2.
  • Zhang et al. (2025a) H. Zhang, J. Song, J. Zhu, Y. Wu, T. Zhang, and C. Niu RAG-Reward: optimizing RAG with reward modeling and RLHF. arXiv preprint arXiv:2501.13264. Cited by: §2.
  • Zhang et al. (2025b) J. Zhang, C. Gao, W. Shi, X. Chen, J. Wang, X. Cai, and F. Feng Leveraging unpaired feedback for long-term LLM-based recommendation tuning. In Findings of the Association for Computational Linguistics: EMNLP 2025, pp. 24507–24521. Cited by: §2.
  • Zhong et al. (2020) W. Zhong, J. Xu, D. Tang, Z. Xu, N. Duan, M. Zhou, J. Wang, and J. Yin Reasoning over semantic-level graph for fact checking. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 6170–6180. Cited by: §2.
  • Zhong et al. (2023) Z. Zhong, Z. Huang, A. Wettig, and D. Chen Poisoning retrieval corpora by injecting adversarial passages. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Cited by: §1, §4.1.
  • Zhou et al. (2003) D. Zhou, O. Bousquet, T. N. Lal, J. Weston, and B. Schölkopf Learning with local and global consistency. In Advances in Neural Information Processing Systems, Vol. 16, pp. 321–328. Cited by: §1.
  • Zhou et al. (2025) H. Zhou, K. Lee, Z. Zhan, Y. Chen, Z. Li, Z. Wang, H. Haddadi, and E. Yilmaz TrustRAG: enhancing robustness and trustworthiness in retrieval-augmented generation. arXiv preprint arXiv:2501.00879. Cited by: §2, §4.1.
  • Zhu et al. (2003) X. Zhu, Z. Ghahramani, and J. D. Lafferty Semi-supervised learning using Gaussian fields and harmonic functions. In Proceedings of the 20th International Conference on Machine Learning, pp. 912–919. Cited by: §1.
  • Zou et al. (2025) W. Zou, R. Geng, B. Wang, and J. Jia PoisonedRAG: knowledge corruption attacks to retrieval-augmented generation of large language models. In 34th USENIX Security Symposium (USENIX Security 25), pp. 3827–3844. Cited by: Appendix B, §1, §3, §4.1.

Appendix A Implementation Details

Models and computational resources. Among the LLMs we evaluate, Llama-3.1-8B-Instruct and Mistral-7B-Instruct are open-source models with 8B and 7B parameters, respectively; GPT-4o-mini, GPT-5, Gemini 2.5 Flash, and Gemini 2.5 Pro are proprietary models accessed through their respective APIs. The NLI model (nli-deberta-v3-small) has approximately 44M parameters, and the dense retrievers (Contriever and all-MiniLM-L6-v2) are sentence-encoder models with approximately 110M and 22M parameters respectively. All local computation, including NLI-based edge construction, dense retrieval, and trust score optimization, was run on an NVIDIA RTX 5070 GPU. For our largest corpus (1,037 documents, TriviaQA), this took approximately 0.3 GPU-hours. Trust score optimization via projected gradient descent is lightweight, converging within 1,000 iterations and taking under 1 minute per corpus. We implement retrieval, NLI inference, and trust score optimization using standard Python libraries, and provide exact package versions and scripts in the released code.

Hyperparameters. We use the same hyperparameters across all datasets, retrievers, and LLMs. For trust score initialization, we set τmin=0.3\tau_{\min}=0.3 and τmax=0.7\tau_{\max}=0.7. We set feedback loss weight λf=1.0\lambda_{f}{=}1.0. We set the feedback values τ+=0.9\tau^{+}=0.9 and τ−=0.1\tau^{-}=0.1.

Appendix B Number of Retrieved Documents kk

The number of retrieved documents kk controls the size of the context provided to the LLM. A smaller kk yields a more selective context but may omit relevant evidence, while a larger kk provides broader coverage at the cost of allowing more unreliable documents into the prompt. We evaluate TrustPropRAG across k∈{1,3,5,10,20}k\in\{1,3,5,10,20\} on MS MARCO with MiniLM as the retriever, Gemini 2.5 Flash as the backbone LLM, and FB==30%. All other hyperparameters follow Section 4.

kk=1 kk=3 kk=5 kk=10 kk=20
EM 0.41 0.43 0.46 0.45 0.42
FP@kk 95% 92% 88% 84% 73%
Table 7: Impact of the number of retrieved documents kk on EM and FP@kk (MS MARCO, MiniLM, Gemini 2.5 Flash, FB==30%).

FP@kk falls steadily from 95% (kk=1) to 73% (kk=20) as lower-trust documents enter the top-kk selected documents. EM is less sensitive, staying within 0.41–0.46 and peaking at kk=5 and kk=10. Trust-aware rescoring places reliable evidence at the top, so small kk already provides sufficient factual support; at larger kk, the trust-aware prompt helps the LLM put less weight to lower-trust documents that are provided in the context. EM is smaller at kk=1, where the retrieved context may provide insufficient context, and at kk=20, where the larger retrieval set exposes the LLM to more contradictory context.

Considering that EM peaks at kk=5 and kk=10, we adopt kk=5, matching prior work (Zou et al., 2025; Shen et al., 2025).

Appendix C Proofs

We provide proofs of the two lemmas stated in Section 3.2. Both lemmas rely on the fact that the trust optimization objective ℒ=ℒpair+λf​ℒfb\mathcal{L}=\mathcal{L}_{\mathrm{pair}}+\lambda_{f}\mathcal{L}_{\mathrm{fb}} is a quadratic in 𝝉\boldsymbol{\tau}, so the first-order optimality condition for the unconstrained quadratic problem is given by the linear system A​𝝉∗=𝐛,A\boldsymbol{\tau}^{*}=\mathbf{b}, where A=∇2ℒA=\nabla^{2}\mathcal{L} is the Hessian. We first verify that AA is positive definite under the assumption that every connected component of GG contains at least one document with feedback, then prove the two lemmas.

Positive definiteness of AA.

Let 𝐞i∈ℝN\mathbf{e}_{i}\in\mathbb{R}^{N} denote the ii-th standard basis vector, whose ii-th entry is one and all other entries are zero. Expanding the Hessian gives

A\displaystyle A =∑ei​j∈Eℓi​j=+1wi​j​(𝐞i−𝐞j)​(𝐞i−𝐞j)⊤\displaystyle=2\!\!\sum_{\begin{subarray}{c}e_{ij}\in E\\ \ell_{ij}=+1\end{subarray}}\!\!w_{ij}\,(\mathbf{e}_{i}-\mathbf{e}_{j})(\mathbf{e}_{i}-\mathbf{e}_{j})^{\top}
+∑ei​j∈Eℓi​j=−1wi​j(𝐞i+𝐞j)(𝐞i+𝐞j)⊤\displaystyle\quad+2\!\!\sum_{\begin{subarray}{c}e_{ij}\in E\\ \ell_{ij}=-1\end{subarray}}\!\!w_{ij}\,(\mathbf{e}_{i}+\mathbf{e}_{j})(\mathbf{e}_{i}+\mathbf{e}_{j})^{\top}
+2λf∑i∈ℱ𝐞i𝐞i⊤,\displaystyle\quad+2\lambda_{f}\!\sum_{i\in\mathcal{F}}\mathbf{e}_{i}\mathbf{e}_{i}^{\top}, (2)

which is a sum of positive semidefinite matrices, hence A⪰0A\succeq 0. Next, we prove A≻0A\succ 0. A vector 𝐯\mathbf{v} lies in ker⁡(A)\ker(A) iff (a) vi=vjv_{i}=v_{j} on every supportive edge (ℓi​j=+1\ell_{ij}=+1), (b) vi=−vjv_{i}=-v_{j} on every contradictory edge (ℓi​j=−1\ell_{ij}=-1), and (c) vi=0v_{i}=0 on every node with feedback. If a node with value zero is connected to another node by a supportive edge, condition (a) implies that the neighboring node also has value zero. If it is connected by a contradictory edge, condition (b) again implies that the neighboring node has value zero. Under our assumption that every connected component contains at least one feedback node, we have 𝐯=𝟎\mathbf{v}=\mathbf{0} globally, so A≻0A\succ 0. We write its eigenvalues as σmin≤⋯≤σmax\sigma_{\min}\leq\cdots\leq\sigma_{\max} and its condition number as κ=σmax/σmin\kappa=\sigma_{\max}/\sigma_{\min}.

C.1 Proof of Lemma 1 (Trust Propagation)

We fix the set of documents with feedback ℱ={f1,…,fm}\mathcal{F}=\{f_{1},\ldots,f_{m}\} and consider perturbations of their feedback values. Let Δ​yfk\Delta y_{f_{k}} denote the perturbation applied to the feedback value yfky_{f_{k}}. These perturbations leave AA unchanged and only modify the right-hand side 𝐛\mathbf{b} at entries corresponding to documents in ℱ\mathcal{F}. Let Δ​𝐛\Delta\mathbf{b} denote the resulting change in 𝐛\mathbf{b}, and let Δ​bfk\Delta b_{f_{k}} denote its entry corresponding to feedback document fkf_{k}. We have

Δ​bfk=2​λf​Δ​yfk.\Delta b_{f_{k}}=2\lambda_{f}\Delta y_{f_{k}}. (3)

Step 1: Linearity and superposition.

Because the perturbations only change the feedback target values, the matrix AA remains fixed. Therefore, the unconstrained optimum satisfies 𝝉∗=A−1​𝐛.\boldsymbol{\tau}^{*}=A^{-1}\mathbf{b}. After the perturbation, the right-hand side becomes 𝐛+Δ​𝐛\mathbf{b}+\Delta\mathbf{b}, and the corresponding change in the optimized trust scores is

Δ​𝝉∗=A−1​Δ​𝐛.\Delta\boldsymbol{\tau}^{*}=A^{-1}\Delta\mathbf{b}.

Since Δ​𝐛\Delta\mathbf{b} is nonzero only at documents with feedback, the change in the trust score of document did_{i} can be written as

Δ​τi=∑k=1m(A−1)i​fk​Δ​bfk,\Delta\tau_{i}=\sum_{k=1}^{m}(A^{-1})_{if_{k}}\Delta b_{f_{k}}, (4)

where (A−1)i​fk(A^{-1})_{if_{k}} denotes the entry in the ii-th row and fkf_{k}-th column of A−1A^{-1}.

Step 2: Entry-wise decay of A−1A^{-1}.

The feedback term contributes only to the diagonal entries of AA. The off-diagonal entry Ai​jA_{ij} (i≠ji\neq j) is nonzero when documents ii and jj are joined by an edge (ℓi​j=±1\ell_{ij}=\pm 1). Specifically, from Eq. (2), each pairwise relation term adds ±2​wi​j\pm 2w_{ij} at positions (i,j)(i,j) and (j,i)(j,i) of AA. The sparsity pattern of AA therefore coincides with the adjacency structure of GG. For a symmetric positive definite matrix with this property, the Demko–Moss–Smith decay bound (Demko et al., 1984; Benzi and Razouk, 2007) gives an entry-wise decay of A−1A^{-1} that is exponential in graph distance:

|(A−1)i​j|≤γ​qdG​(i,j),q=κ−1κ+1,\lvert(A^{-1})_{ij}\rvert\;\leq\;\gamma q^{\,d_{G}(i,j)},\qquad q=\frac{\sqrt{\kappa}-1}{\sqrt{\kappa}+1}, (5)

where γ=(1+κ)2/(2​σmax)=O⁡(1/σmin)\gamma=(1+\sqrt{\kappa})^{2}/(2\sigma_{\max})=O(1/\sigma_{\min}).

Step 3: Superposition bound.

Combining Eqs. (3)-(5) and using the triangle inequality, we obtain

|Δ​τi|\displaystyle\lvert\Delta\tau_{i}\rvert ≤∑k=1m|(A−1)i​fk|⋅|Δ​bfk|\displaystyle\leq\sum_{k=1}^{m}\left|(A^{-1})_{if_{k}}\right|\cdot\left|\Delta b_{f_{k}}\right|
≤∑k=1m2​γ​λf​|Δ​yfk|​qdG​(i,fk).\displaystyle\leq\sum_{k=1}^{m}2\gamma\lambda_{f}\lvert\Delta y_{f_{k}}\rvert q^{d_{G}(i,f_{k})}.

C.2 Proof of Lemma 2 (Convergence)

Let 𝝉~(t)=𝝉(t−1)−η∇ℒ(𝝉(t−1))\widetilde{\boldsymbol{\tau}}^{(t)}=\boldsymbol{\tau}^{(t-1)}-\eta\nabla\mathcal{L}\!\left(\boldsymbol{\tau}^{(t-1)}\right) denote the gradient descent step before projection, and let 𝝉(t)=P[0,1]N​(𝝉~(t))\boldsymbol{\tau}^{(t)}=P_{[0,1]^{N}}\!\left(\widetilde{\boldsymbol{\tau}}^{(t)}\right) denote the projected iterate. Since ℒ\mathcal{L} is quadratic with Hessian AA, for any 𝝉\boldsymbol{\tau} we have ∇ℒ​(𝝉)=A⁡(𝝉−𝝉∗)\nabla\mathcal{L}(\boldsymbol{\tau})=A(\boldsymbol{\tau}-\boldsymbol{\tau}^{*}). Therefore,

𝝉~(t)−𝝉∗\displaystyle\widetilde{\boldsymbol{\tau}}^{(t)}-\boldsymbol{\tau}^{*} =𝝉(t−1)−η∇ℒ(𝝉(t−1))−𝝉∗\displaystyle=\boldsymbol{\tau}^{(t-1)}-\eta\nabla\mathcal{L}\!\left(\boldsymbol{\tau}^{(t-1)}\right)-\boldsymbol{\tau}^{*} (6)
=𝝉(t−1)−𝝉∗−η​A​(𝝉(t−1)−𝝉∗)\displaystyle=\boldsymbol{\tau}^{(t-1)}-\boldsymbol{\tau}^{*}-\eta A\left(\boldsymbol{\tau}^{(t-1)}-\boldsymbol{\tau}^{*}\right)
=(I−η​A)​(𝝉(t−1)−𝝉∗).\displaystyle=(I-\eta A)\left(\boldsymbol{\tau}^{(t-1)}-\boldsymbol{\tau}^{*}\right).

With the fixed step size η⋆=2/(σmin+σmax)\eta^{\star}=2/(\sigma_{\min}+\sigma_{\max}), each eigenvalue σ\sigma of AA is mapped to an eigenvalue 1−η⋆​σ1-\eta^{\star}\sigma of I−η⋆​AI-\eta^{\star}A. Since σ∈[σmin,σmax]\sigma\in[\sigma_{\min},\sigma_{\max}], these eigenvalues lie in [−κ−1κ+1,κ−1κ+1].\left[-\frac{\kappa-1}{\kappa+1},\frac{\kappa-1}{\kappa+1}\right]. Thus, the largest absolute eigenvalue of I−η⋆​AI-\eta^{\star}A is ρ=κ���1κ+1.\rho=\frac{\kappa-1}{\kappa+1}. Because I−η⋆​AI-\eta^{\star}A is symmetric, its spectral norm is equal to its largest absolute eigenvalue, so

‖I−η⋆​A‖2=ρ.\left\|I-\eta^{\star}A\right\|_{2}=\rho.

Taking norms in (6) then gives

‖𝝉~(t)−𝝉∗‖≤ρ⁡‖𝝉(t−1)−𝝉∗‖.\left\|\widetilde{\boldsymbol{\tau}}^{(t)}-\boldsymbol{\tau}^{*}\right\|\leq\rho\left\|\boldsymbol{\tau}^{(t-1)}-\boldsymbol{\tau}^{*}\right\|.

The projection operator P[0,1]N​(⋅)P_{[0,1]^{N}}(\cdot) is non-expansive, and hence

∥𝝉(t)−𝝉∗∥\displaystyle\lVert\boldsymbol{\tau}^{(t)}-\boldsymbol{\tau}^{*}\rVert ≤∥𝝉~(t)−𝝉∗∥≤ρ⁡∥𝝉(t−1)−𝝉∗∥.\displaystyle\leq\lVert\widetilde{\boldsymbol{\tau}}^{(t)}-\boldsymbol{\tau}^{*}\rVert\leq\rho\,\lVert\boldsymbol{\tau}^{(t-1)}-\boldsymbol{\tau}^{*}\rVert.

Appendix D Impact of Mislabeled Edges

A mislabeled edge between documents ii and jj changes the Hessian AA in Eq. (2) only at the entries involving ii and jj. Under the same assumptions used in Lemma 1, the entry-wise decay bound on A−1A^{-1} from Demko et al. (1984) implies that the resulting perturbation to the optimized trust scores decreases exponentially with the shortest-path distance from the mislabeled edge. Therefore, the influence of an edge error diminishes rapidly as it propagates to more distant documents.

Appendix E Prompt Templates

Trust-aware prompt.

Each retrieved document is annotated with both its trust score τi\tau_{i} and a trust label. We assign the label based on the trust score: high if τi≥0.7\tau_{i}\geq 0.7, low if τi≤0.3\tau_{i}\leq 0.3, and medium otherwise. The system instruction is:

You are a QA system. Each document has a trust score. Strongly prefer HIGH-trust documents. Ignore or discount LOW-trust documents, they may contain deliberate misinformation. Give a short, direct answer.

Documents are formatted as:

[Document 1] (Trust: HIGH, score=0.92): <text> [Document 2] (Trust: LOW, score=0.15): <text> ...

Baseline prompt. VanillaRAG uses the following prompt. The other baselines use the prompts specified in their original papers.

Answer the following question based on the provided documents. Give a short, direct answer. Documents: [Document 1]: <text> ... Question: <query> Short answer:

Appendix F Quality of NLI-derived Edges

We evaluate the NLI-derived edges against the known factual and contradictory structure of the corpus on the three datasets. We treat an edge between a factual document and a contradictory document as a ground-truth contradictory relation, and an edge between two documents of the same category as a supportive relation. Table 8 reports the precision and recall of the contradictory and supportive edges produced by the NLI model.

Dataset Contradictory Supportive
Precision Recall Precision Recall
MS MARCO 0.65 0.18 0.71 0.14
NQ 0.69 0.21 0.68 0.17
TriviaQA 0.58 0.13 0.62 0.12
Table 8: Precision and recall of the NLI-derived contradictory and supportive edges.

The NLI-derived edges have moderate precision and low recall, as the NLI model predicts most document pairs as neutral. The low recall indicates that many relations are omitted, resulting in a sparse graph, which reflects practical settings where only a limited subset of inter-document relations can be identified. The imperfect precision indicates that the graph also contains some incorrectly labeled edges. Nevertheless, the NLI-only ablation in Table 4 shows that TrustPropRAG still outperforms the baselines when relying only on these imperfect edges, which suggests that trust propagation is robust to noise and sparsity in graph construction.

Appendix G Performance on QACC

We evaluate TrustPropRAG on naturally occurring conflicts without any synthetic injection. We used QACC Liu et al. (2025a), a human-annotated open-domain QA dataset in which unambiguous questions are paired with real web contexts retrieved via Google Search. The conflicts among these contexts arise naturally on the web rather than from injected documents. We randomly sampled 100 questions from QACC and constructed the evaluation corpus from their retrieved contexts. As shown in Table 9, TrustPropRAG achieves the best EM (0.59) on this naturally conflicting corpus, outperforming the strongest baseline ReliabilityRAG (0.56) and improving over VanillaRAG by 0.11. These results indicate that the effectiveness of TrustPropRAG is not limited to synthetically injected contradictions and extends to naturally occurring conflicts.

Method EM
VanillaRAG 0.48
InstructRAG 0.54
AstuteRAG 0.50
ReliabilityRAG 0.56
TrustRAGstage 1 0.50
TrustPropRAG 0.59
Table 9: EM on QACC (GPT-4o-mini, MiniLM, k=5k{=}5, FB=30%).