Feedback-Assisted Trust Propagation over Document Relation Graphs for Retrieval-Augmented Generation
Abstract
Retrieval-augmented generation (RAG) systems rely on external corpora that may contain outdated, contradictory, noisy, or unreliable documents, introducing reliability risks. Prior work has leveraged document relations to improve the answer reliability of RAG. To propagate reliability signals beyond directly compared document pairs, we propose TrustPropRAG, which structures document relations as a graph and estimates document reliability through multi-hop propagation across the graph. TrustPropRAG anchors this propagation with a limited set of human feedback on document reliability, extending these costly-to-collect feedback-based reliability signals across the whole corpus. Specifically, based on the constructed document relation graph, TrustPropRAG estimates a trust score for each document by formulating and solving an optimization problem that jointly captures pairwise document relations and user feedback. These scores are then used to improve the selection of reliable documents and support trust-aware answer generation. Evaluation results show that TrustPropRAG improves both retrieval quality and exact match over baselines, and remains robust under sparse and noisy feedback.
1 Introduction
Retrieval-augmented generation (RAG) has emerged as an effective approach for mitigating the limitations of large language models (LLMs), whose parametric knowledge may be incomplete or outdated Lewis et al. (2020); Borgeaud et al. (2022); Izacard et al. (2023). By retrieving knowledge from external corpora, LLMs can incorporate relevant and up-to-date information during answer generation. However, RAG’s reliance on external corpora can introduce reliability risks, as the corpora may contain noisy, outdated, or contaminated content (Zhong et al., 2023; Zou et al., 2025), as well as knowledge conflicts (Xie et al., 2024; Chen et al., 2022).
Prior work aims to improve RAG reliability in both the retrieval Ma et al. (2023) and answer generation Wang et al. (2025a) phases. Notably, a line of work leverages relationships among documents to improve the answer reliability of RAG, such as aggregating answers from different documents through majority agreement Xiang et al. (2024), or selecting a subset of documents that are mutually consistent before answer generation Shen et al. (2025). Document relations can be inferred by using natural language inference (NLI) models or leveraging source metadata.
We propose TrustPropRAG, a framework that structures document relations as a graph, as graph representations enable information to propagate through multi-hop relations. Recent studies have explored graph-enhanced RAG, but they mainly focus on organizing knowledge graphs Edge et al. (2024), or constructing index graphs for retrieval Guo et al. (2025). In contrast, we use the graph structure to estimate document reliability and propagate these estimates across the corpus through multi-hop relations. Graph propagation, however, requires anchoring, as document relations reveal consistency or conflict, but not document reliability itself. We anchor this propagation with a set of human feedback on document reliability. In practice, explicit signals (e.g., thumbs-up or thumbs-down ratings) and implicit signals (e.g., clicks or dwell time) can be attributed to the documents that contribute to the generated answer using traceback methods Wang et al. (2025b); Cohen-Wang et al. (2024), thereby providing positive or negative signals for those documents. Since user feedback is costly to collect, TrustPropRAG leverages document relation graphs to propagate sparse feedback across the corpus, similar to how label propagation methods in semi-supervised learning (Zhu et al., 2003; Zhou et al., 2003) spread a few labeled examples across many unlabeled ones. TrustPropRAG then turns the propagated document reliability scores into performance improvements in retrieval and answer generation.
To estimate document reliability and propagate estimates across the corpus, TrustPropRAG first constructs a document relation graph, where nodes represent documents and edges capture document relations (e.g., shared source, mutual support, or contradiction). Based on this graph, TrustPropRAG estimates a trust score for each document by formulating and solving an optimization problem. The problem includes two objective function terms to jointly capture pairwise document relations and incorporate user feedback. At query time, these trust scores are used for trust-aware rescoring, which helps select more reliable documents, and for trust-aware prompting, where the LLM is informed of document reliability during answer generation. Trust score optimization is performed offline over the corpus, amortizing its computational cost across future queries.
Our contributions are threefold. First, we propose a framework that estimates document trust scores by propagating human feedback over a document relation graph, and uses these scores to improve both retrieval and answer generation. Second, we formulate and solve an optimization problem for trust score estimation. We also provide a theoretical analysis of how feedback influences trust score propagation, as well as the convergence rate of the propagation. Third, we evaluate TrustPropRAG with three open-domain question-answering (QA) benchmarks, three retrievers, and six LLMs, and further test its robustness under sparse and noisy feedback. Our code is publicly available at https://github.com/zhliOvO/TrustPropRAG.
2 Related Work
RAG reliability. A growing body of work has studied how to improve RAG reliability. For example, InstructRAG guides LLMs to denoise retrieved contexts using rationales Wei et al. (2025). AstuteRAG compares retrieved evidence with the LLM’s parametric knowledge to handle knowledge conflicts (Wang et al., 2025a). A notable line of work uses relationships (agreements or contradictions) among documents to improve reliability. For instance, RobustRAG isolates retrieved documents, generates per-document answers independently, and aggregates them through majority agreement Xiang et al. (2024). TrustRAG leverages similarity relations among retrieved documents by clustering them and then filtering out contaminated documents before generation Zhou et al. (2025). ReliabilityRAG detects contradictions between documents, selects a subset of documents that are mutually consistent, and produces the final answer via keyword aggregation (Shen et al., 2025). However, these relationships are considered within a subset of documents, and reliability is not propagated beyond the directly compared documents. To move beyond this one-hop relation, TrustPropRAG structures document relations as a graph and captures multi-hop relations among documents through the graph representation.
Graph-empowered RAG. Graphs have long been used to model relationships among textual evidence in tasks such as multi-document summarization (Christensen et al., 2013) and fact verification (Zhong et al., 2020). Recent work has also incorporated graph structures into RAG. For example, GraphRAG organizes corpora with entity knowledge graphs, and leverages community summaries for all groups of closely related entities Edge et al. (2024). LightRAG incorporates graph structures into text indexing and relevant information retrieval for efficient RAG (Guo et al., 2025). NodeRAG introduces heterogeneous graph structures to better leverage the structural nature of graphs Xu et al. (2025). However, these methods primarily use graphs to organize knowledge and guide retrieval. In contrast, we use graph structure to propagate document reliability estimates across the corpus that may contain noisy, outdated, or contradictory information.
Human feedback for enhancing LLMs. Human feedback has been used to align LLM systems with user preferences or intents. For example, InstructGPT trains language models with human feedback using reinforcement learning (Ouyang et al., 2022). RAG-Reward uses preference-based reward models to improve the quality of generated answers (Zhang et al., 2025a). Pistis-RAG aligns human feedback by training a ranking model to improve content ranking and retrieval mechanisms (Bai et al., 2024). Some methods further study feedback in interactive settings. For example, Liu et al. (2025b) leverages implicit feedback in human–LLM dialogues. Zhang et al. (2025b) leverages engagement and disengagement signals for improving LLM-based recommenders. These works mainly use feedback to fine-tune model behavior or improve outputs in human-LLM interactions. TrustPropRAG instead uses limited feedback to estimate document reliability across the large-scale corpus without model fine-tuning.
3 Methodology
Our objective is to enhance RAG by selecting and using more reliable documents from an external corpus for answer generation. The external document corpus may contain unreliable documents, such as those with misinformation, outdated information, or conflicting claims (Xu et al., 2024; Zou et al., 2025). User feedback can indicate whether certain documents are reliable or unreliable. For example, correct and incorrect RAG-generated answers, reported by users or flagged by fact-verification systems, can be traced back to their contributing documents (e.g., using traceback methods Wang et al. (2025b); Cohen-Wang et al. (2024)), providing positive or negative feedback for the corresponding documents. However, user feedback can cover only a small fraction of a large external corpus. To address this, TrustPropRAG takes advantage of document relations to propagate sparse feedback across the corpus.
As shown in Figure 1, TrustPropRAG consists of four main stages. First, we construct a graph where each node represents a document, and each edge encodes a relation between two documents, such as a same-source relation, mutual support, or contradiction (Section 3.1). Second, based on the document relation graph, we formulate and solve optimization problems for estimating trust scores (Section 3.2). We estimate a trust score for each document by jointly optimizing these scores over all documents in the corpus with documents. Third, we perform trust-aware rescoring (Section 3.3). Given a user query , the objective is to produce a ranked list of documents that are provided to the LLM for answer generation, favoring query-relevant documents with higher trust scores. Finally, we perform trust-aware prompting (Section 3.4) by leveraging document trust scores during answer generation.
TrustPropRAG separates offline trust estimation from query-time answer generation. Graph construction and trust score optimization are performed offline over the corpus, amortizing their computational cost across future queries, while trust-aware rescoring and prompting are performed for each query with limited additional overhead.
3.1 Document Relation Graph Construction
Given the corpus , we construct a document relation graph , where each node represents a document, and each edge encodes a relation between documents and . Each edge is assigned a relation label . Specifically, means that the two documents are expected to receive similar trust scores, means that they contain conflicting claims and should not both be highly trusted, and means that no reliable relation can be identified. Each edge also has a confidence weight , which reflects the confidence associated with the relation label .
The relation label and confidence weight can be obtained in different ways. For example, we can use source metadata to connect two documents and from the same source group with supportive edges (), because they are expected to have correlated trustworthiness and thus similar trust scores. Document relations can also be inferred using an NLI model, which derives whether documents and are supportive (), contradictory (), or neutral (). For edges derived by NLI, the confidence weight is set to the softmax probability associated with the relation label.
3.2 Trust Score Optimization
Pairwise consistency loss. We encourage two documents with positive relation labels to have similar trust scores, while enforcing that contradictory documents have diverging trust scores. Specifically, if and are mutually supportive (), their trust scores and should be close to each other, i.e., we aim to minimize ; whereas if , meaning that and are contradictory, their trust scores should be far apart so that at most one document can be trusted, i.e., we aim to minimize . Hence, we define the pairwise consistency loss :
Feedback loss. Over time, we maintain a feedback set : for documents marked as reliable through user feedback, and for documents marked as unreliable. We incorporate this feedback into the optimization problem of trust score estimation. Nodes in are encouraged to have high trust scores, e.g., by minimizing , where is close to . In contrast, nodes in are pushed toward lower trust scores, e.g., by minimizing , where is close to , to limit their influence on answer generation. We implement this by adding a feedback term to the optimization problem of trust score estimation. is formulated as
where is the feedback value.
Formulating trust score optimization problem. To capture pairwise consistency of documents and incorporate user feedback, we jointly optimize objective function terms and . We formulate the optimization problem to assign trust scores to all documents in the corpus:
| (1) |
where is the set of trust scores for all documents, and controls the trade-off between the two terms.
Solving trust score optimization problem. Eq. (1) is a quadratic programming (QP) problem. We solve the problem with projected gradient descent (PGD). The document relation graph can contain many nodes but remains sparse (i.e., each node has supportive or contradictory relation labels with only a small number of other nodes). As a result, each PGD step only needs to compute the loss and gradients for document pairs connected by these relations, rather than all possible document pairs. Section 4.2 reports the measured cost of each stage in TrustPropRAG.
Before running PGD, we initialize each document’s trust score using its relations with other documents in the document relation graph. Specifically, for each document , we compute an initial score by aggregating the signed weights of its connected edges: . Values are min-max normalized to , with . Intuitively, documents connected to others through stronger supportive relations receive higher initial trust scores, while documents involved in stronger contradictory relations receive lower initial trust scores. Graph construction and trust score optimization are performed offline over the corpus, amortizing their computational cost across future queries.
Theoretical analysis. Trust scores exhibit multi-hop propagation. Feedback on a document shifts its trust score; through pairwise document relations, this effect then spreads to its neighbors, which further propagate the influence to their neighbors. We formalize this behavior in two lemmas below. Since is quadratic in , its Hessian is symmetric positive semidefinite. We assume that every connected component of contains at least one document with feedback, which is sufficient to make strictly positive definite (see proof in Appendix C). Let denote its eigenvalues and its condition number. The unconstrained minimizer satisfies the linear system . We use this unconstrained solution to analyze trust propagation and convergence rate.
Lemma 1 (Trust propagation)
Let be the feedback documents in . Suppose the feedback value at each is perturbed with a change of . Then the resulting change in the trust score of document satisfies
where , , and is the shortest-path distance between nodes and in .
Proof. See Appendix C.1.
This lemma shows that trust propagation is distance-weighted and cumulative. Each document receives influence from all documents with feedback. Each influence term is weighted by , which depends on its graph distance to the corresponding document with feedback. The result suggests that sparse feedback can still be effective when the documents with feedback are well positioned in the relation graph, allowing their influence to reach other documents through short relation paths.
Lemma 2 (Convergence)
With the fixed step size , the projected gradient descent iterates satisfy
Proof. See Appendix C.2.
At each iteration, the error contracts by a factor of at most , so the number of iterations needed to reach an -accurate solution (i.e. ) scales as . In our experiments, PGD converges within iterations on tested datasets.
3.3 Trust-aware Rescoring
We select documents with both high trust scores and high query relevance, and provide them to the LLM as retrieved context for answer generation. To this end, we perform trust-aware rescoring for every document by using its estimated trust score and its similarity score. The similarity score is computed by the retriever, e.g., as a dot product score, to measure how relevant the document is to the query. Specifically, for document , we combine its trust score with its normalized similarity score and define the trust-aware ranking score as , where controls the trade-off between document trustworthiness and its relevance to the query. We select the top- documents with the largest and provide them to the LLM for answer generation.
3.4 Trust-aware Answer Generation
Apart from using trust-aware rescoring for document retrieval, we further leverage the estimated trust scores during answer generation. Specifically, for the top- documents with the largest trust-aware ranking scores, we format the prompt so that each document is accompanied by its trust score. We also include an instruction in the prompt to guide the LLM to place greater reliance on documents with higher trust scores during answer generation. This trust-aware prompting complements trust-aware rescoring: while rescoring filters out less reliable documents before answer generation, this trust-aware prompting helps the LLM account for document reliability during answer generation.
4 Evaluation
4.1 Evaluation Setup
Datasets. We evaluate TrustPropRAG with three open-domain QA datasets: MS MARCO (Bajaj et al., 2016), Natural Questions (NQ) (Kwiatkowski et al., 2019), and TriviaQA (Joshi et al., 2017). These datasets are widely adopted in retrieval-augmented and open-domain QA evaluation, and together they cover diverse question types. Following the evaluation settings commonly used in prior RAG reliability and robustness studies (Shen et al., 2025; Zhong et al., 2023; Wei et al., 2025; Wang et al., 2025a), we randomly sample 300 query–answer instances for each dataset. Half of the sampled queries are paired with synthetic contradictory documents generated by following the method in PoisonedRAG (Zou et al., 2025). These documents are written to remain fluent and relevant to the query while supporting answers that conflict with the ground-truth answers. The remaining queries are paired only with factual documents from the original datasets. In this way, we construct the evaluation corpus from the relevant documents associated with the sampled queries, resulting in 805, 893, and 1,037 documents for MS MARCO, NQ, and TriviaQA, respectively. This setup evaluates whether the end-to-end system can produce correct answers when the corpus contains both reliable evidence and fluent but misleading contradictory content.
Retrievers. We evaluate with both sparse and dense retrievers. We use BM25 (Robertson and Zaragoza, 2009), a sparse lexical retriever, and two dense bi-encoder retrievers: Contriever (Izacard et al., 2022) and MiniLM based on all-MiniLM-L6-v2 (Reimers and Gurevych, 2019). For each query, we retrieve documents.
Graph construction. We construct a document relation graph using two types of information: document sources and semantic relations between documents. Documents and from the same source are connected with supportive edges (). In TriviaQA, we determine whether two documents share the same source based on their source Uniform Resource Locators (URLs). For MS MARCO and NQ, which do not provide source metadata, we use simulated source grouping: factual documents are partitioned into multiple source groups, and contradictory documents are partitioned into another set of source groups. Documents in the same group are treated as sharing the same source. We adopt this simulation as a proxy for real source metadata, following recent source-reliability-aware RAG work that simulates sources with varying reliability (Hwang et al., 2025). We also apply an NLI model (nli-deberta-v3-small) to each document pair. If the model predicts that two documents and support each other, we add a supportive edge (). If it predicts that they contradict each other, we add a contradictory edge (). Appendix F reports the precision and recall of the NLI-derived edges on all three datasets. Table 4 evaluates the impact of removing the source-group edges to isolate the contribution of the NLI-derived edges.
Feedback construction. We construct the feedback set by randomly sampling a fraction of factual and contradictory documents and simulating feedback labels for them. Sampled factual documents are assigned positive feedback labels, while sampled contradictory documents are assigned negative feedback labels. We refer to this sampling fraction as the feedback ratio (FB); for instance, FB = 30% means that feedback labels are provided for 30% of the documents. As real-world feedback may be imperfect, we evaluate TrustPropRAG under noisy feedback in Figure 2, where a portion of the feedback labels are incorrect.
Trust optimization. For PGD, we use a step size of , run for at most iterations, and stop when the convergence tolerance reaches . We set feedback loss weight . Trust scores are optimized over all corpus documents, resulting in large-scale optimization problems over 805, 893, and 1,037 documents for MS MARCO, NQ, and TriviaQA, respectively.
Rescoring. We set the trust coefficient to . We also test the impact of in Section 4.2. The top- documents, where in our experiments, are provided as retrieved context to the LLM. We also test the impact of in Appendix B.
Answer generation. We use proprietary and open-source LLMs, including GPT-4o-mini, GPT-5, Gemini 2.5 Flash, Gemini 2.5 Pro, Llama-3.1-8B-Instruct, and Mistral-7B-Instruct. TrustPropRAG leverages the trust-aware prompt that annotates each document with its trust score (as detailed in Section 3.4).
Evaluation metrics. We adopt exact match (EM) and fact precision@ (FP@) as evaluation metrics. EM is a widely used metric in open-domain QA. EM checks whether the normalized generated answer contains the ground-truth answer, with normalization including lowercasing and the removal of articles and punctuation. FP@ evaluates retrieval quality, which measures the reliability of the top-ranked documents used for answer generation. Specifically, let and denote the sets of factual and contradictory documents, respectively, and let denote the top- documents for query . FP@ is defined as the fraction of factual documents among these top- documents:
Baselines. We compare TrustPropRAG against the following methods: (1) VanillaRAG ranks documents using the relevance scores produced by the retriever and feeds the top-ranked documents to the LLM. (2) InstructRAG (Wei et al., 2025) guides the LLM to better use retrieved evidence during generation. It decomposes retrieved documents into atomic claims, and prompts the LLM to generate a rationale connecting relevant claims to the final answer. (3) AstuteRAG (Wang et al., 2025a) combines the LLM’s internal knowledge with retrieved evidence. It first prompts the LLM to produce an answer-relevant document from its internal knowledge, and then compares this document with the retrieved documents and resolves conflicts before generating the answer. (4) ReliabilityRAG (Shen et al., 2025) uses NLI to identify contradictions among retrieved documents. It then solves a maximum independent set problem to select a subset of documents with no detected pairwise contradictions and uses keyword aggregation to generate the final answer. (5) TrustRAG (Zhou et al., 2025) includes a first-stage filtering procedure that clusters the retrieved documents by embedding similarity and filters out clusters identified as suspicious. We use this first-stage procedure as the baseline.
4.2 Evaluation Results
| Dataset | Method | Contriever | BM25 | MiniLM |
| MS MARCO | VanillaRAG | 0.27 | 0.24 | 0.26 |
| InstructRAG | 0.35 | 0.40 | 0.37 | |
| AstuteRAG | 0.38 | 0.39 | 0.37 | |
| ReliabilityRAG | 0.42 | 0.40 | 0.39 | |
| TrustRAGstage 1 | 0.36 | 0.35 | 0.33 | |
| TrustPropRAG | 0.45 | 0.47 | 0.44 | |
| NQ | VanillaRAG | 0.37 | 0.37 | 0.39 |
| InstructRAG | 0.46 | 0.48 | 0.43 | |
| AstuteRAG | 0.44 | 0.50 | 0.42 | |
| ReliabilityRAG | 0.45 | 0.50 | 0.48 | |
| TrustRAGstage 1 | 0.44 | 0.46 | 0.41 | |
| TrustPropRAG | 0.55 | 0.59 | 0.59 | |
| TriviaQA | VanillaRAG | 0.57 | 0.62 | 0.49 |
| InstructRAG | 0.68 | 0.67 | 0.69 | |
| AstuteRAG | 0.70 | 0.71 | 0.72 | |
| ReliabilityRAG | 0.75 | 0.72 | 0.67 | |
| TrustRAGstage 1 | 0.66 | 0.65 | 0.63 | |
| TrustPropRAG | 0.78 | 0.75 | 0.77 |
| Dataset | Retriever | Baselines | TrustPropRAG |
| MS MARCO | Contriever | 44% | 77% |
| BM25 | 57% | 92% | |
| MiniLM | 54% | 88% | |
| NQ | Contriever | 53% | 79% |
| BM25 | 46% | 88% | |
| MiniLM | 48% | 96% | |
| TriviaQA | Contriever | 38% | 77% |
| BM25 | 57% | 90% | |
| MiniLM | 49% | 92% |
TrustPropRAG outperforms baselines on EM and FP@5. Tables 1 and 2 report the EM and FP@5 averaged over GPT-4o-mini, Gemini 2.5 Flash, and GPT-5. Table 1 shows that TrustPropRAG achieves the best EM across all datasets and retrievers, with absolute EM gains of 0.03–0.11 compared with the strongest baseline. Table 2 shows FP@5 (the fraction of factual documents among top- documents) of TrustPropRAG and baselines. All baselines have the same FP@5 as VanillaRAG under this retrieval-level metric: InstructRAG, AstuteRAG, and TrustRAG operate after the retriever selects the top- documents based on query relevance and thus do not change the ranked top- documents; although ReliabilityRAG filters documents before final answer generation, its filtering is driven by intermediate LLM-generated answers and their consistency, rather than by producing a new top- list. The baseline methods have an FP@5 of 38–57%, while TrustPropRAG raises it to 77–96%. The improvement is observed with all three retrievers, indicating that the method is not tied to a specific retriever. These improvements in FP@5 help explain the EM gains of TrustPropRAG. Specifically, a higher FP@5 indicates that the top- list contains a larger fraction of factual documents, which suggests that trust score optimization assigns higher trust to more reliable documents, and trust-aware rescoring promotes them into the top- list. As a result, the LLM inputs contain more documents that support the correct answer.
Ablation study. We conduct an ablation study to examine the contribution of each component in TrustPropRAG. Table 3 evaluates five variants on MS MARCO with FB=30% using Gemini 2.5 Flash. We observe that the combination of trust score optimization and trust-aware rescoring leads to a substantial improvement. Compared with VanillaRAG, adding these two components increases FP@5 from 54% to 87% and improves EM from 0.28 to 0.41. This suggests that the optimized trust scores are effective for identifying more trustworthy documents and that rescoring can promote them into the top- context used for answer generation. We also observe that trust-aware prompting provides an additional benefit when combined with trust score optimization and trust-aware rescoring, increasing EM from 0.41 to 0.46. This suggests that some low-trust documents may still remain in the top- list after rescoring, and trust-aware prompting helps the LLM reduce their influence during answer generation. Finally, we observe that using trust-aware prompting alone does not improve performance. Without trust score optimization, FP@5 remains at 54%, and EM slightly decreases from 0.28 to 0.27 compared with VanillaRAG. This suggests the trust-aware prompt becomes useful only when the trust scores have been optimized and can provide a meaningful signal for answer generation.
| Variant | Opt. | Rescore | Prompt | EM | FP@5 |
| VanillaRAG | – | – | – | 0.28 | 54% |
| Opt.+Prompt | ✓ | – | ✓ | 0.35 | 54% |
| Opt.+Rescore | ✓ | ✓ | – | 0.41 | 87% |
| Prompt only | – | – | ✓ | 0.27 | 54% |
| Full TrustPropRAG | ✓ | ✓ | ✓ | 0.46 | 88% |
| Variant | EM | FP@5 |
| VanillaRAG | 0.31 | 51% |
| TrustPropRAG (trivial feedback) | 0.36 | 62% |
| TrustPropRAG (NLI-only) | 0.44 | 79% |
| TrustPropRAG | 0.49 | 92% |
| Model | MS MARCO | NQ | TriviaQA | ||||
| VanillaRAG | TrustPropRAG | VanillaRAG | TrustPropRAG | VanillaRAG | TrustPropRAG | ||
| Proprietary | |||||||
| GPT-4o-mini | 0.21 | 0.40 | 0.41 | 0.58 | 0.50 | 0.76 | +0.21 |
| GPT-5 | 0.30 | 0.49 | 0.43 | 0.62 | 0.53 | 0.85 | +0.23 |
| Gemini 2.5 Flash | 0.28 | 0.46 | 0.33 | 0.56 | 0.44 | 0.69 | +0.22 |
| Gemini 2.5 Pro | 0.32 | 0.48 | 0.37 | 0.60 | 0.46 | 0.77 | +0.23 |
| Open-source | |||||||
| Llama-3.1-8B-Instruct | 0.19 | 0.32 | 0.38 | 0.55 | 0.48 | 0.77 | +0.20 |
| Mistral-7B-Instruct | 0.22 | 0.35 | 0.33 | 0.47 | 0.51 | 0.74 | +0.17 |
The effects of trust propagation and source-group edges. To separate the benefit of trust propagation from simply having access to feedback, the trivial feedback variant uses the same feedback set as TrustPropRAG but does not propagate the feedback through the graph. Documents with positive feedback receive a boost to their retrieval scores, while documents with negative feedback are not included in the top- list. To isolate the contribution of the NLI-derived edges from source-group edges, the NLI-only variant removes all simulated source-group edges and retains only the NLI-derived edges.
From Table 4, the trivial feedback variant achieves an EM of 0.36 and an FP@5 of 62%. This confirms that directly using feedback is beneficial. However, its performance remains substantially below TrustPropRAG, which achieves an EM of 0.49 and an FP@5 of 92%, respectively. This gap shows that the gain of TrustPropRAG cannot be attributed only to access to the feedback set. Propagating feedback through document relation graphs provides trust estimates for documents without direct feedback. Table 4 also shows that the NLI-only variant achieves an EM of 0.44 and an FP@5 of 79%. Although removing the source-group edges leads to a performance drop, the NLI-only variant still outperforms the trivial feedback variant, showing that propagation remains beneficial even when the graph contains only NLI-derived relations.
Performance under noisy feedback. Figure 2 evaluates the robustness of TrustPropRAG by varying the accuracy of feedback labels from 60% to 100%. Feedback accuracy denotes the fraction of feedback labels that are correct. For example, 80% feedback accuracy means that 20% of the feedback labels are flipped: a factual document may be labeled as low-trust, or a contradictory document may be labeled as high-trust. Figure 2 reports EM on MS MARCO with Gemini 2.5 Flash. We observe that TrustPropRAG is robust to noisy feedback. Even when feedback accuracy is only 60% (i.e., 40% of feedback labels are incorrect), TrustPropRAG still improves EM over the setting without feedback across all feedback ratios. This is because individual feedback errors can be mitigated through trust propagation over the document relation graph, where relations among documents help mitigate the impact of noisy feedback signals.
The impact of feedback ratio. Figure 2 also shows EM as a function of the feedback ratio on MS MARCO with Gemini 2.5 Flash. Compared with the setting without feedback, TrustPropRAG improves sharply when only 10% of documents receive feedback, after which the improvement becomes more gradual. This indicates that a small portion of feedback is sufficient for trust propagation to produce reliable trust estimates across the document relation graph.
The impact of trust coefficient . Figure 3 shows EM and FP@5 as the trust coefficient varies from 0 to 1. When , rescoring uses only the retrieval relevance score. As increases from , both EM and FP@5 improve, showing that incorporating trust scores helps promote documents that support the correct answer into the top- context. EM reaches its best value at , and remains stable between and . However, when is too large, the top- list tends to include documents with high trust scores but lower query relevance, leading to a slight decrease in EM. As TrustPropRAG is not sensitive to the exact choice of within a range (e.g., between 0.5 and 0.8), we set for the rest of the experiments.
Performance with different LLMs. We evaluate the performance of TrustPropRAG with a diverse set of proprietary and open-source LLMs of different model scales in Table 5. Since FP@5 depends only on retrieval and rescoring, it is unchanged across LLMs; therefore, we report EM as the main metric. We use MiniLM as the retriever and set and FB=30%. Table 5 shows that TrustPropRAG consistently outperforms VanillaRAG across all evaluated LLMs. The average EM improvement ranges from 0.17 to 0.23 across both proprietary models (GPT-4o-mini, GPT-5, Gemini 2.5 Flash, Gemini 2.5 Pro) and smaller open-source models (Llama-3.1-8B-Instruct, Mistral-7B-Instruct). Notably, TrustPropRAG still achieves an average EM gain of 0.23 with strong LLMs such as GPT-5 and Gemini 2.5 Pro. This suggests that improving the reliability of the retrieved context with trust propagation remains useful even for strong LLMs.
Computational cost. Table 6 reports the computational cost of TrustPropRAG, including offline graph construction and trust optimization, as well as query-time retrieval. The computational resources are detailed in Appendix A. Constructing the graph accounts for most of the offline cost, taking 578.3–1077.1 seconds. Given the constructed graph, trust score optimization is efficient, requiring only 0.3–0.4 seconds, indicating that trust propagation itself introduces negligible computational overhead. Both graph construction and trust optimization are performed offline and their costs are therefore amortized across future queries. At query time, retrieval takes 4.4–4.8 seconds and is shared by TrustPropRAG and the baselines. Beyond retrieval, TrustPropRAG only performs lightweight trust-aware rescoring using the precomputed trust scores, introducing limited additional query-time overhead.
| Stage | MS MARCO | NQ | TriviaQA |
| Offline | |||
| Graph construction | 578.3 s | 852.4 s | 1077.1 s |
| Trust optimization (PGD) | 0.3 s | 0.3 s | 0.4 s |
| Query time | |||
| Retrieval | 4.4 s | 4.8 s | 4.8 s |
5 Conclusions
We present TrustPropRAG, a framework that improves RAG reliability by propagating sparse human feedback over a document relation graph. TrustPropRAG estimates a trust score for each document by optimizing an objective that combines pairwise consistency loss and feedback loss. The resulting trust scores are used to select more reliable documents and guide trust-aware answer generation. TrustPropRAG achieves absolute EM gains of 0.03–0.11 over the strongest baseline, and demonstrates robustness to noisy feedback.
Acknowledgments
We thank the anonymous reviewers for insightful feedback. This work was supported by Seed Grant of IST, and the National Science Foundation under Grants 2550742, 2623125, 2555329, and 2549266.
Limitations
TrustPropRAG incorporates human feedback as binary labels, marking documents as either reliable or unreliable. In practice, feedback can be graded rather than binary. Studying TrustPropRAG under graded feedback is left to future work. Following common practice in RAG evaluation, we focus on unreliable content that is injected as synthetic contradictions rather than drawn from naturally occurring conflicts. Finally, we focus on open-domain QA with short factual answers. Applying trust propagation to long-form or multi-hop generation, where reliability interacts with compositional reasoning, is left to future work.
Ethical Considerations
This work addresses the problem of unreliable content in retrieval-augmented generation. Our experiments use synthetically generated contradictions. All datasets and models used in this work are publicly released and licensed for research use, and our usage is consistent with their intended research purposes. MS MARCO, NQ, and TriviaQA are publicly available open-domain QA benchmarks released for research. The retrievers (BM25, Contriever, and the all-MiniLM-L6-v2 encoder) and the NLI model (nli-deberta-v3-small) are also publicly released.
References
- Pistis-RAG: enhancing retrieval-augmented generation with human feedback. arXiv preprint arXiv:2407.00072. Cited by: §2.
- MS MARCO: a human generated machine reading comprehension dataset. In Proceedings of the Workshop on Cognitive Computation: Integrating Neural and Symbolic Approaches (CoCo@NIPS), Cited by: §4.1.
- Decay bounds and O() algorithms for approximating functions of sparse matrices. Electronic Transactions on Numerical Analysis 28, pp. 16–39. Cited by: §C.1.
- Improving language models by retrieving from trillions of tokens. In Proceedings of the 39th International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, Vol. 162, pp. 2206–2240. Cited by: §1.
- Rich knowledge sources bring complex knowledge conflicts: recalibrating models to reflect conflicting evidence. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Abu Dhabi, United Arab Emirates, pp. 2292–2307. Cited by: §1.
- Towards coherent multi-document summarization. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Atlanta, Georgia, pp. 1163–1173. Cited by: §2.
- ContextCite: attributing model generation to context. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 37. Cited by: §1, §3.
- Decay rates for inverses of band matrices. Mathematics of Computation 43 (168), pp. 491–499. Cited by: §C.1, Appendix D.
- From local to global: a graph RAG approach to query-focused summarization. arXiv preprint arXiv:2404.16130. Cited by: §1, §2.
- LightRAG: simple and fast retrieval-augmented generation. In Findings of the Association for Computational Linguistics: EMNLP 2025, pp. 10746–10761. Cited by: §1, §2.
- Retrieval-augmented generation with estimation of source reliability. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 34279–34303. Cited by: §4.1.
- Unsupervised dense information retrieval with contrastive learning. Transactions on Machine Learning Research. Cited by: §4.1.
- Few-shot learning with retrieval augmented language models. Journal of Machine Learning Research 24 (251), pp. 1–43. Cited by: §1.
- TriviaQA: a large scale distantly supervised challenge dataset for reading comprehension. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Cited by: §4.1.
- Natural questions: a benchmark for question answering research. Transactions of the Association for Computational Linguistics 7, pp. 452–466. Cited by: §4.1.
- Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems, Vol. 33. Cited by: §1.
- Open domain question answering with conflicting contexts. In Findings of the Association for Computational Linguistics: NAACL 2025, pp. 1838–1854. Cited by: Appendix G.
- User feedback in human-LLM dialogues: a lens to understand users but noisy as a learning signal. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 2666–2681. Cited by: §2.
- Query rewriting in retrieval-augmented large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pp. 5303–5315. Cited by: §1.
- Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, Vol. 35. Cited by: §2.
- Sentence-BERT: sentence embeddings using siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, pp. 3982–3992. Cited by: §4.1.
- The probabilistic relevance framework: BM25 and beyond. Foundations and Trends in Information Retrieval 3 (4), pp. 333–389. Cited by: §4.1.
- ReliabilityRAG: effective and provably robust defense for RAG-based web-search. In Advances in Neural Information Processing Systems, Cited by: Appendix B, §1, §2, §4.1, §4.1.
- Astute RAG: overcoming imperfect retrieval augmentation and knowledge conflicts for large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria, pp. 30553–30571. Cited by: §1, §2, §4.1, §4.1.
- TracLLM: a generic framework for attributing long context llms. In USENIX Security Symposium, Cited by: §1, §3.
- InstructRAG: instructing retrieval-augmented generation via self-synthesized rationales. In Proceedings of the Thirteenth International Conference on Learning Representations (ICLR), Cited by: §2, §4.1, §4.1.
- Certifiably robust RAG against retrieval corruption. arXiv preprint arXiv:2405.15556. Cited by: §1, §2.
- Adaptive chameleon or stubborn sloth: revealing the behavior of large language models in knowledge conflicts. In Proceedings of the Twelfth International Conference on Learning Representations (ICLR), Cited by: §1.
- Knowledge conflicts for LLMs: a survey. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, Florida, USA, pp. 8541–8565. Cited by: §3.
- NodeRAG: structuring graph-based rag with heterogeneous nodes. arXiv preprint arXiv:2504.11544. Cited by: §2.
- RAG-Reward: optimizing RAG with reward modeling and RLHF. arXiv preprint arXiv:2501.13264. Cited by: §2.
- Leveraging unpaired feedback for long-term LLM-based recommendation tuning. In Findings of the Association for Computational Linguistics: EMNLP 2025, pp. 24507–24521. Cited by: §2.
- Reasoning over semantic-level graph for fact checking. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pp. 6170–6180. Cited by: §2.
- Poisoning retrieval corpora by injecting adversarial passages. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Cited by: §1, §4.1.
- Learning with local and global consistency. In Advances in Neural Information Processing Systems, Vol. 16, pp. 321–328. Cited by: §1.
- TrustRAG: enhancing robustness and trustworthiness in retrieval-augmented generation. arXiv preprint arXiv:2501.00879. Cited by: §2, §4.1.
- Semi-supervised learning using Gaussian fields and harmonic functions. In Proceedings of the 20th International Conference on Machine Learning, pp. 912–919. Cited by: §1.
- PoisonedRAG: knowledge corruption attacks to retrieval-augmented generation of large language models. In 34th USENIX Security Symposium (USENIX Security 25), pp. 3827–3844. Cited by: Appendix B, §1, §3, §4.1.
Appendix A Implementation Details
Models and computational resources. Among the LLMs we evaluate, Llama-3.1-8B-Instruct and Mistral-7B-Instruct are open-source models with 8B and 7B parameters, respectively; GPT-4o-mini, GPT-5, Gemini 2.5 Flash, and Gemini 2.5 Pro are proprietary models accessed through their respective APIs. The NLI model (nli-deberta-v3-small) has approximately 44M parameters, and the dense retrievers (Contriever and all-MiniLM-L6-v2) are sentence-encoder models with approximately 110M and 22M parameters respectively. All local computation, including NLI-based edge construction, dense retrieval, and trust score optimization, was run on an NVIDIA RTX 5070 GPU. For our largest corpus (1,037 documents, TriviaQA), this took approximately 0.3 GPU-hours. Trust score optimization via projected gradient descent is lightweight, converging within 1,000 iterations and taking under 1 minute per corpus. We implement retrieval, NLI inference, and trust score optimization using standard Python libraries, and provide exact package versions and scripts in the released code.
Hyperparameters. We use the same hyperparameters across all datasets, retrievers, and LLMs. For trust score initialization, we set and . We set feedback loss weight . We set the feedback values and .
Appendix B Number of Retrieved Documents
The number of retrieved documents controls the size of the context provided to the LLM. A smaller yields a more selective context but may omit relevant evidence, while a larger provides broader coverage at the cost of allowing more unreliable documents into the prompt. We evaluate TrustPropRAG across on MS MARCO with MiniLM as the retriever, Gemini 2.5 Flash as the backbone LLM, and FB30%. All other hyperparameters follow Section 4.
| =1 | =3 | =5 | =10 | =20 | |
| EM | 0.41 | 0.43 | 0.46 | 0.45 | 0.42 |
| FP@ | 95% | 92% | 88% | 84% | 73% |
FP@ falls steadily from 95% (=1) to 73% (=20) as lower-trust documents enter the top- selected documents. EM is less sensitive, staying within 0.41–0.46 and peaking at =5 and =10. Trust-aware rescoring places reliable evidence at the top, so small already provides sufficient factual support; at larger , the trust-aware prompt helps the LLM put less weight to lower-trust documents that are provided in the context. EM is smaller at =1, where the retrieved context may provide insufficient context, and at =20, where the larger retrieval set exposes the LLM to more contradictory context.
Considering that EM peaks at =5 and =10, we adopt =5, matching prior work (Zou et al., 2025; Shen et al., 2025).
Appendix C Proofs
We provide proofs of the two lemmas stated in Section 3.2. Both lemmas rely on the fact that the trust optimization objective is a quadratic in , so the first-order optimality condition for the unconstrained quadratic problem is given by the linear system where is the Hessian. We first verify that is positive definite under the assumption that every connected component of contains at least one document with feedback, then prove the two lemmas.
Positive definiteness of .
Let denote the -th standard basis vector, whose -th entry is one and all other entries are zero. Expanding the Hessian gives
| (2) |
which is a sum of positive semidefinite matrices, hence . Next, we prove . A vector lies in iff (a) on every supportive edge (), (b) on every contradictory edge (), and (c) on every node with feedback. If a node with value zero is connected to another node by a supportive edge, condition (a) implies that the neighboring node also has value zero. If it is connected by a contradictory edge, condition (b) again implies that the neighboring node has value zero. Under our assumption that every connected component contains at least one feedback node, we have globally, so . We write its eigenvalues as and its condition number as .
C.1 Proof of Lemma 1 (Trust Propagation)
We fix the set of documents with feedback and consider perturbations of their feedback values. Let denote the perturbation applied to the feedback value . These perturbations leave unchanged and only modify the right-hand side at entries corresponding to documents in . Let denote the resulting change in , and let denote its entry corresponding to feedback document . We have
| (3) |
Step 1: Linearity and superposition.
Because the perturbations only change the feedback target values, the matrix remains fixed. Therefore, the unconstrained optimum satisfies After the perturbation, the right-hand side becomes , and the corresponding change in the optimized trust scores is
Since is nonzero only at documents with feedback, the change in the trust score of document can be written as
| (4) |
where denotes the entry in the -th row and -th column of .
Step 2: Entry-wise decay of .
The feedback term contributes only to the diagonal entries of . The off-diagonal entry () is nonzero when documents and are joined by an edge (). Specifically, from Eq. (2), each pairwise relation term adds at positions and of . The sparsity pattern of therefore coincides with the adjacency structure of . For a symmetric positive definite matrix with this property, the Demko–Moss–Smith decay bound (Demko et al., 1984; Benzi and Razouk, 2007) gives an entry-wise decay of that is exponential in graph distance:
| (5) |
where .
Step 3: Superposition bound.
C.2 Proof of Lemma 2 (Convergence)
Let denote the gradient descent step before projection, and let denote the projected iterate. Since is quadratic with Hessian , for any we have . Therefore,
| (6) | ||||
With the fixed step size , each eigenvalue of is mapped to an eigenvalue of . Since , these eigenvalues lie in Thus, the largest absolute eigenvalue of is Because is symmetric, its spectral norm is equal to its largest absolute eigenvalue, so
Taking norms in (6) then gives
The projection operator is non-expansive, and hence
Appendix D Impact of Mislabeled Edges
A mislabeled edge between documents and changes the Hessian in Eq. (2) only at the entries involving and . Under the same assumptions used in Lemma 1, the entry-wise decay bound on from Demko et al. (1984) implies that the resulting perturbation to the optimized trust scores decreases exponentially with the shortest-path distance from the mislabeled edge. Therefore, the influence of an edge error diminishes rapidly as it propagates to more distant documents.
Appendix E Prompt Templates
Trust-aware prompt.
Each retrieved document is annotated with both its trust score and a trust label. We assign the label based on the trust score: high if , low if , and medium otherwise. The system instruction is:
Documents are formatted as:
Baseline prompt. VanillaRAG uses the following prompt. The other baselines use the prompts specified in their original papers.
Appendix F Quality of NLI-derived Edges
We evaluate the NLI-derived edges against the known factual and contradictory structure of the corpus on the three datasets. We treat an edge between a factual document and a contradictory document as a ground-truth contradictory relation, and an edge between two documents of the same category as a supportive relation. Table 8 reports the precision and recall of the contradictory and supportive edges produced by the NLI model.
| Dataset | Contradictory | Supportive | ||
| Precision | Recall | Precision | Recall | |
| MS MARCO | 0.65 | 0.18 | 0.71 | 0.14 |
| NQ | 0.69 | 0.21 | 0.68 | 0.17 |
| TriviaQA | 0.58 | 0.13 | 0.62 | 0.12 |
The NLI-derived edges have moderate precision and low recall, as the NLI model predicts most document pairs as neutral. The low recall indicates that many relations are omitted, resulting in a sparse graph, which reflects practical settings where only a limited subset of inter-document relations can be identified. The imperfect precision indicates that the graph also contains some incorrectly labeled edges. Nevertheless, the NLI-only ablation in Table 4 shows that TrustPropRAG still outperforms the baselines when relying only on these imperfect edges, which suggests that trust propagation is robust to noise and sparsity in graph construction.
Appendix G Performance on QACC
We evaluate TrustPropRAG on naturally occurring conflicts without any synthetic injection. We used QACC Liu et al. (2025a), a human-annotated open-domain QA dataset in which unambiguous questions are paired with real web contexts retrieved via Google Search. The conflicts among these contexts arise naturally on the web rather than from injected documents. We randomly sampled 100 questions from QACC and constructed the evaluation corpus from their retrieved contexts. As shown in Table 9, TrustPropRAG achieves the best EM (0.59) on this naturally conflicting corpus, outperforming the strongest baseline ReliabilityRAG (0.56) and improving over VanillaRAG by 0.11. These results indicate that the effectiveness of TrustPropRAG is not limited to synthetically injected contradictions and extends to naturally occurring conflicts.
| Method | EM |
| VanillaRAG | 0.48 |
| InstructRAG | 0.54 |
| AstuteRAG | 0.50 |
| ReliabilityRAG | 0.56 |
| TrustRAGstage 1 | 0.50 |
| TrustPropRAG | 0.59 |