ISO-RAG: Isoperimetric Noise Control for Retrieval-Augmented Generation
Abstract
Retrieval-Augmented Generation (RAG) mitigates large language models (LLMs) hallucinations, yet conventional dense retrieval struggles with the complex reasoning paths of multi-hop question answering (QA). Graph-based RAG captures multi-step relationships but suffers from severe semantic drift and high online latency due to noisy global graph traversals. Thus, we propose ISO-RAG (ISOperimetric Retrieval-Augmented Generation), a training-free, purely topology-driven RAG framework. Breaking away from computationally expensive continuous geometric embeddings, ISO-RAG leverages discrete graph theory, specifically the local Cheeger ratio (topological expansion rate), to compute node-wise isoperimetric profiles. By identifying and pruning spurious shortcut edges that lead to combinatorial explosion, ISO-RAG restricts the search space to a strictly localized, contextually safe subgraph. This topological purification regulates deterministic Personalized PageRank (PPR) diffusion during retrieval, ensuring exact and low-latency convergence without probability leakage. Experiments on multi-hop QA benchmarks demonstrate that ISO-RAG outperforms state-of-the-art baselines by average absolute gains of 10.0% in retrieval recall and 4.3% in downstream exact match, achieving a superior accuracy-efficiency trade-off by fundamentally eliminating the latency bottleneck of global traversals. Our source code is available at https://github.com/ZaiizaiZHANG/ISO-RAG.git.
1University of Technology Sydney 2University of New South Wales
1 Introduction
Retrieval-Augmented Generation (RAG) (Lewis et al. 2020; Guu et al. 2020) is a standard paradigm for improving the factuality and controllability of large language models (LLMs) (Kaplan et al. 2020; Vaswani et al. 2017). Grounding generation in external evidence substantially reduces hallucinations (Ji et al. 2023; Gao et al. 2023; Jiang et al. 2023) and improves answer reliability. While effective for single-hop factual lookup, RAG remains less reliable for multi-hop question answering, which requires connecting multiple evidence pieces scattered across different documents. In such settings, retrieval quality becomes the dominant bottleneck, as LLMs cannot reliably infer answers (Wei et al. 2022; Yao et al. 2022; Yao et al. 2023) without the complete reasoning chain.
Existing sparse and dense retrievers (e.g., BM25 (Robertson, Zaragoza et al. 2009), MDR (Xiong et al. 2021), (Karpukhin et al. 2020)) rely on flat similarity matching, often failing to reconstruct the full reasoning chains required by compositional benchmarks (Trivedi et al. 2022; Chen et al. 2023). To address this, recent graph-based (Edge et al. 2024; Jimenez Gutierrez et al. 2024; Cao et al. 2025) methods organize corpora into interconnected structures for multi-hop evidence aggregation. Notably, HyperbolicRAG (Cao et al. 2025) embeds document networks into continuous hyperbolic spaces to model their inherent scale-free hierarchies. This non-Euclidean mapping enables capturing complex multi-hop dependencies with low structural distortion.
Despite these advantages, conventional graph paradigms lack explicit topological intervention, manifesting in three limitations (Figure 1): (1) Unconstrained diffusion leads to semantic drift. Graph-based frameworks (Edge et al. 2024; Jimenez Gutierrez et al. 2024; Cao et al. 2025) utilizing Personalized PageRank (PPR) (Page et al. 1999) propagate probabilities over dense graphs. As illustrated by the MuSiQue query, answering requires traversing a specific two-hop path to the target entity (e.g., Poptropica to Pearson Education). However, unconstrained PPR retrieves misleading cues through spurious edges (e.g., release year 2007). Because these distractors exhibit high textual overlap with the query context, they bypass downstream re-rankers, causing the LLM to hallucinate the incorrect year. Crucially, if unconstrained diffusion breaks down on a mere two-hop path, this semantic drift compounds for more complex three- or four-hop queries. (2) Continuous models fail to prune noise. Methods like HyperbolicRAG (Cao et al. 2025) rely on continuous node embeddings without explicitly pruning noisy edges (depicted as the scissors in Figure 1), thereby retaining erroneous pathways to distractors. This fails to isolate the specific reasoning branches necessary for accurate multi-hop deduction. (3) Dense graphs degrade efficiency. Computing random walks over the entire graph explores unrelated entities (e.g., distant brands like Pepsi-Cola), which incurs significant computational overhead and dilutes the probability mass of the actual target entity. Consequently, these paradigms struggle to balance signal fidelity and retrieval latency.
In response to these limitations, we propose ISO-RAG (ISOperimetric Retrieval-Augmented Generation), a geometry-aware local graph retrieval framework for multi-hop QA. The core intuition of ISO-RAG is transitioning from unconstrained probability diffusion to geometrically regulated local diffusion. After routing a query to semantically aligned seed nodes to form a candidate subgraph, the framework maps these nodes into a hyperbolic space using the Poincaré ball model (Nickel and Kiela 2017; Chami et al. 2019; Balazevic, Allen, and Hospedales 2019; Gulcehre et al. 2019; Nickel and Kiela 2018; Peng et al. 2022). Document networks in multi-hop QA naturally form hierarchical structures, where a few generic entities act as dense hubs connecting numerous specific facts. Because hyperbolic space expands exponentially, it embeds these scale-free topologies with low geometric distortion. Crucially, this non-Euclidean embedding structurally highlights intrinsic bottlenecks (Alon and Yahav 2021; Topping et al. 2022), such as the dense generic hubs (Girvan and Newman 2002) that misguide retrieval. In spectral graph theory, the classical Cheeger constant (Chung 1997) identifies global graph bottlenecks. Building on this geometric foundation, ISO-RAG introduces a localized isoperimetric control mechanism (Andersen, Chung, and Lang 2006; Krioukov et al. 2010).Because computing a strict set-level isoperimetric constant is computationally prohibitive for dynamic online retrieval, we design a numerically stable, node-wise proxy. This proxy acts as a structural filter, explicitly identifying and severing incompatible edges prior to propagation. By anchoring the Personalized PageRank diffusion to initial query-aligned seed passages and executing it strictly within this geometrically bounded subgraph, the framework mitigates semantic drift and facilitates noise-controlled evidence aggregation.
Our contributions are summarized as follows. (1) We propose ISO-RAG, a novel geometry-aware RAG framework that uses explicit topological control to mitigate spurious diffusion. (2) At the core of this framework, we introduce an isoperimetric control mechanism that explicitly prunes misleading connections to dense hubs prior to diffusion. (3) Extensive evaluations demonstrate that ISO-RAG achieves a highly favorable balance between retrieval efficiency and downstream QA performance, delivering robust average absolute gains of nearly 10% in Recall@5 and 4.3% in Exact Match over competitive baselines.
2 Related Works
Sparse and Dense Retrieval for Multi-Hop QA. Conventional retrieval encompasses sparse methods like BM25 (Robertson, Zaragoza et al. 2009) and dense models ranging from flat bi-encoders (Karpukhin et al. 2020) to multi-hop extensions like MDR (Xiong et al. 2021). Whether utilizing exact keyword matching or cosine similarity, these approaches operate in fundamentally flat search spaces. Compressing documents into isolated points optimized for direct semantic overlap, dense vectors cannot explicitly model relationships between intermediate entities. This structural limitation fragments reasoning by allowing lexically similar yet logically disconnected distractors to overshadow critical evidence. Consequently, flat search spaces struggle with the compositional reasoning paths required by increasingly difficult multi-hop QA datasets like HotpotQA (Yang et al. 2018), 2WikiMultihopQA (Ho et al. 2020), and MuSiQue (Trivedi et al. 2022).
Conventional Graph-based RAG Systems. To overcome flat semantic spaces, graph-based retrieval structures corpora into networks capturing multi-hop dependencies. Notable architectures include GraphRAG (Edge et al. 2024) utilizing hierarchical summaries with optional heuristic routing, LightRAG (Guo et al. 2024) employing dual-level structures, and HippoRAG2 (Jimenez Gutierrez et al. 2025) leveraging neurobiologically-inspired memory networks with continuous activation spreading. Despite improving recall, these baselines rely on heuristic edge weighting and unconstrained probability diffusion; for instance, HippoRAG2 executes unrestricted global PPR. Consequently, this uncontrolled diffusion introduces severe topological noise during retrieval (Alon and Yahav 2021; Topping et al. 2022). Without rigorous mathematical bounds pruning the search space, these frameworks inevitably retrieve spurious subgraphs and misleading entities before generation.
Hyperbolic Geometry and Continuous Aggregation. The inherent hierarchical structure of knowledge graphs makes them poorly suited for Euclidean embeddings (Nickel and Kiela 2017; Sala et al. 2018). Foundational graph neural networks therefore extend representation learning into hyperbolic space through architectures like HGCN (Chami et al. 2019) and related hyperbolic networks (Gulcehre et al. 2019; Peng et al. 2022). Mainstream methodologies use the Poincar’e ball model (Nickel and Kiela 2017) for its intuitive conformal geometry, or the Lorentz model (Nickel and Kiela 2018) for its numerical advantages in distance optimization. Both models provide exponential capacity to embed complex networks with minimal structural distortion. Recent works like HyperbolicRAG (Cao et al. 2025) attempt to leverage this by introducing hyperbolic representations into RAG. However, despite operating in hyperbolic space, its probability diffusion remains fundamentally unconstrained. By performing continuous neighborhood aggregation without explicit discrete pruning, HyperbolicRAG fails to resolve the topological bottleneck, inevitably accumulating topological noise and causing severe semantic drift during retrieval.
3 Methodology
We present ISO-RAG (ISOperimetric Retrieval-Augmented Generation), a geometry-aware local graph retrieval framework for multi-hop question answering. The core idea is to avoid broad graph-wide diffusion by combining four components: (i) query-aware seed routing, (ii) local candidate graph construction, (iii) isoperimetric structural filtering derived from hyperbolic structure, and (iv) deterministic personalized PageRank on the filtered local graph.
3.1 Problem Formulation
Let denote a corpus of passages. Given a multi-hop query , the retrieval objective is to extract a top- subset , such that the retrieved passages jointly cover the evidence required to answer .
We model the corpus as an undirected retrieval graph , where each node represents a passage equipped with a dense semantic embedding . serves as a question-induced passage co-occurrence graph: an edge is established whenever passages and co-occur within a training instance or retrieval context.
Such co-occurrence graphs effectively expose latent multi-hop dependencies, but they also inherently introduce noisy shortcuts and hub-like regions. As a result, unconstrained diffusion processes can prematurely leak probability mass into generic yet weakly useful passages, reducing both retrieval precision and downstream QA performance.
3.2 Seeded Local Graph Construction
Given a query , ISO-RAG first retrieves its dense embedding from a precomputed embedding cache and measures its cosine similarity to every node embedding:
| (1) |
Let denote the set of top- seed nodes selected according to :
| (2) |
Rather than propagating over the full graph, ISO-RAG restricts the search space to a compact local candidate set . To ensure both high recall and structural connectivity, we construct this set by integrating two components: a dense semantic pool and a topology-aware neighborhood.
Specifically, let denote the top- passages retrieved via (where , thus containing the seed set ). Let denote the -hop structural neighborhood expanded from . The local candidate set is defined as their union:
| (3) |
The induced local subgraph is .
This query-conditioned local graph substantially reduces the search space and transforms retrieval from global diffusion vulnerable to topological noise into local propagation anchored at semantically aligned entry points.
3.3 Hyperbolic Isoperimetric Edge Filtering
Hyperbolic structural signal. To characterize local graph structure, we map node representations into the Poincaré ball (Nickel and Kiela 2017; Chami et al. 2019; Balazevic, Allen, and Hospedales 2019):
| (4) |
where
| (5) |
and denotes the standard Euclidean norm. The mapping , which includes a projection onto the open unit ball, directly maps precomputed Euclidean text embeddings into the hyperbolic manifold. To align textual semantics with discrete multi-hop topology, is trained offline via a graph-supervised margin-based triplet loss and a radial depth regularizer (detailed in Appendix). This helps ensure the representations encapsulate both raw semantics and hierarchical structures.
For a node , let
| (6) |
denote the conformal factor of the Poincaré metric at . In Riemannian geometry, this factor dictates the volume expansion of the local space. Therefore, the geometric volume occupied by a node is intrinsically driven by this conformal scaling. We thus define the local volume proxy as:
| (7) |
Geometrically, quantifies this continuous spatial occupancy, which acts as an inverse indicator of semantic breadth. Due to the exponential outward expansion of the Poincaré ball, generic semantic hubs at the origin are tightly compressed into minimal conformal volumes, whereas highly specific factual entities at the periphery occupy vast spatial regions. In strict Riemannian geometry, the true volume scales with the dimensionality . However, computing for high-dimensional embeddings () inevitably leads to numerical overflow. To ensure system reliability, we introduce a tunable scaling exponent . This engineering formulation provides a numerically stable proxy for the node volume that flexibly controls the dynamic range of the structural signal. By doing so, we preserve the core monotonic conformal scaling property of hyperbolic space while guaranteeing robust online computation.
To construct a principled structural ratio, we define a localized geometric measure based on the classical Cheeger constant. For any node , let denote its closed 1-hop neighborhood set. We define the internal geometric volume of this local region as the sum of the conformal volume proxies of its constituent nodes:
| (8) |
Next, we identify the topological boundary of this region. Let denote the 2-hop boundary shell of , consisting of all nodes adjacent to that are not contained within . The geometric volume of this boundary is analogously defined as:
| (9) |
Having formalized both the internal and boundary volumes using the exact same hyperbolic measure, we define the node-wise structural pruning score as the localized geometric Cheeger ratio:
| (10) |
where is a small stability constant. Unlike heuristic formulations that mix mismatched topological and geometric scales, this definition strictly preserves the isoperimetric nature of the score. The proxy measures the relative geometric expansion of a node’s local neighborhood. Generic hubs exhibit explosive boundary volumes () compared to their highly compressed internal neighborhood volumes (), yielding extremely large values. By evaluating this rigorous volume-to-volume ratio, the framework can explicitly identify structural bottlenecks without requiring computationally prohibitive global graph partitioning.
Discriminative Power of the Proxy. In scale-free hyperbolic embeddings, a node’s radial distance inversely tracks its topological degree, while its conformal volume grows exponentially with radius. This dual scaling creates a geometric mismatch for generic hubs: they reside near the origin with heavily compressed individual volumes, yet their massive topological connectivity bridges to numerous specific nodes at the periphery. Consequently, for a central hub, its 2-hop boundary shell reaches into the expansive periphery, accumulating an explosive boundary volume , while its internal 1-hop neighborhood volume remains heavily constrained. This extreme volume-to-volume divergence triggers massive anomalies for spurious edges connecting factual nodes to unrelated hubs, enabling ISO-RAG to structurally isolate probability leakage. Detailed mathematical formulations are provided in Appendix.
Edge Compatibility Filtering. To prune spurious edges bridging structurally dissimilar nodes, we enforce structural consistency between endpoints. An edge is retained only if:
| (11) |
where is a validation-tuned tolerance threshold. Equivalently, in log-scale:
| (12) |
Here, bounds the maximum allowable structural divergence. Valid reasoning steps between nodes of comparable specificity exhibit small -divergences. Conversely, edges bridging opposite structural extremes (e.g., direct transitions between specific factual leaves and generic hubs) inevitably violate this threshold and are systematically pruned. This dual-sided filtering mitigates semantic drift from two directions: it blocks forward probability leakage into hubs during propagation, and isolates erroneously retrieved hub anchors before diffusion begins. Because all node-wise structural scores can be fully precomputed and cached, this geometric filtering introduces negligible online computational latency. We provide details in Appendix.
Let
| (13) |
denote the resulting geometrically filtered local graph, which provides a structurally coherent and bounded manifold for the subsequent localized PageRank.
3.4 Localized Deterministic Personalized PageRank
After edge filtering, retrieval operates strictly on the compact graph . We re-normalize its adjacency matrix to derive a valid column-stochastic transition matrix .
The personalization vector distributes the initial probability mass exclusively across the selected seed nodes via a temperature-scaled softmax:
| (14) |
where controls the mass concentration sharpness, and denotes the initial dense semantic similarity score.
The localized PPR vector is defined as the unique fixed point of the diffusion process:
| (15) |
where is the damping factor (with acting as the teleportation probability). During diffusion, the probability mass of dangling nodes is intrinsically redistributed via , which falls back to a uniform distribution if seed weights degrade to zero. Rather than relying on stochastic random walks that introduce approximation variance, we compute deterministically via power iteration. Because our structural filtering strictly bounds the diffusion space to the compact local subgraph , this exact computation remains highly efficient and yields stable structural scores.
Finally, the converged structural scores are fused with the initial dense semantic similarities. Denoting the scalar structural score for node as , which is extracted from the vector , the final retrieval score for each passage is computed via a linear combination:
| (16) |
where and are min-max normalized over the candidate set prior to fusion, and and are tunable balancing weights. The candidate passages are subsequently ranked by to extract the optimal top- evidence set for downstream QA. This dual-signal fusion elegantly couples the semantic recall of dense models with the structural multi-hop precision of our isoperimetric framework, ensuring that the final ranking is both contextually relevant and topologically coherent.
4 Experiments
| Model | Retriever | HotpotQA | 2Wiki- MultihopQA | MuSiQue | |||
| F1 | EM | F1 | EM | F1 | EM | ||
| Qwen2.5 | GraphRAG | 79.1 | 72.3 | 54.1 | 52.8 | 28.8 | 23.2 |
| GraphRAG+PPR | 75.0 | 68.4 | 54.5 | 52.9 | 30.6 | 24.2 | |
| LightRAG | 79.8 | 72.3 | 72.2 | 68.9 | 34.0 | 28.4 | |
| HippoRAG2 | 80.6 | 73.1 | 62.3 | 59.7 | 36.9 | 30.8 | |
| HyperbolicRAG | 79.6 | 72.3 | 60.0 | 58.2 | 30.6 | 25.6 | |
| ISO-RAG | 81.1 | 74.1 | 76.9 | 73.2 | 40.1 | 33.9 | |
| Qwen- Plus | GraphRAG | 80.4 | 72.8 | 56.7 | 55.2 | 28.7 | 23.3 |
| GraphRAG+PPR | 78.1 | 70.8 | 57.5 | 55.2 | 30.1 | 24.0 | |
| LightRAG | 82.3 | 75.0 | 72.8 | 68.8 | 30.9 | 24.6 | |
| HippoRAG2 | 82.2 | 74.6 | 63.7 | 60.9 | 34.2 | 27.7 | |
| HyperbolicRAG | 81.4 | 73.7 | 61.6 | 59.4 | 32.1 | 25.4 | |
| ISO-RAG | 82.5 | 75.0 | 80.4 | 75.8 | 42.8 | 31.4 | |
| Qwen3- Max | GraphRAG | 83.1 | 76.0 | 60.9 | 59.2 | 28.7 | 23.3 |
| GraphRAG+PPR | 81.0 | 74.1 | 59.5 | 57.7 | 30.1 | 24.0 | |
| LightRAG | 84.6 | 77.4 | 77.0 | 73.8 | 33.6 | 28.3 | |
| HippoRAG2 | 84.8 | 77.7 | 67.4 | 64.6 | 36.9 | 30.9 | |
| HyperbolicRAG | 83.9 | 76.5 | 64.5 | 62.5 | 34.6 | 28.8 | |
| ISO-RAG | 85.2 | 78.2 | 84.7 | 80.7 | 40.2 | 33.7 | |
In this section, we comprehensively evaluate ISO-RAG to answer the following Research Questions (RQs): RQ1 (Overall Performance): Does ISO-RAG outperform existing dense and graph-based retrieval methods in both retrieval accuracy and downstream multi-hop QA? RQ2 (Efficiency): Can the localized deterministic routing paradigm achieve better retrieval-time efficiency compared to unconstrained graph traversals? RQ3 (Geometric Filtering): How does the isoperimetric signal () explicitly contribute to noise-aware structural filtering?
| Method | HotpotQA | 2WikiMultihopQA | MuSiQue | |||||||||
| R@5 | R@10 | P@5 | P@10 | R@5 | R@10 | P@5 | P@10 | R@5 | R@10 | P@5 | P@10 | |
| BM25 | 62.70 | 75.25 | 25.08 | 15.05 | 39.80 | 48.20 | 18.42 | 11.08 | 27.00 | 32.15 | 10.64 | 6.34 |
| Flat Dense | 81.90 | 87.95 | 32.76 | 17.59 | 69.43 | 71.58 | 31.82 | 16.39 | 52.40 | 58.95 | 20.54 | 11.57 |
| MDR | 78.80 | 88.15 | 31.52 | 17.63 | 61.65 | 68.55 | 27.70 | 15.53 | 51.70 | 59.45 | 20.34 | 11.70 |
| Vanilla PPR | 81.75 | 94.40 | 32.70 | 18.88 | 69.05 | 78.95 | 31.80 | 18.96 | 70.10 | 77.30 | 27.12 | 14.98 |
| GraphRAG | 85.25 | 96.20 | 34.10 | 19.24 | 71.15 | 88.45 | 32.84 | 21.34 | 59.00 | 72.30 | 22.84 | 13.99 |
| GraphRAG+PPR | 71.45 | 94.20 | 28.58 | 18.84 | 69.78 | 81.95 | 32.38 | 19.91 | 63.20 | 76.15 | 24.40 | 14.74 |
| LightRAG | 87.30 | 97.20 | 34.92 | 19.44 | 77.93 | 88.80 | 35.92 | 20.95 | 55.05 | 67.05 | 21.60 | 13.16 |
| HippoRAG2 | 86.65 | 97.85 | 34.66 | 19.57 | 70.18 | 82.80 | 32.30 | 19.88 | 63.85 | 75.40 | 26.14 | 14.64 |
| HyperbolicRAG | 85.30 | 96.85 | 34.12 | 19.37 | 70.45 | 78.38 | 32.34 | 18.55 | 55.35 | 76.05 | 21.70 | 14.70 |
| ISO-RAG | 88.05 | 98.05 | 35.22 | 19.61 | 88.10 | 97.15 | 40.10 | 22.90 | 74.75 | 84.00 | 28.98 | 16.31 |
4.1 Experimental Setup
Datasets & Graph Construction. We evaluate ISO-RAG on three standard multi-hop QA benchmarks with varying reasoning complexities: HotpotQA (Yang et al. 2018) (primarily 2-hop), 2WikiMultihopQA (Ho et al. 2020) (2–4 hop), and MuSiQue (Trivedi et al. 2022) (up to 4-hop compositional reasoning). To establish a unified evaluation setting for both structural retrieval and end-to-end QA, we randomly sample 1,000 instances from the validation set of each benchmark. For graph construction, rather than building a traditional entity-relation knowledge graph, we construct a question-induced passage co-occurrence graph (details in Appendix. Passages are treated as nodes, with edges connecting passages that co-occur in the same training instance. Node texts are embedded offline using the text-embedding-v3 encoder (Zhang et al. 2025). To strictly prevent data leakage, the passage co-occurrence graphs are constructed exclusively using the training splits of the respective datasets. Validation and test sets are completely excluded from the graph construction phase, ensuring that the retrieval framework does not benefit from any benchmark-specific structural shortcuts.
Table 3 compares retrieval latency and token usage under Qwen3-Max to assess the computational advantage of our localized pipeline.
Baselines. To comprehensively evaluate retrieval quality, we categorize baselines into non-graph methods (BM25 (Robertson, Zaragoza et al. 2009), Flat Dense (Karpukhin et al. 2020), MDR (Xiong et al. 2021)) and graph-based paradigms. The latter ranges from foundational algorithms (Vanilla PPR (Page et al. 1999)) to state-of-the-art frameworks (GraphRAG (Edge et al. 2024), GraphRAG+PPR (Page et al. 1999), LightRAG (Guo et al. 2024), HippoRAG2 (Jimenez Gutierrez et al. 2024), and HyperbolicRAG (Cao et al. 2025)). As Vanilla PPR is a foundational retrieval algorithm rather than an end-to-end RAG pipeline, we focus our QA assessment exclusively on the aforementioned state-of-the-art frameworks designed for generative tasks. To execute this generation, the retrieval outputs are paired with multiple LLM backbones (Qwen2.5, Qwen-Plus (Team 2024), and Qwen3-Max (Team 2025)).
Implementation Details. We evaluate downstream QA performance using standard Exact Match (EM) and F1 scores. For fair comparison, all graph frameworks supply the raw text of their retrieved top- passages to the LLM generator using an identical prompt template. Comprehensive implementation details, including hyperparameter tuning (e.g., PageRank ), exact grid-search ranges, and full prompt templates, are detailed in Appendix.
4.2 Main Results: Retrieval and QA Performance (RQ1)
We first evaluate the fundamental retrieval capability (Tables 2) and the end-to-end QA generation quality (Table 1). ISO-RAG consistently yields the best results across all settings.
Consistent Gains in Shallow Reasoning. While performance on HotpotQA is nearing saturation for high-capacity models, ISO-RAG still guarantees stable improvements, achieving 88.05 R@5 (+0.75 absolute points over the strongest baseline) and peaking at 85.2 F1 score with Qwen3-Max. Notably, the improvements are robust across model scales, suggesting that our retrieval framework itself fundamentally drives the observed gains.
Superiority in Complex Topologies. Table 1 demonstrates that ISO-RAG outperforms all baselines on 2WikiMultihopQA. Notably, it surpasses the strongest baseline, LightRAG, achieving an R@5 of 88.10 with a 10.17-point absolute improvement. Compared to recent topology-driven frameworks (i.e., HippoRAG2 and HyperbolicRAG), this gap widens, yielding a relative R@5 gain exceeding 25%. Furthermore, QA evaluation using Qwen3-Max yields an F1 score of 84.7 and an EM score of 80.7. These results confirm the efficacy of ISO-RAG in structured multi-hop scenarios, where preserving valid intermediate hops and suppressing spurious diffusion are critical.
Robustness against High Noise. The MuSiQue dataset presents a severe challenge of deep compositional reasoning amidst dense distractors. In this regime, while ISO-RAG achieves a state-of-the-art R@10 score of 84.00 (+11.4% relative gain over HippoRAG2) alongside highly competitive precision of 16.31, its full potential materializes in the QA phase. Evaluated with Qwen-Plus, the framework peaks at an F1 score of 42.8, securing an approximate 25% relative gain over HippoRAG2. This disproportionate amplification, transitioning from a steady retrieval improvement to a drastic QA leap, demonstrates that ISO-RAG does not merely accumulate disjoint relevant documents. Instead, driven by its high retrieval precision, it successfully isolates the precise, noise-free multi-hop evidence pathways required to cross the reasoning threshold of the LLM.
4.3 Efficiency Analysis (RQ2)
| Dataset | Method | Efficiency | |
| Retr/Q (ms) | Avg Prompt | ||
| HotpotQA | GraphRAG | 10.8 | 1619.91 |
| GraphRAG+PPR | 15.1 | 1652.10 | |
| LightRAG | 22.9 | 1644.83 | |
| HippoRAG2 | 52.0 | 895.16 | |
| HyperbolicRAG | 92.7 | 889.61 | |
| ISO-RAG | 12.3 | 886.44 | |
| 2Wiki- MultihopQA | GraphRAG | 6.2 | 1389.09 |
| GraphRAG+PPR | 8.7 | 1225.13 | |
| LightRAG | 14.4 | 1271.18 | |
| HippoRAG2 | 49.1 | 720.89 | |
| HyperbolicRAG | 71.1 | 758.91 | |
| ISO-RAG | 11.5 | 793.49 | |
| MuSiQue | GraphRAG | 9.9 | 1579.64 |
| GraphRAG+PPR | 16.2 | 1588.76 | |
| LightRAG | 65.6 | 1561.58 | |
| HippoRAG2 | 117.2 | 912.46 | |
| HyperbolicRAG | 375.1 | 900.95 | |
| ISO-RAG | 14.2 | 927.07 | |
Latency Reduction via Local Subgraphs. ISO-RAG maintains stable millisecond-level speeds across all benchmarks, ranking as the second fastest retriever on both HotpotQA and MuSiQue. While marginally trailing the vanilla GraphRAG baseline in raw speed, it provides a substantially more precise reasoning context. Compared to recent topology-driven frameworks, the latency gap is particularly striking: on MuSiQue, ISO-RAG delivers an approximate 8 speedup over HippoRAG2 and operates over 25 faster than HyperbolicRAG. This confirms that bounding PageRank within a geometrically filtered subgraph successfully bypasses the heavy overhead of global traversals.
Favorable Accuracy-Efficiency Trade-off. ISO-RAG exhibits highly competitive token efficiency, incurring the lowest prompt token overhead among all baselines on HotpotQA. While structurally complex datasets require marginally more tokens than HippoRAG2 and HyperbolicRAG, this slight increment is fully offset by substantial gains in retrieval accuracy. Ultimately, this confirms that ISO-RAG delivers a strictly higher density of actionable reasoning chains per token, ensuring a highly cost-effective retrieval process.
4.4 Ablation: Isoperimetric Geometric Filtering (RQ3)
We isolate the isoperimetric filtering module on 2WikiMultihopQA, where the impact of geometric guidance is most pronounced, to compare true geometric guidance against unconstrained, random, or topologically decoupled edge pruning under highly noisy conditions (Figure 3).
We define three core variables to control the ablation space: Candidate , the over-retrieval factor applied to the initial dense pool prior to -filtering; , the threshold controlling geometric edge pruning strictness; and four mapping strategies: (1) real_phi applies actual learned scores to capture true local continuity; (2) uniform_phi assigns a constant value, disabling the filter to revert to vanilla PPR; (3) shuffled_phi randomly permutes real_phi values, preserving global distribution but destroying correlation with graph topology; and (4) random_phi applies uniform random values to test arbitrary edge pruning. The performance is evaluated using recall, precision, and the percentage of removed edges.
Necessity of the Real Geometric Signal. real_phi dominates all variants, achieving 87.28% peak recall and 40.44% peak precision at . Conversely, uniform_phi (no filtering) and random_phi (random scalar field) stagnate at 71-72%. This demonstrates that retrieval gains stem from the learned geometric signal, not merely baseline graph topology (uniform_phi) or random edge dropout regularization (random_phi).
Necessity of Topological Alignment. shuffled_phi preserves the true numerical distribution of isoperimetric scores but decouples them from graph topology. Consequently, it aggressively removes up to 75% of edges and suffers a massive performance drop, whereas ISO-RAG achieves peak results by pruning only 32%. This proves multi-hop retrieval requires topology-aware filtering, not indiscriminate pruning. By strictly aligning the isoperimetric signal with graph structure, ISO-RAG selectively prunes edges leaking probability mass into irrelevant neighborhoods.
4.5 Qualitative Case Study
To illustrate how geometric filtering prevents probability leakage in PPR, we examine a multi-hop query from 2WikiMultihopQA (Figure 2). This query requires a parallel 2-hop reasoning chain: identifying directors for two films, retrieving their biographies, and comparing their birth dates.
Preventing Probability Leakage. In the top-5 contexts, baseline methods fail to retrieve the crucial passages; their unconstrained PageRank diffusion is hijacked by dense, semantically adjacent hub nodes (e.g., unrelated films or directors). Consequently, HippoRAG2 outputs "Unknown", while HyperbolicRAG hallucinates the wrong film. Conversely, ISO-RAG’s isoperimetric control severs spurious edges to these hubs. By bounding probability mass within the local manifold, it retrieves both required passages, enabling the LLM to deduce the correct answer.
Decoupling Retrieval and Generation Errors. To further illustrate the vulnerability of downstream LLMs to contextual noise, we present a "Perfect Retrieval, Failed Generation" case in Figure 2. Although ISO-RAG successfully retrieved 100% of the ground-truth entities, the LLM still failed. This demonstrates that degraded F1/EM scores are not exclusively indicative of retrieval failure, but can also stem from LLM reasoning bottlenecks.
5 Conclusion
In this work, we presented ISO-RAG, a novel retrieval-augmented generation framework that imposes strict topological control to mitigate spurious diffusion during graph-based retrieval. By mapping the graph into hyperbolic space and training a geometry-aware encoder, ISO-RAG enables controlled retrieval that is both efficient and effective for downstream question answering. Across benchmarks, it yields consistent absolute improvements over competitive baselines, underscoring the importance of topology-aware control for multi-hop reasoning.
References
- Alon and Yahav (2021) Alon, U.; and Yahav, E. 2021. On the Bottleneck of Graph Neural Networks and Its Practical Implications. In International Conference on Learning Representations.
- Andersen, Chung, and Lang (2006) Andersen, R.; Chung, F.; and Lang, K. 2006. Local Graph Partitioning Using PageRank Vectors. In 47th Annual IEEE Symposium on Foundations of Computer Science (FOCS’06), 475–486. IEEE.
- Balazevic, Allen, and Hospedales (2019) Balazevic, I.; Allen, C.; and Hospedales, T. 2019. Multi-Relational Poincaré Graph Embeddings. In Advances in Neural Information Processing Systems, volume 32.
- Cao et al. (2025) Cao, L.; Wang, R.; Li, J.; Zhou, Z.; and Yang, M. 2025. HyperbolicRAG: Enhancing Retrieval-Augmented Generation with Hyperbolic Representations. arXiv:2511.18808.
- Chami et al. (2019) Chami, I.; Ying, Z.; Ré, C.; and Leskovec, J. 2019. Hyperbolic Graph Convolutional Neural Networks. In Advances in Neural Information Processing Systems, volume 32.
- Chen et al. (2023) Chen, J.; Lin, H.; Han, X.; and Sun, L. 2023. Benchmarking Large Language Models in Retrieval-Augmented Generation. arXiv:2309.01431.
- Chung (1997) Chung, F. R. 1997. Spectral Graph Theory, volume 92. American Mathematical Society.
- Edge et al. (2024) Edge, D.; Trinh, H.; Cheng, N.; Bradley, J.; Chao, A.; Mody, A.; Truitt, S.; and Larson, J. 2024. From Local to Global: A Graph RAG Approach to Query-Focused Summarization. arXiv:2404.16130.
- Gao et al. (2023) Gao, Y.; Xiong, Y.; Gao, X.; Jia, K.; Pan, J.; Bi, Y.; Dai, Y.; Sun, J.; and Wang, H. 2023. Retrieval-Augmented Generation for Large Language Models: A Survey. arXiv:2312.10997.
- Girvan and Newman (2002) Girvan, M.; and Newman, M. E. 2002. Community Structure in Social and Biological Networks. Proceedings of the National Academy of Sciences, 99(12): 7821–7826.
- Gulcehre et al. (2019) Gulcehre, C.; Denil, M.; Cabukuetli, M.; Pfau, D.; Pascanu, R.; Hoffman, M. W.; and Nando, d. F. 2019. Hyperbolic Attention Networks. In International Conference on Learning Representations.
- Guo et al. (2024) Guo, Z.; Xia, L.; Yu, Y.; Ao, T.; and Huang, C. 2024. LightRAG: Simple and Fast Retrieval-Augmented Generation. arXiv:2410.05779.
- Guu et al. (2020) Guu, K.; Gurkurun, T.; Knecht, D.; Chang, M.-W.; and Salakhutdinov, R. 2020. REALM: Retrieval-Augmented Language Model Pre-training. In International Conference on Machine Learning, 3929–3938. PMLR.
- Ho et al. (2020) Ho, X.; Nguyen, A.-K. D.; Sugawara, S.; and Aizawa, A. 2020. Constructing a Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps. In Proceedings of the 28th International Conference on Computational Linguistics, 6609–6625.
- Ji et al. (2023) Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y. J.; Madotto, A.; and Fung, P. 2023. Survey of Hallucination in Natural Language Generation. ACM Computing Surveys, 55(12): 1–38.
- Jiang et al. (2023) Jiang, Z.; Xu, F. F.; Gao, L.; Sun, Z.; Liu, Q.; Dwivedi-Yu, J.; Yang, Y.; Jamie, C.; and Neubig, G. 2023. Active Retrieval Augmented Generation. arXiv:2305.06983.
- Jimenez Gutierrez et al. (2024) Jimenez Gutierrez, B.; Shu, Y.; Gu, Y.; Yasunaga, M.; and Su, Y. 2024. HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models. In Advances in Neural Information Processing Systems, volume 37, 59532–59569.
- Jimenez Gutierrez et al. (2025) Jimenez Gutierrez, B.; Shu, Y.; Gu, Y.; Yasunaga, M.; and Su, Y. 2025. From RAG to Memory: Non-Parametric Continual Learning for Large Language Models. arXiv:2502.14802.
- Kaplan et al. (2020) Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T. B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; and Amodei, D. 2020. Scaling Laws for Neural Language Models. arXiv:2001.08361.
- Karpukhin et al. (2020) Karpukhin, V.; Oguz, B.; Min, S.; Lewis, P.; Wu, L.; Edunov, S.; Chen, D.; and Yih, W.-t. 2020. Dense Passage Retrieval for Open-Domain Question Answering. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 6769–6781.
- Krioukov et al. (2010) Krioukov, D.; Papadopoulos, F.; Kitsak, M.; Vahdat, A.; and Boguná, M. 2010. Hyperbolic Geometry of Complex Networks. Physical Review E, 82(3): 036106.
- Lewis et al. (2020) Lewis, P.; Perez, E.; Piktus, A.; Petroni, F.; Karpukhin, V.; Goyal, N.; Küttler, H.; Lewis, M.; Yih, W.-t.; Rocktäschel, T.; Riedel, S.; and Kiela, D. 2020. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. In Advances in Neural Information Processing Systems, volume 33, 9459–9474.
- Nickel and Kiela (2017) Nickel, M.; and Kiela, D. 2017. Poincaré Embeddings for Learning Hierarchical Representations. In Advances in Neural Information Processing Systems, volume 30.
- Nickel and Kiela (2018) Nickel, M.; and Kiela, D. 2018. Learning Continuous Hierarchies in the Lorentz Model of Hyperbolic Geometry. In International Conference on Machine Learning, 3779–3788.
- Page et al. (1999) Page, L.; Brin, S.; Motwani, R.; and Winograd, T. 1999. The PageRank Citation Ranking: Bringing Order to the Web. Technical report, Stanford InfoLab.
- Peng et al. (2022) Peng, W.; Varanka, T.; Mostafa, A.; Shi, H.; and Zhao, G. 2022. Hyperbolic Deep Neural Networks: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(12): 10023–10044.
- Robertson, Zaragoza et al. (2009) Robertson, S.; Zaragoza, H.; et al. 2009. The Probabilistic Relevance Framework: BM25 and Beyond. Foundations and Trends® in Information Retrieval, 3(4): 333–389.
- Sala et al. (2018) Sala, F.; De Sa, C.; Gu, A.; and Ré, C. 2018. Representation Tradeoffs for Hyperbolic Embeddings. In International Conference on Machine Learning, 4460–4469. PMLR.
- Team (2024) Team, Q. 2024. Qwen2.5 Technical Report. arXiv:2412.15115.
- Team (2025) Team, Q. 2025. Qwen3 Technical Report. arXiv:2505.09388.
- Topping et al. (2022) Topping, J.; Di Giovanni, F.; Chamberlain, B. P.; Dong, X.; and Bronstein, M. M. 2022. Understanding Over-Squashing and Bottlenecks on Graphs via Ricci Curvature. In International Conference on Learning Representations.
- Trivedi et al. (2022) Trivedi, H.; Balasubramanian, N.; Khot, T.; and Sabharwal, A. 2022. MuSiQue: Multihop Questions via Single-hop Question Composition. Transactions of the Association for Computational Linguistics, 10: 539–554.
- Vaswani et al. (2017) Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017. Attention Is All You Need. In Advances in Neural Information Processing Systems, 5998–6008.
- Wei et al. (2022) Wei, J.; Wang, X.; Schuurmans, D.; Maeda, M.; Xia, F.; Chi, E.; Le, Q. V.; and Zhou, D. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. In Advances in Neural Information Processing Systems, volume 35, 24824–24837.
- Xiong et al. (2021) Xiong, W.; Li, X. L.; Iyer, S.; Du, J.; Lewis, P.; Wang, W. Y.; Yashar, M.; Yih, W.-t.; Riedel, S.; Douze, M.; et al. 2021. Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval. In International Conference on Learning Representations (ICLR).
- Yang et al. (2018) Yang, Z.; Qi, P.; Zhang, S.; Bengio, Y.; Cohen, W. W.; Salakhutdinov, R.; and Manning, C. D. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 2369–2380.
- Yao et al. (2023) Yao, S.; Yu, D.; Zhao, J.; Shafran, I.; McManus, T. G.; Isbell, R.; and Narasimhan, K. 2023. Tree of Thoughts: Deliberate Problem Solving with Large Language Models. In Advances in Neural Information Processing Systems.
- Yao et al. (2022) Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; and Cao, Y. 2022. ReAct: Synergizing Reasoning and Acting in Language Models. In International Conference on Learning Representations.
- Zhang et al. (2025) Zhang, Y.; Li, M.; Long, D.; Zhang, X.; Lin, H.; Yang, B.; Xie, P.; Yang, A.; Liu, D.; Lin, J.; Huang, F.; and Zhou, J. 2025. Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models. arXiv:2506.05176.