SpecMind: Enabling Spectrum Intelligence via Multi-Agent Hybrid Retrieval-Augmented Generation
Thanks: This work was supported in part by SpectrumX, the National Science Foundation (NSF) Spectrum Innovation Center, through grant AST 2132700 operated under Cooperative Agreement by the University of Notre Dame.Thanks: ∗The first two authors contributed equally to this work.
Abstract
The exponential growth of wireless devices is driving unprecedented spectrum demand, pushing spectrum management toward more fine-grained decisions across space, time, and device constraints. As a result, spectrum policymakers and engineers must process large volumes of data that come from diverse sources and take many different forms, such as text and tables. These data sources are often disaggregated and require significant time and effort to integrate, search, and interpret. Furthermore, most of this information is formatted for human understanding and is not readily accessible to automated systems. To address this challenge, we propose SpecMind, a novel Multi-Agent Retrieval-Augmented Generation (RAG) system for spectrum intelligence that performs reasoning over heterogeneous data sources. This system enables autonomous agents to coordinate specialized sub-agents that retrieve and synthesize knowledge across policy proceedings, legal regulations, and license databases. We develop SpecBench, a question and answer (Q&A) dataset based on real-world license records and policy proceedings, addressing the lack of evaluation resources for RAG systems in the spectrum domain. Experimental results demonstrate that SpecMind outperforms traditional, general-purposed RAG systems across spectrum-related tasks, achieving over 80% win rate against strong baselines. The agent-based design enables more accurate retrieval, better contextual reasoning, and improved task completion across diverse query types.
Index Terms:
spectrum management, agentic AI, retrieval-augmented generation, large language modelsI Introduction
The wireless spectrum faces increasing demand as modern services and devices compete for limited bandwidth. Spectrum management has consequently become more complex, involving constraints across frequency, time, location, users, and regulatory requirements [1]. This complexity drives rapid growth in spectrum data in both scale and heterogeneity. Routine tasks such as transmission compliance verification and spectrum policy development require integrating fragmented, heterogeneous information. These challenges necessitate scalable and context-aware methods for spectrum data interpretation and decision support [2].
Large language models (LLMs) have demonstrated strong capabilities in domain-specific understanding and reasoning [3, 4], while retrieval-augmented generation (RAG) improves grounding by incorporating external knowledge sources [5, 6, 7]. Recent work has begun to extend RAG to spectrum and telecommunications domains. SpectrumRAG [8] introduces an iterative framework that leverages query rewriting to refine retrieval. TelcoRAG [9] enhances retrieval over telecommunications standards through query augmentation and specialized pipelines. Radio Regulations GPT [10] proposes a domain-specific RAG pipeline for regulatory understanding. However, these approaches remain largely centered on text-based retrieval and lack adaptive mechanisms for integrating heterogeneous, multi-source data, limiting their effectiveness on increasingly complex tasks.
To address these limitations, we propose SpecMind, an Agentic RAG framework for spectrum intelligence over heterogeneous data. Our key contributions include the construction of a spectrum-specific corpus, the design of specialized tools and prompts, and the development of an agentic framework that orchestrates these components for complex reasoning and task execution that supports tasks such as incumbent investigation, stakeholder analysis, and regulatory compliance verification. We further introduce a benchmark for joint license and proceeding tasks and conduct a systematic evaluation of SpecMind. Our contributions are summarized as follows:
- •
Reusable, extensible spectrum knowledge databases: We construct a unified corpus of FCC licensing records, proceeding documents, and regulatory texts as complementary machine-readable resources. Licensing data are organized in a relational SQL database, proceedings are modeled as graphs capturing entities and cross-document relations, and regulatory texts are indexed via embeddings for retrieval. The resulting corpus is reusable and extensible, enabling efficient integration of new records.
- •
An expert-designed benchmark for heterogeneous spectrum-focused RAG: We introduce SpecBench, an expert-designed benchmark dataset for evaluating RAG systems over heterogeneous spectrum data. It comprises 450 curated question–answer pairs reflecting realistic spectrum analysis tasks, including fact verification, cross-source synthesis, and multi-hop reasoning. Built upon the proposed databases, SpecBench enables standardized and reproducible evaluation and serves as a valuable resource for spectrum-oriented RAG.
- •
A multi-agent hybrid RAG system for heterogeneous sources: To the best of our knowledge, SpecMind is the first multi-agent hybrid RAG framework for spectrum intelligence that enables coordinated retrieval and reasoning over heterogeneous spectrum data, including tabular, graph-structured, and textual sources. By supporting modality-aware retrieval and cross-source reasoning, SpecMind improves response accuracy and completeness over existing spectrum-oriented RAG pipelines, achieving over 80% win rate against strong baselines.
We have publicly released the SpecBench dataset, the reusable spectrum databases, and the SpecMind framework implementation to support reproducibility and community research [11].
II Methodology
II-A Database Construction
The construction process is illustrated in Figure 1. We consider three primary data sources in the spectrum domain: (1) license data, which appear as structured tables with numerical attributes; (2) proceeding documents, released by the FCC, each centered on a specific topic and containing diverse comments from individuals and organizations during regulatory deliberation; and (3) regulatory documents, which codify existing rules and policies. Applying RAG to such heterogeneous sources requires constructing modality-aware databases tailored to their distinct characteristics, rather than relying on a general-purposed semantic retrieval pipeline.
For regulatory texts, we adopt a standard dense retrieval paradigm. Documents are segmented into fixed-length chunks (1000 tokens with 200-token overlap) to preserve semantic continuity. Given a query, we retrieve top- candidates via embedding similarity and further apply a reranking stage to refine relevance. Specifically, we select the top 20 candidates from initial retrieval and rerank them to obtain the top 5 final contexts. This design aligns with the nature of regulatory texts, which are concise and legally precise, making queries rely heavily on accurate semantic matching.
In contrast, FCC proceeding comments and reply comments exhibit high volume, topic diversity, and complex multi-entity interactions. A single entity may express inconsistent positions across proceedings, and many queries require cross-document synthesis (e.g., identifying agency stances or comparing stakeholder views). Such requirements are poorly served by chunk-level semantic retrieval. We therefore enhance GraphRAG [12], which models documents as graphs with entities as nodes and relations as edges, and construct hierarchical summaries to support multi-hop reasoning. We build a separate graph database for each proceeding to ensure structural consistency (e.g., avoiding cross-document conflicts introduced by merging multiple proceedings into a single graph) and enable modular extensibility without rebuilding the entire corpus when new proceedings are introduced. Detailed statistics of the graph corpora, including node distributions across hierarchy levels, are summarized in Table I.
License data present a fundamentally different challenge, as they are dominated by structured numerical fields with weak semantic signals [13]. Traditional RAG methods perform poorly in this setting due to their reliance on semantic similarity. We instead model license data using relational databases and perform retrieval via SQL queries to enable exact matching over structured attributes. The problem is thus reformulated as generating executable SQL queries from natural language questions, which can be effectively handled by LLMs. Concretely, we extract license records into intermediate JSON representations for preprocessing, and then convert them into a relational SQL database. We further organize the database into service-specific subtables to accommodate heterogeneous schemas across different license types.
| Proceeding | Number of Nodes | ||||
| Level 0 | Level 1 | Level 2 | Level 3 | Level 4 | |
| FCC 19 - 38 | 477 | 445 | 133 | 57 | - |
| FCC 24 - 72 | 375 | 329 | 163 | 23 | - |
| FCC 25 - 59 | 1045 | 1008 | 419 | 40 | - |
| FCC 22 - 352 | 2094 | 2051 | 1599 | 571 | - |
| FCC 23 - 158 | 902 | 868 | 405 | 53 | - |
| FCC 23 - 232 | 871 | 794 | 338 | 82 | 58 |
| NTIA NSS | 4079 | 3960 | 3074 | 679 | 345 |
II-B Framework Design
We propose SpecMind, a multi-agent hybrid RAG system for reasoning over heterogeneous spectrum data. The overall system design is illustrated in Figure 2. Each task instance is represented as
where denotes a user query, is the ground-truth answer, and comprises heterogeneous knowledge sources, including structured license databases, graph-structured proceeding documents, and unstructured regulatory texts. The agent set is defined as
Each agent is associated with an action policy . The Supervisor Agent coordinates task decomposition, agent invocation, and result integration, while specialized agents perform domain-specific actions.
Given a query , the system maintains a global task state that accumulates intermediate evidence and partial results across multiple agent interactions. The Supervisor Agent coordinates the execution of specialized agents and iteratively refines the task state until a final answer is produced.
The objective of the system is to generate an answer that is factually grounded in and consistent with the ground-truth answer , while enabling accurate and flexible reasoning across heterogeneous spectrum data sources.
II-C Prompt Engineering
Design principle. We formulate prompt engineering as the explicit specification of agent-level action policies , which govern task decomposition, tool usage, and cross-agent coordination. Inspired by the ReAct paradigm [14], which interleaves reasoning and action, our design extends this principle to a multi-agent setting tailored to spectrum-specific tasks. Rather than relying on implicit reasoning in LLMs, we enforce structured decision-making through prompt-level control, enabling reliable interaction with heterogeneous data sources.
Supervisor prompt. The Supervisor Agent is prompted to follow a structured Think–Act–Observe loop for multi-step reasoning. The prompt explicitly instructs the agent to decompose the query into subtasks, decide whether to invoke a specialized agent or perform an internal operation, and iteratively update the global task state. This design enables adaptive planning and dynamic routing across heterogeneous tools, while ensuring that intermediate results are incorporated into subsequent reasoning steps.
License agent prompt. For structured license data, we design a tool-aware prompt that explicitly specifies the usage of four SQL tools, including schema inspection, query execution, and query validation. The action policy enforces a stepwise interaction pattern, where the agent first inspects database structure, then generates executable SQL queries, and finally refines results based on returned outputs. This design enables precise handling of numerical and table-centric queries beyond the capability of semantic retrieval.
Proceeding agent prompt. For graph-structured proceeding data, the prompt encodes a hierarchical retrieval strategy. The action policy first performs proceeding selection via topic matching, and then dynamically chooses among three GraphRAG tools (basic, local, and global) according to query granularity. Specifically, local retrieval is preferred for entity-centric queries, whereas global retrieval is used for higher-level summarization. To improve robustness under incomplete graph construction, the prompt further incorporates a fallback rule that invokes basic retrieval when structured retrieval fails to return sufficient evidence.
Regulation agent prompt. For regulatory documents, we design a minimal high-precision prompt that focuses on retrieval refinement rather than tool selection. The action policy adopts a two-stage retrieve-and-rerank pipeline, where embedding-based retrieval is followed by lightweight reranking with Qwen3-Reranker-0.6B. This design improves evidence precision for regulation-oriented queries, where correctness is critical.
III SpecBench Dataset
We introduce SpecBench, a real-world benchmark dataset for evaluating RAG systems in the spectrum domain. SpecBench addresses the lack of benchmarks for heterogeneous spectrum data sources, including proceedings, license databases, and regulations. SpecBench adopts a question and answer (Q&A) format, a standard paradigm for evaluating both retrieval accuracy and generation faithfulness [15, 10].
Following established RAG evaluation dimensions [16], we consider three core capabilities: (i) noise robustness, i.e., the ability to identify correct evidence under noisy retrieval results; (ii) information integration, i.e., the ability to synthesize evidence across multiple sources; and (iii) negative rejection, i.e., the ability to detect insufficient supporting evidence and abstain from answering, thereby avoiding hallucinated responses.
To unify evaluation across modalities, we define an evidence unit as the minimal retrievable element, corresponding to a document for unstructured data and a table cell for structured data. A question is classified as single-source if it can be answered using one evidence unit, and multi-source otherwise. This distinction enables controlled evaluation of noise robustness and information integration.
| Category | Subtype | Data Source | Evaluated Capability |
| Proceeding | Single-cell | Proceedings | Noise Robustness |
| Multi-cells | Information Integration | ||
| License | Single-cell | Licenses | Noise Robustness |
| Multi-cells | Information Integration | ||
| Regulation | WiLL [15] | FCC Title 47 | Information Integration |
| Compound | Parallel | Multi-source | Information Integration |
| Sequential | |||
| Unanswerable | - | - | Negative Rejection |
As summarized in Table II, SpecBench organizes questions by data source and task structure. Single-source questions primarily evaluate noise robustness, while multi-source and compound questions assess information integration. Compound questions further test cross-agent coordination and are categorized into parallel and sequential types. In parallel questions, required evidence can be retrieved independently from different sources and combined at the final stage. By contrast, sequential questions involve inter-step dependencies, where intermediate results from one retrieval step are required to formulate subsequent queries, making them more susceptible to error propagation. Unanswerable questions contain no valid supporting evidence and evaluate negative rejection.
We construct SpecBench based on question types identified through real-world domain experts interviews, ensuring coverage of practical tasks. For each question, we manually retrieve supporting documents and provide a evidence-grounded reference answer. For regulation questions, we incorporate adapted Q&A samples from the WiLL benchmark [15] to improve coverage of regulation-focused queries. The dataset contains 450 Q&A pairs spanning proceeding (31.1%), license (31.1%), regulation (13.3%), compound (14.4%), and unanswerable (10.0%), balancing realism and annotation quality. The distribution of questions across data sources and task types is illustrated in Fig. 3.
IV Experiments
IV-A Baselines
Web-search RAG. This baseline uses Google Search as the retriever, collecting the top-20 results per query as generation context. It serves as a practical baseline for spectrum question answering using publicly accessible resources.
SpectrumRAG. We include SpectrumRAG [8] as a competitive baseline, which is an iterative RAG framework that leverages LLM-based query rewriting to refine retrieval for spectrum policy question answering. This design improves retrieval quality and downstream generation, making it a strong representative of advanced RAG systems in this domain.
| Model | Method | Question Type | |||||||||||
| Proceeding | License | Regulation | Compound | Unanswerable | Overall | ||||||||
| Win | Success | Win | Success | Win | Success | Win | Success | Win | Success | Win | Success | ||
| Qwen3-8B | Web-search RAG | 2.9 | 70.2 | 0.0 | 15.5 | 29.7 | 88.3 | 4.8 | 9.2 | - | 100 | 5.4 | 37.8 |
| SpectrumRAG | 12.5 | 71.5 | 0.8 | 6.7 | 6.7 | 52.4 | 0.0 | 0.0 | - | 100 | 6.1 | 42.6 | |
| SpecMind | 60.4 | 91.8 | 79.3 | 87.9 | 42.0 | 96.7 | 70.5 | 81.7 | - | 100 | 66.2 | 89.8 | |
| GPT-5.2 | Web-search RAG | 4.7 | 78.3 | 0.0 | 25.3 | 38.3 | 95.8 | 8.3 | 13.3 | - | 100 | 7.9 | 50.5 |
| SpectrumRAG | 18.8 | 78.8 | 1.2 | 8.2 | 8.3 | 60.6 | 0.0 | 0.0 | - | 100 | 8.5 | 56.5 | |
| SpecMind | 72.9 | 100 | 97.6 | 100 | 51.7 | 100 | 86.8 | 96.7 | - | 100 | 81.1 | 99.6 | |
IV-B Experimental Setup
Models. To ensure a fair comparison, we evaluate all methods, including the baselines and SpecMind, using two backbone LLMs of different scales, Qwen3-8B and GPT-5.2. Each method uses the same backbone for both RAG and reasoning, and text-embedding-3-small is used for retrieval across all methods.
Metrics. We report success rate and win rate for each query. The success rate evaluates whether a method produces a valid response under one of the two settings: (i) for standard question types, including license, proceeding, regulation, and compound queries, the method must return a factually grounded answer when sufficient supporting evidence exists; and (ii) for negative-rejection questions, where no valid evidence is available, the method must correctly abstain and avoid hallucination by issuing a rejection. The win rate measures comparative answer quality. We use a strong proprietary LLM, Gemini 3.1-Pro-Preview, as a unified evaluator, which is distinct from the backbone models and fixed across all experiments to ensure consistent comparisons. For each query, the evaluator selects the best answer among all candidates; if all answers are incorrect, no method is assigned a win.
IV-C Main Results
As shown in Table III, SpecMind consistently outperforms the baselines, including Web-search RAG and SpectrumRAG, across question types under both backbone LLMs, achieving substantial win-rate gains, especially on license and compound queries, demonstrating effective structured retrieval and cross-source reasoning.
In terms of success rate, SpecMind maintains strong performance across all categories, achieving near-perfect results with GPT-5.2 and consistently high accuracy with Qwen3-8B. In contrast, the baselines achieve significantly lower success rates, particularly on license and compound tasks, highlighting their limitations in precise retrieval and coordinated reasoning. While Web-search RAG performs relatively well on regulation queries due to strong semantic retrieval, it fails to generalize to structured and multi-source settings.
For unanswerable questions, all methods achieve perfect success rates. This is largely because the task primarily requires abstention rather than evidence synthesis, and modern backbone LLMs already exhibit strong intrinsic capabilities in recognizing insufficient evidence and avoiding unsupported generation. As a result, performance on this category is less sensitive to the retrieval framework and more dependent on the underlying model’s calibration.
IV-D Ablation Studies
| Method | Question Type | |||||||||||
| Proc. | Lic. | Reg. | Comp. | Neg. | All | |||||||
| W | S | W | S | W | S | W | S | W | S | W | S | |
| SpecMind | 72.9 | 100 | 97.6 | 100 | 51.7 | 100 | 86.8 | 96.7 | - | 100 | 81.1 | 99.6 |
| w/o License Agent (RAG) | 70.2 (-2.7) | 98.5 (-1.5) | 9.8 (-87.8) | 17.2 (-82.8) | 48.9 (-2.8) | 97.6 (-2.4) | 66.5 (-20.3) | 80.1 (-16.6) | - | 100(-0.0) | 66.8 (-14.3) | 84.4 (-15.2) |
| w/o Proceeding Agent (RAG) | 45.6 (-27.3) | 89.7 (-10.3) | 94.8 (-2.8) | 98.7 (-1.3) | 49.5 (-2.2) | 97.9 (-2.1) | 63.2 (-23.6) | 79.4 (-17.3) | - | 100(-0.0) | 63.1 (-18.0) | 87.2 (-12.4) |
| w/o Regulation Agent (RAG) | 71.8 (-1.1) | 99.3 (-0.7) | 96.2 (-1.4) | 99.1 (-0.9) | 46.7 (-5.0) | 93.6 (-6.4) | 84.7 (-2.1) | 95.1 (-1.6) | - | 100(-0.0) | 80.3 (-0.8) | 99.0 (-0.6) |
To analyze the source of performance gains, we conduct ablation studies on the three specialized agents under the GPT-5.2 backbone by individually replacing each sub-agent with a naive RAG module (without re-ranking) while preserving the overall framework. Results are reported in Table IV.
We observe consistent performance degradation across all ablations, confirming that each agent contributes to the overall effectiveness. The impact is most pronounced on the corresponding question types, followed by compound queries, reflecting error propagation in multi-step reasoning.
The degradation is particularly severe for the license agent, with substantial drops in both win rate and success rate (up to and ), highlighting the importance of SQL-based retrieval for structured data. The proceeding agent also plays a critical role, with notable declines on proceeding questions ( win, success), primarily due to failures in multi-document reasoning, underscoring the benefit of graph-based retrieval. In contrast, the regulation agent shows relatively smaller impact, as its improvements mainly arise from prompt design and reranking over text-based retrieval, which offer more limited gains than structured and graph-based mechanisms.
V Qualitative Analysis
V-A Illustrative Example
To illustrate how SpecMind operates, we present a compound Q&A example jointly supported by all three specialized agents: Within the 12.7–13.25 GHz band, which frequencies are allocated to satellite services? Is Intelsat an incumbent in this band, and if so, do Intelsat and other incumbents share the same position on spectrum sharing with terrestrial services?
The reasoning process is illustrated in Fig. 4. The supervisor agent operates in a sequential, state-driven manner, iteratively selecting which agent to invoke and what subtask to issue. It first queries the regulation agent to identify the relevant frequency band (e.g., 12.7–13.25 GHz), establishing a constraint for subsequent steps. Based on this result, it invokes the license agent to retrieve corresponding licensees (e.g., INTELSAT, GLOBECOMM) via SQL queries, grounding the task in structured data. Finally, the proceeding agent performs graph-based retrieval to compare these entities’ positions on spectrum sharing and reveal consistent opposition among incumbents. At each step, intermediate results are incorporated into the task state, allowing the supervisor to adaptively refine subsequent actions and resolve inter-step dependencies. This example demonstrates how SpecMind supports coherent cross-source reasoning through dynamic coordination of heterogeneous retrieval mechanisms.
V-B Failure Analysis
To better understand the small set of remaining non-optimal cases, we manually inspect queries where SpecMind does not achieve the best result or produces a partially complete answer. We observe three recurring edge conditions. (i) Fine-grained attribution. In some proceeding graphs, entity–filing relations indicate co-occurrence but do not always explicitly encode whether the entity authored the filing or was referenced by another party, which can make fine-grained attribution more challenging. (ii) Sparse evidence for long-tail entities. Stakeholders that appear only a few times in a proceeding naturally have sparser graph representations, and retrieval for such entities may surface broader contextual information rather than the most specific statement of position. (iii) Ambiguity in structured queries. Some identifiers are only unique within a service-specific subtable, so queries that omit the service type can admit multiple reasonable interpretations and lead to broader aggregation than intended. Overall, these cases largely reflect edge conditions in evidence representation and query specification rather than systematic limitations of the coordination framework. They also point to natural extensions, including richer authorship metadata in graph construction and lightweight clarification mechanisms for ambiguous queries.
VI Conclusions
In this paper, we introduced SpecMind, a multi-agent hybrid RAG framework for heterogeneous spectrum data, together with SpecBench, a 450-question benchmark spanning licenses, regulations, proceedings, and cross-source tasks. The results show that SpecMind is particularly effective when conventional semantic retrieval is poorly matched to the underlying data structure: SQL-based retrieval is critical for license records, graph-based retrieval benefits multi-document proceedings, and their coordination yields substantial gains on compound queries. In contrast, the smaller improvement on regulation questions suggests that conventional RAG remains competitive when the source is already well suited to semantic retrieval. These findings indicate that the value of SpecMind lies in selecting and composing retrieval mechanisms according to the structure of the evidence rather than relying on a single universal retrieval pipeline.
Our analysis further characterizes the remaining non-optimal cases and highlights opportunities for refinement, particularly in fine-grained evidence representation and query disambiguation. Future work can build on these directions while expanding the scale and coverage of spectrum policy datasets to support more comprehensive evaluation across broader regulatory domains, frequency bands, and stakeholder interactions. To facilitate such studies, we publicly release SpecBench, the reusable spectrum databases, and the SpecMind framework, providing a common basis for evaluating retrieval-augmented reasoning over heterogeneous spectrum data[11].
Acknowledgment
The authors would like to thank Zhiyu Shen, Yuxi Chen, Omkar Mujumdar, Yankai Peng, Hasan Nazim Bicer, Christopher Wahl, Solbee Kang, and Lucas Scholler in the Notre Dame Wireless Institute for their valuable contributions to data collection and the evaluation process.
In addition, we thank Caleb Reinking, Le Li Kruczek, Connor Howington and Paul Brenner in the Notre Dame Center for Research Computing for their contributions to code optimization and the implementation of the system user interface. Last but not the least, this project was conceived during a SpectrumX center meeting, and subsequently received a seed fund to carry out the research. We thank NSF and SpectrumX for the support.
References
- [1] (2017) Three-tier shared spectrum, shared infrastructure, and a path to 5G. Cambridge University Press. Cited by: §I.
- [2] (2024) Accelerating regulatory processes with large language models: a spectrum management case study. In 2024 IEEE Pacific Rim Conference on Communications, Computers and Signal Processing (PACRIM), pp. 1–7. Cited by: §I.
- [3] (2023) Understanding telecom language through large language models. In GLOBECOM 2023-2023 IEEE Global Communications Conference, pp. 6542–6547. Cited by: §I.
- [4] (2020) Language models are few-shot learners. arXiv preprint arXiv:2005.14165 1 (3), pp. 3. Cited by: §I.
- [5] (2025) Agentic retrieval-augmented generation: a survey on agentic RAG. arXiv preprint arXiv:2501.09136. Cited by: §I.
- [6] (2023) Retrieval-augmented generation for large language models: a survey. arXiv preprint arXiv:2312.10997 2 (1). Cited by: §I.
- [7] (2020) REALM: retrieval-augmented language model pre-training. External Links: 2002.08909, Link Cited by: §I.
- [8] (2025) Initial evaluation of retrieval-augmented generation approaches in spectrum policy research. In 2025 IEEE International Symposium on Dynamic Spectrum Access Networks (DySPAN), pp. 1–5. Cited by: §I, §IV-A.
- [9] (2024) Telco-RAG: navigating the challenges of retrieval augmented language models for telecommunications. In GLOBECOM 2024-2024 IEEE Global Communications Conference, pp. 2359–2364. Cited by: §I.
- [10] (2025) Retrieval-augmented generation for reliable interpretation of radio regulations. arXiv preprint arXiv:2509.09651. Cited by: §I, §III.
- [11]
Specmind and SpecBench.
Note: https://specmindrag.github.io
Cited by: §I,
§VI,
SpecMind: Enabling Spectrum Intelligence via Multi-Agent Hybrid Retrieval-Augmented Generation
. - [12] (2025) From local to global: a graph RAG approach to query-focused summarization. External Links: 2404.16130, Link Cited by: §II-A.
- [13] (2019) Do NLP models know numbers? probing numeracy in embeddings. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), K. Inui, J. Jiang, V. Ng, and X. Wan (Eds.), Hong Kong, China, pp. 5307–5315. External Links: Link, Document Cited by: §II-A.
- [14] (2023) ReAct: synergizing reasoning and acting in language models. External Links: 2210.03629, Link Cited by: §II-C.
- [15] (2025) Can we make FCC experts out of LLMs?. In Proceedings of the 26th International Workshop on Mobile Computing Systems and Applications, pp. 85–90. Cited by: TABLE II, §III, §III.
- [16] (2024) Benchmarking large language models in retrieval-augmented generation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, pp. 17754–17762. Cited by: §III.