1]University of British Columbia
2]Vector Institute
3]Harvard University
4]Rice University
5]University of Chicago
6]BC Cancer Agency
\contribution∗ Equal contribution
† Correspondence to Xiaoxiao Li at xiaoxiao.li@ece.ubc.ca
SlideBank: A Persistent Hierarchical Evidence Bank for
Consistent Whole-Slide Reasoning
Abstract
Whole-slide images (WSIs) are challenging for vision-language reasoning because diagnostically relevant morphology is sparse, heterogeneous, and distributed across gigapixel-scale images and multiple spatial resolutions. Existing WSI models and pathology agents can aggregate slide features or actively acquire evidence, but the information retained after exploration is often difficult to access semantically while preserving its connection to the original visual evidence. We introduce SlideBank, a training-free framework that represents each WSI as a persistent, concept-indexed, and spatially grounded evidence bank. SlideBank performs question-independent coarse-to-fine exploration to identify informative regions and multi-scale views, converts them into explicit morphological observations, and grounds pathology signals to their supporting patches and WSI coordinates. At inference time, questions are routed to relevant signals and evidence scales, and the linked global, regional, and patch evidence is integrated through confidence-based cross-level consensus. Experiments on WSI-VQA and SlideBench-BCNB show that with Patho-R1, SlideBank reaches 52.77% on WSI-VQA and with Quilt-LLaVA, it reaches 50.92% average accuracy on SlideBench-BCNB, while structured signal-guided retrieval consistently outperforms random evidence sampling. Reusing the same bank across repeated queries further achieves over 99% rephrasing consistency and substantially reduces amortized inference cost through persistent evidence reuse.
1 Introduction
Whole-slide images (WSIs) are among the most challenging visual inputs in computational pathology. Unlike conventional natural or medical images, a single WSI can contain billions of pixels and span tissue structures ranging from large-scale architectural patterns to fine-grained cellular morphology. Diagnostically relevant findings are often sparse, spatially dispersed, and highly heterogeneous: a small focus of atypical cells, an invasive tissue interface, or a mitotic hotspot may occupy only a tiny fraction of the entire slide while being decisive for diagnosis. Consequently, directly compressing a WSI into a thumbnail or uniformly processing a limited set of patches can easily discard critical evidence. Effective WSI reasoning therefore requires not only understanding what is visible, but also determining where diagnostically informative evidence lies, at what scale it should be examined, and how evidence distributed across the slide should be organized for downstream reasoning.
Recent work has approached this challenge from two major directions. WSI-level vision-language models Chen et al. (2025b); Liang et al. (2025); Ding et al. (2025); Alawode et al. (2026) aggregate precomputed patch or slide features into compact latent representations, enabling end-to-end reasoning over gigapixel images. More recently, pathology agents Ghezloo et al. (2025); Chen et al. (2025a); Yang et al. (2026) have treated WSI understanding as an active evidence-acquisition problem, using planning, region selection, and multi-scale navigation to locate diagnostically relevant fields before generating an answer. Several methods further introduce hierarchical representations, cached observations, or slide-specific memory to improve efficiency and reasoning over large visual contexts Alawode et al. (2026); Chen et al. (2025a); Yang et al. (2026); Huang et al. (2025); Li et al. (2026). These approaches have substantially advanced WSI understanding, but the information retained after slide exploration is typically represented either as latent visual features, flat collections of observations, or states optimized for a particular navigation or reasoning trajectory. How to organize visually established evidence from the current slide into an explicit representation that can be semantically accessed by different downstream questions remains comparatively underexplored.
This motivates the central question of this work: how can evidence acquired from a gigapixel WSI be transformed into a persistent representation that supports diverse downstream questions while preserving explicit traceability to the underlying visual observations? This perspective shifts the focus from evidence acquisition alone to how acquired evidence is organized and accessed after exploration. A useful evidence representation should expose clinically meaningful findings in a form that can bridge different question formulations, preserve the multi-scale structure of pathology evidence, and maintain a traceable connection between semantic findings and the fine-grained areas and WSI locations that support them.
Constructing such a representation introduces several challenges. First, exhaustive inspection is infeasible at gigapixel scale, yet aggressive sampling risks missing sparse or heterogeneous abnormalities; the system must therefore identify a compact but diagnostically diverse set of fields across both architectural and cellular scales. Second, the language used in downstream questions does not necessarily match the language of raw visual observations. Questions about nuclear grade, pleomorphism, or cellular atypia, for example, may rely on overlapping morphological signals even when their surface forms differ, making direct retrieval over free-form descriptions unreliable. Third, clinically meaningful evidence is inherently multi-scale and spatially grounded: a semantic finding should remain linked not only to a textual description, but also to the anchor region, magnification, visual field, and WSI coordinates from which it was derived. Finally, different questions may require different combinations of global context, regional architecture, and cellular detail, requiring an evidence interface that remains persistent while supporting query-adaptive composition.
To address these challenges, we propose SlideBank, an agent-driven framework that requires no task-specific training and converts each WSI into a concept-indexed, spatially grounded hierarchical evidence bank. To construct the bank, SlideBank performs question-independent coarse-to-fine exploration, using a global survey and category-aware sampling to identify diagnostically informative anchor regions and complementary architecture- and cell-level views. A pathology generative model converts the selected views into explicit morphological observations, which are subsequently normalized into clinically meaningful signals by a concept-guided pathology signal layer. Each signal is associated with its status, confidence, supporting anchors, multi-scale views, coordinates, and magnifications, establishing a traceable path from pathology signal to supporting evidence, visual field, and ultimately WSI location. At inference time, a signal-aware reader routes each question to a relevant pathology concept, uses the associated signals to retrieve supporting anchors and multi-scale views, and generates candidates independently at the global, anchor, and patch levels. These candidates are combined through confidence-based selection and cross-level consensus. In this way, SlideBank explores each slide once, organizes the observed morphology into a persistent evidence bank, and reuses the same bank to answer different downstream questions.
Our contributions are three-fold:
- •
We introduce SlideBank, a structured, slide-specific hierarchical evidence representation that organizes observed morphology through pathology signals while preserving explicit links to supporting regions, multi-scale views, and WSI coordinates.
- •
We develop a concept-to-signal retrieval method that couples concept-level question routing with signal-grounded evidence retrieval, providing a consistent semantic interface for indexing and reasoning across diverse question formulations while retaining traceability to the underlying visual evidence.
- •
We evaluate SlideBank on two public WSI question-answering benchmarks in terms of accuracy, efficiency, and cross-turn consistency. A controlled comparison using the same evidence bank shows the benefit of concept-to-signal retrieval over random evidence sampling.
2 Related Work
2.1 Vision–Language Models and Agentic Reasoning for Whole-Slide Pathology
Early pathology vision–language models mainly learned image–text alignment or instruction following from localized fields. CONCH learns transferable visual–language representations, while Quilt-LLaVA and PathChat support instruction following and open-ended interaction with histology images Lu et al. (2024a); Seyfioglu et al. (2024); Lu et al. (2024b). Because these models cannot directly process a complete WSI at native resolution, subsequent methods aggregate pre-extracted patch features into slide-level representations. WSI-VQA introduced generative question answering over WSIs, TITAN learned a multimodal whole-slide representation, and SlideChat and WSI-LLaVA connected slide features to language models for slide-level dialogue and reasoning Chen et al. (2024); Ding et al. (2025); Chen et al. (2025b); Liang et al. (2025). More recent systems improve hierarchical or large-scale slide understanding: MLLM-HWSI aligns cell-, patch-, region-, and slide-level tokens with pathology language, while ALPaCA and PRISM2 scale slide-level supervision through question answering and clinical dialogue Alawode et al. (2026); Gao et al. (2026); Vorontsov et al. (2026). In parallel, agentic systems make evidence acquisition more adaptive. PathFinder, WSI-Agents, SlideSeek, GIANT, and PathAgent introduce specialized agents, planning, tool use, iterative navigation, or evidence synthesis for gigapixel slides Ghezloo et al. (2025); Lyu et al. (2025); Weishaupt et al. (2025); Buckley et al. (2025); Chen et al. (2025a); HistoSelect performs question-guided coarse-to-fine selection, BEACON acquires patches according to expected information gain, and AdaptivePath learns question-agnostic multiscale navigation Huang et al. (2026); Wong et al. (2026); Chen et al. (2026). EviPathBench further shows that locating diagnostic regions remains substantially harder than reasoning over preselected evidence Liao et al. (2026). PathNavigate is particularly related to our setting because it performs a question-independent scan and maintains an online memory over frozen slide features before question-conditioned search Yang et al. (2026). However, existing methods primarily represent acquired evidence as model-internal features or transient reasoning states. SlideBank instead materializes selected multiscale image fields, morphological descriptions, pathology signals, and WSI coordinates as an independently inspectable bank that can be retrieved across multiple questions about the same slide.
2.2 Hierarchical Evidence and Memory
Hierarchical modeling provides a natural way to represent gigapixel WSIs. HIPT learns nested patch- and region-level representations, while H2-MIL organizes heterogeneous instances into a hierarchy for slide analysis Chen et al. (2022); Hou et al. (2022); however, both primarily structure latent visual features for downstream prediction. Concept bottleneck models expose human-interpretable intermediate variables, and post-hoc variants add such structure to pretrained models Koh et al. (2020); Yuksekgonul et al. (2023). Persistent memory has also been studied in long-horizon agents: Generative Agents store episodic experiences, MemGPT manages context through external memory, and A-MEM and G-Memory organize memories through linked or hierarchical structures Park et al. (2023); Packer et al. (2023); Xu et al. (2025); Zhang et al. (2025). Related ideas have recently appeared in pathology. SurvAgent constructs a multimodal case bank for cross-case survival prediction, whereas PathMem organizes structured diagnostic knowledge as long-term memory and transfers relevant content into working memory Huang et al. (2025); Li et al. (2026). These approaches focus on latent hierarchies, general episodic memory, cross-case experience, or domain knowledge rather than persistent spatial evidence from repeated examination of one WSI. SlideBank combines hierarchical visual evidence with an interpretable routing interface: concepts represent coarse question-side categories, signals represent localized image-derived findings, and a predefined ontology maps concepts to relevant signals. Each signal record retains its status and links to associated anchors, multiscale patch views, descriptions, and WSI coordinates.The resulting slide-specific bank is reused across questions, so every turn is answered from the persistent evidence bank, not only affected by the accumulated conversation history.
3 Method
3.1 Overview
Whole-slide reasoning is challenging since a gigapixel WSI exceeds the accepted maximum image size of current vision-language models (VLMs) and clinically relevant findings may occupy only a small fraction of the slide. SlideBank therefore separates slide exploration from question answering. As illustrated in Fig. 1, the framework consists of three components: Agent-Driven Hierarchical Evidence Bank Construction organizes multi-scale slide observations into a persistent, spatially grounded evidence bank; Hierarchical Evidence-Augmented WSI Reasoning retrieves and integrates question-relevant evidence from the bank; and Persistent Evidence Reuse Across Queries enables the same slide representation to support subsequent questions without reconstructing the underlying evidence.
Let a whole slide image be represented by a resolution pyramid , where and denote the highest- and lowest-resolution views, respectively. In Section 3.2, SlideBank performs question-independent coarse-to-fine exploration to construct a slide-specific evidence bank that organizes multi-scale observations and links pathology signals to their supporting evidence and WSI locations. In Section 3.3, the online reader maps a question to relevant signals and evidence scales, retrieves the linked evidence, and combines global, regional, and patch-level predictions through confidence-based consensus. Finally, Section 3.4 reuses the same across subsequent queries.
3.2 Agent-Driven Hierarchical Evidence Bank Construction
Global Survey and Coordinate Normalization. To capture slide-level contextual cues and facilitate global morphological measurements, such as tumor size and disease extent, we first construct a global thumbnail. For pyramidal WSIs, we use the lowest-resolution level; for non-pyramidal or single-resolution slides, we construct an equivalent bounded thumbnail via chunk-wise subsampling. All spatial coordinates identified on the thumbnail are proportionally mapped back to the level-0 coordinate system of the source WSI. Foreground tissue is separated from the glass background using HSV thresholding, and the valid tissue area is partitioned into a non-overlapping grid. Each tile is characterized using handcrafted morphological statistics and foreground occupancy, and assigned to one of five coarse tissue classes: tumor, necrosis, adipose, stroma, or other. When physical pixel spacing is available, we additionally estimate the maximum Feret diameter and principal axis of the primary tumor candidate. These slide-level observations, including the natural-language summary, tissue-composition statistics, and geometric diagnostics, are stored in a global record for slide .
Coarse-to-fine Agentic Sampling. Tiles are first ranked according to coarse diagnostic priority, triage confidence, and tissue coverage, and retained under a category-aware budget that allocates more samples to suspicious regions while preserving coverage of the remaining tissue classes. Sampling then proceeds through recursive refinement. For each selected tile, the agent scores all 16 cells of its downsampled representation to propose anchor regions. Within each anchor region, the same procedure is repeated to identify informative and patches. At both stages, scoring is guided by a fixed set of pathology-derived cues provided in the prompt (prompts used throughout SlideBank are provided in Appendix).
Multi-Scale Visual–Textual Observations. The resulting hierarchy separates coarse triage from fine-grained evidence: tiles provide the coarse sampling prior, whereas anchor regions and their associated patches provide the observed morphological evidence. For each patch , a pathology-specific VLM generates a concise morphological description , together with a self-reported confidence score and a quality-control flag, conditioned on the patch image and its magnification. Descriptions with confidence below a predefined threshold or invalid quality-control flags are discarded.
For each retained anchor region , the descriptions of its multi-magnification patches are collected as
| (1) |
where denotes the patches associated with anchor . This anchor-level textual observation summarizes complementary morphology across magnifications while retaining links to the originating patches and their WSI coordinates. We write for the set of anchor regions retained for slide , and for all multi-scale patch views collected under them; each stores its image, magnification, level-0 coordinates, and description .
Pathology Signal Grounding. Free-form descriptions may express the same finding in different ways. We therefore map to a predefined vocabulary of localized pathology signals, such as marked nuclear pleomorphism, high mitotic activity, low tubule formation, and tumor necrosis. For each applicable signal , we store
| (2) |
where is the inferred status, is its confidence, and contains the patch views supporting a positive assessment. This standardized signal layer preserves links to the underlying descriptions, images, and WSI locations.
Persistent Hierarchical Storage. The signal records across all anchors are collected as
| (3) |
where is the set of signals applicable to anchor . We also define an ontology that maps each coarse pathology concept to a set of relevant signals . The slide bank is then represented as
| (4) |
comprising the global record, the retained anchor regions, their multi-scale patch views, the grounded signal records, and the concept-to-signal ontology. The complete pathology concept vocabulary, concept-to-signal mapping and evidence-bank schema are provided in Appendix.
3.3 Hierarchical Evidence-Augmented WSI Reasoning
Question-to-Concept Routing and Signal Retrieval. A pathology concept represents the coarse intent of a question, such as tumor grading, necrosis assessment, or margin status, whereas a pathology signal represents a localized image-derived finding. Given a question and its options , a deterministic router returns the top- pathology concepts
| (5) |
is based on lexical rules over the pathology concept vocabulary (details in Appendix). We use by default, sensitivity to n is analyzed in Sec. 4.3. The target signals are then used to retrieve
| (6) |
For each target signal, SlideBank selects the anchor regions in which that signal is most strongly supported. Records labeled as present are preferred over uncertain or absent records, and records with the same status are ranked by confidence. The associated anchor images, patch views, and morphological descriptions are then retrieved as the question-specific evidence .
Evidence-Conditioned Branch Predictions. The question-specific evidence is organized into three complementary branches . The global branch captures slide-wide context, the anchor branch represents regional architecture, and the patch branch provides localized morphology. Anchor and patch inputs include their stored morphological descriptions, whereas the global branch uses the slide thumbnail without a local description.
Each evidence item is evaluated independently by the inference VLM. Let denote the answer-option set and the logit assigned to option for item from branch . Its option probability is
| (7) |
The corresponding prediction and confidence are
| (8) |
Within the anchor and patch branches, candidates are ranked by , and the top- candidates are retained as .
Level-Weighted Answer CSonsensus. For each local branch , the retained candidates are aggregated by majority voting. The resulting support for option is
| (9) |
For the global branch, the option probabilities are used directly:
| (10) |
The three branches are combined using
| (11) |
and the final answer is
| (12) |
We use , , and , placing greater emphasis on localized morphological evidence while retaining regional and global context. The retained candidates, signal records, and spatial evidence links provide a trace from the final answer to its supporting WSI regions.
3.4 Persistent Evidence Reuse Across Queries
SlideBank constructs the evidence bank once and reuses it across all turns. For each question , the reader performs question-dependent routing to obtain evidence from the bank and predicts the answer with the evidence, avoiding repeated WSI exploration and reducing evidence drift. Routing depends only on the current question, so the retrieved evidence for a given question is identical whether it is asked in isolation or within a conversation.
To preserve conversational continuity without accumulating a long history, the reader carries only the previous question–answer pair,
| (13) |
which is supplied to the inference VLM alongside the retrieved evidence. The visual evidence therefore always comes from the persistent bank, while provides bounded dialogue context.
4 Experiments
Datasets. We evaluate SlideBank on two public WSI-level benchmarks: WSI-VQA Chen et al. (2024), which contains 85 WSIs with slide-level visual question answering annotations, and SlideBench-BCNB Chen et al. (2025b), which contains 1,058 breast cancer slides covering tumor typing, grading, and molecular subtyping. The two benchmarks provide complementary settings for evaluating general WSI question answering and clinically relevant diagnostic reasoning. For repeated-query evaluation, we additionally construct a multi-turn setting from WSI-VQA by grouping all questions associated with the same slide into one conversation, resulting in an average of 29 questions per WSI. We further generate five semantically equivalent formulations of each question with shuffled answer options using Gemini-2.5-Flash Comanici et al. (2025) to evaluate the stability of predictions under different question formulations.
Baselines. We compare against three families of methods. Zero-shot VLMs include general-purpose and pathology-specific models that receive only a WSI thumbnail. WSI-trained models include WSI-LLaVA, TITAN, and SlideChat, which are trained or adapted for slide-level understanding. Agentic systems include Med-Agents and PathAgent, which introduce additional inference-time reasoning or WSI exploration. All methods are evaluated on the same question sets and answer spaces. Thumbnail-based baselines receive no additional high-resolution patches. We evaluate SlideBank with Quilt-LLaVA and Patho-R1 to examine backbone transferability.
Metrics. For standard WSI QA, we report accuracy on WSI-VQA and question-count-weighted average over the three BCNB tasks (tumor, histological grading and molecular subtype). For repeated-query evaluation, we additionally report accuracy, the change from single-query accuracy (Acc), and rephrase consistency () across semantically equivalent question formulations.
Implementation Details. Within each experimental setting, we use the same pathology VLM backbone for both evidence-bank construction and online inference, without additional training. We show the result with Quilt-LLaVA Seyfioglu et al. (2024) and Patho-R1 Zhang et al. (2026) as the backbone. For evidence-bank construction, we partition the slide thumbnail into tiles, retain those with at least tissue coverage, and perform both tile-level anchor localization and within-anchor patch selection on a grid. We select three anchors from each tumor-like tile and two from each retained non-tumor tile. For each anchor, we extract an adaptive context crop capped at 4096 level-0 pixels, a 512-pixel fine-detail view corresponding to approximately on slides scanned at , and an additional architectural view at approximately . All views are resized to for VLM processing. During online inference, we use all retrieved anchor-level evidence and sample at most 32 patch-level evidence items with a fixed random seed 42. We retain the top five (=5) candidates from the anchor and patch branches according to token-probability confidence and fuse the global, anchor, and patch predictions using . All experiments use a single NVIDIA A100 GPU with 80 GB memory, with FP16 inference for Quilt-LLaVA and BF16 inference for Patho-R1. Additional dataset, baseline, multi-turn evaluation, and answer-parsing details are provided in Appendix.
| Method | Inference Input | SlideBench-BCNB Accuracy (%) | WSI-VQA (%) | |||
| Tumor Type | Grading | Subtype | Weighted Avg. | Accuracy | ||
| Vision-Language Models (Zero-shot) | ||||||
| Qwen2.5-VL-7B Bai et al. (2025b) | Thumbnail | 58.13 | 27.11 | 16.07 | 34.06 | 33.43 |
| Qwen3-VL-8B Bai et al. (2025a) | Thumbnail | 55.01 | 27.65 | 16.92 | 33.43 | 39.92 |
| Qwen3.5-4B Qwen Team (2026) | Thumbnail | 39.89 | 38.34 | 17.30 | 31.56 | 44.53 |
| LLaVA-Med Li et al. (2023) | Thumbnail | 26.56 | 27.75 | 22.40 | 25.48 | 31.87 |
| Quilt-LLaVA Seyfioglu et al. (2024) | Thumbnail | 44.33 | 32.18 | 25.05 | 33.93 | 27.65 |
| PathGen-LLaVA Sun et al. (2025) | Thumbnail | 45.18 | 33.05 | 22.68 | 33.66 | 39.90 |
| Patho-R1 Zhang et al. (2026) | Thumbnail | 61.11 | 29.93 | 27.57 | 39.95 | 44.28 |
| WSI-trained Models | ||||||
| WSI-LLaVA-7B Liang et al. (2025) | Slide | 42.53 | 30.02 | 32.42 | 35.21 | 49.61 |
| TITAN Ding et al. (2025) | Slide | 83.55 | 27.00 | 21.27 | 44.67 | 43.12 |
| SlideChat Chen et al. (2025b) | Slide | 87.99 | 17.86 | 20.41 | 43.14 | 52.05 |
| Agentic Systems | ||||||
| MedAgents Tang et al. (2024) | Thumbnail | 89.38 | 37.37 | 15.12 | 47.72 | 42.75 |
| PathAgent Chen et al. (2025a) | Slide | 50.76 | 46.00 | 27.08 | 41.08 | 52.05 |
| SlideBank (w/ Quilt-LLaVA) | Thumbnail + Evidence | 89.82 | 33.56 | 27.22 | 50.92 | 43.75 |
| SlideBank (w/ Patho-R1) | Thumbnail + Evidence | 83.27 | 40.06 | 26.47 | 50.36 | 52.77 |
4.1 Standard WSI Question Answering
Table 1 evaluates whether the question-independent evidence bank preserves sufficient slide information for downstream WSI reasoning. Despite requiring no task-specific training, SlideBank achieves strong performance on both benchmarks. On SlideBench-BCNB, SlideBank with Quilt-LLaVA obtains 89.82% accuracy on tumor typing and the highest overall average of 50.92%, substantially improving over the thumbnail-only Quilt-LLaVA backbone. On WSI-VQA, SlideBank with Patho-R1 reaches 52.77%, outperforming its thumbnail-only counterpart by 8.49 percentage points and achieving the best performance among the evaluated methods. These results show that organizing gigapixel WSI observations into a reusable structured evidence bank preserves and can improve standard single-query reasoning performance.
4.2 Effect of Evidence Organization and Access
Multi-scale Hierarchy. We first examine how different levels of the evidence hierarchy contribute to WSI reasoning. As shown in Table 2, local evidence is substantially more informative than the global thumbnail alone. On WSI-VQA, the anchor and patch branches achieve 51.98% and 51.45% accuracy, respectively, compared with 48.55% using only global evidence. A similar trend is observed on SlideBench-BCNB, where anchor-level evidence provides the strongest single-level performance at 50.23%. Combining global, anchor, and patch evidence shows the best overall results on both benchmarks, reaching 52.77% on WSI-VQA and 50.36% on SlideBench-BCNB.
| Variant | Evidence | WSI-VQA | BCNB |
| Evidence hierarchy | |||
| Global only | G | 48.55 | 40.57 |
| Anchor only | A | 51.98 | 50.23 |
| Patch only | P | 51.45 | 48.06 |
| Full hierarchy | G+A+P | 52.77 | 50.36 |
| Evidence access | |||
| Random | G+A+P | 50.13 | 49.62 |
| Signal-guided | G+A+P | 52.77 | 50.36 |
Signal-guided Access. We next show whether the gain arises from structured evidence access rather than simply exposing the model to more evidence. Using the same constructed evidence bank, we replace signal-guided retrieval with uniformly sampling 32 patches at random while keeping other settings unchanged. Signal-guided retrieval improves accuracy from 50.13% to 52.77% on WSI-VQA and from 49.62% to 50.36% on SlideBench-BCNB. The improvement shows the benefit of organizing and accessing slide evidence through pathology signals compared to random sampling.
4.3 Sensitivity and Qualitative Example
Sensitivity of the Number of Concepts.
We study the sensitivity to the number of pathology concepts used for concept-routed retrieval. As shown in Fig. 2, accuracy increases from 49.9% with a single concept to a peak of 52.8% with three concepts, and then gradually decreases as additional concepts are included. This reflects a trade-off between evidence coverage and retrieval noise: too few concepts may omit relevant pathology signals, whereas overly broad routing introduces less informative evidence. Performance remains above random sampling across a broad range of 2–6 concepts, indicating that the method is not sensitive to a narrowly tuned setting. We therefore use three concepts by default.
Qualitative Example. Figure 3 illustrates the complete evidence trace for a representative grading question. The query is first routed to grade-related signals, which retrieve their linked regions and multi-magnification patches. This example demonstrates how SlideBank preserves an explicit path from the final prediction back to the pathology signals and their supporting visual evidence.
4.4 Persistent Evidence Reuse Across Queries
Rephrased-QSuery Stability. We evaluate whether a persistent slide-specific evidence bank provides a stable basis for answering repeated questions about the same WSI. We group all WSI-VQA questions associated with each slide into a single sequence, leading to 29 questions per WSI on average, and present each question in a rephrased form with shuffled options. As shown in Table 3, several baselines suffer substantial accuracy degradation compared with independent single-query inference. In contrast, SlideBank shows only modest drops of 1.38 and 3.83 percentage points with Quilt-LLaVA and Patho-R1, respectively, while achieving 99.47% and 99.21% consistency across semantically equivalent question formulations. Since routing depends only on the current question, these small drops arise from question rephrasing and the bounded dialogue context rather than from re-exploring the slide like baselines.
| Method | Acc | Acc | Rephrase Cons. |
| LLaVA-Med Li et al. (2023) | 16.18 | -15.69 | 62.83 |
| Quilt-LLaVA Seyfioglu et al. (2024) | 25.22 | -2.43 | 84.25 |
| PathGen-LLaVA Sun et al. (2025) | 23.06 | -16.84 | 87.27 |
| Patho-R1 Zhang et al. (2026) | 41.93 | -2.35 | 80.94 |
| WSI-LLaVA Liang et al. (2025) | 17.06 | -32.55 | 73.21 |
| SlideChat Chen et al. (2025b) | 42.50 | -9.55 | 88.69 |
| SlideBank (w/ Quilt-LLaVA) | 42.37 | -1.38 | 99.47 |
| SlideBank (w/ Patho-R1) | 48.94 | -3.83 | 99.21 |
Amortized Efficiency. Figure 4 evaluates the computational benefit of reusing the same evidence bank across repeated queries. Although bank construction introduces a one-time cost, this cost is progressively amortized as more questions are asked about the same slide. The average runtime decreases from 94.3 s for a single query to 5.9 s per query over the question sequence, approaching an online reasoning cost of 2.73 s per query. We restrict the comparison to generative VLMs and WSI assistants with comparable conversational inference interfaces; representation-oriented models such as TITAN and agentic methods whose multi-round computation occurs internally within a single query are excluded. As shown in Fig. 4(b), SlideBank has higher consistency and lower runtime than baselines.
5 Conclusion
We introduced SlideBank, a training-free framework that converts each WSI into a persistent, concept-indexed, and spatially grounded evidence bank. SlideBank separates slide exploration from question answering, organizes multi-scale morphology through pathology signals, and retrieves question-relevant evidence with explicit links to supporting regions and patches. Experiments on WSI-VQA and SlideBench-BCNB show competitive performance, improved stability across repeated queries, and efficiency gains from evidence reuse. These results highlight structured evidence organization as a promising direction for WSI reasoning.
References
- MLLM-hwsi: a multimodal large language model for hierarchical whole slide image understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13732–13743. Cited by: §1, §2.1.
- Qwen3-vl technical report. arXiv preprint arXiv:2511.21631. Cited by: §C.2, Table 1.
- Qwen2.5-vl technical report. External Links: 2502.13923, Link Cited by: §C.2, Table 1.
- Navigating gigapixel pathology images with large multimodal models. arXiv preprint arXiv:2511.19652. Cited by: §2.1.
- Pathagent: toward interpretable analysis of whole-slide pathology images via large language model-based agentic reasoning. arXiv preprint arXiv:2511.17052. Cited by: §C.2, §1, §2.1, Table 1.
- Agentic visual reasoning in whole-slide pathology images via active perception. arXiv preprint arXiv:2608.08648. External Links: Document Cited by: §2.1.
- Wsi-vqa: interpreting whole slide images by generative visual question answering. In European Conference on Computer Vision, pp. 401–417. Cited by: §2.1, §4.
- Scaling vision transformers to gigapixel images via hierarchical self-supervised learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 16144–16155. Cited by: §2.2.
- Slidechat: a large vision-language assistant for whole-slide pathology image understanding. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 5134–5143. Cited by: §C.2, §1, §2.1, Table 1, Table 3, §4.
- Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261. Cited by: §4.
- A multimodal whole-slide foundation model for pathology. Nature medicine, pp. 1–13. Cited by: §C.2, §1, §2.1, Table 1.
- ALPaCA: adapting llama for pathology context analysis to enable slide-level question answering. Nature Communications. External Links: Document Cited by: §2.1.
- PathFinder: a multi-modal multi-agent system for medical diagnostic decision-making applied to histopathology. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 23431–23441. Cited by: §1, §2.1.
- Hˆ 2-mil: exploring hierarchical representation with heterogeneous multiple instance learning for whole slide image analysis. In Proceedings of the AAAI conference on artificial intelligence, Vol. 36, pp. 933–941. Cited by: §2.2.
- SurvAgent: hierarchical cot-enhanced case banking and dichotomy-based multi-agent system for multimodal survival prediction. arXiv preprint arXiv:2511.16635. Cited by: §1, §2.2.
- Act like a pathologist: tissue-aware whole slide image reasoning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6972–6981. Cited by: §2.1.
- Concept bottleneck models. In Proceedings of the 37th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 119, pp. 5338–5348. Cited by: §2.2.
- Llava-med: training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems 36, pp. 28541–28564. Cited by: §C.2, Table 1, Table 3.
- PathMem: toward cognition-aligned memory transformation for pathology mllms. arXiv preprint arXiv:2603.09943. Cited by: §1, §2.2.
- WSI-llava: a multimodal large language model for whole slide image. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 22718–22727. Cited by: §C.2, §1, §2.1, Table 1, Table 3.
- EviPathBench: benchmarking evidence acquisition and reasoning in vision-language models for whole-slide pathology. arXiv preprint arXiv:2607.19261. External Links: Document Cited by: §2.1.
- A visual-language foundation model for computational pathology. Nature medicine 30 (3), pp. 863–874. Cited by: §2.1.
- A multimodal generative ai copilot for human pathology. Nature 634, pp. 466–473. External Links: Document Cited by: §2.1.
- Wsi-agents: a collaborative multi-agent system for multi-modal whole slide image analysis. arXiv preprint arXiv:2507.14680. Cited by: §2.1.
- Memgpt: towards llms as operating systems. arXiv preprint arXiv:2310.08560. Cited by: §2.2.
- Generative agents: interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology, pp. 1–22. Cited by: §2.2.
- Qwen3.5: towards native multimodal agents. External Links: Link Cited by: §C.2, Table 1.
- Quilt-llava: visual instruction tuning by extracting localized narratives from open-source histopathology videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13183–13192. Cited by: §C.2, §2.1, Table 1, Table 3, §4.
- Pathgen-1.6 m: 1.6 million pathology image-text pairs generation through multi-agent collaboration. In International Conference on Learning Representations, Vol. 2025, pp. 94611–94653. Cited by: §C.2, Table 1, Table 3.
- Medagents: large language models as collaborators for zero-shot medical reasoning. In Findings of the Association for Computational Linguistics: ACL 2024, pp. 599–621. Cited by: §C.2, Table 1.
- End-to-end multimodal pathology foundation model with clinical dialogue. Nature Medicine. External Links: Document Cited by: §2.1.
- Evidence-based diagnostic reasoning with multi-agent copilot for human pathology. arXiv preprint arXiv:2506.20964. Cited by: §2.1.
- Beyond relevance: bayesian evidence acquisition for agentic whole-slide image reasoning. arXiv preprint arXiv:2608.05757. External Links: Document Cited by: §2.1.
- A-mem: agentic memory for llm agents. arXiv preprint arXiv:2502.12110. Cited by: §2.2.
- PathNavigate: a training-free pathology agent with surprise-guided scan and shared slide memory for whole-slide image vqa. arXiv preprint arXiv:2605.23559. Cited by: §1, §2.1.
- Post-hoc concept bottleneck models. In International Conference on Learning Representations, Cited by: §2.2.
- G-memory: tracing hierarchical memory for multi-agent systems. arXiv preprint arXiv:2506.07398. Cited by: §2.2.
- Patho-r1: a multimodal reinforcement learning-based pathology expert reasoner. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, pp. 28418–28426. Cited by: §C.2, Table 1, Table 3, §4.
Appendix A Prompts
A.1 Hierachical Evidence Bank Construction Prompts
This section lists the complete prompts used by the offline bank-construction agent (Sec. 3.2). All prompts are issued to the same frozen pathology VLM backbone that is used for online reasoning. Decoding is greedy with at most 128 new tokens. The VLM infers three times during the construction: anchor localization inside a retained tile (P1), view selection inside an anchor context (P2 and P3), and multi-scale view description (P4).
P1: anchor localization within a tile.
The tile is resized to and partitioned into a grid. The agent scores all 16 cells and returns the top {target_count} cells, which is 3 for tumor-like tiles and 2 otherwise. P1 shows the anchor localization prompt.
P2: fine-detail view selection within an anchor.
The anchor region is resized to and again partitioned into a grid. The agent first scores all 16 cells with a quality-control label and morphology tags, and then commits to {target_count} cells that are extracted as 512-pixel level-0 crops. The tag vocabulary is the image-side entry point of the signal ontology. We show the example of the breast used for WSI-VQA and BCNB in P2. The corresponding tag list can be replaced by the ontology of the target cohort for other organs.
P3: architectural context view selection.
Each anchor additionally receives one independently selected low-magnification view (). Selection is deliberately decoupled from P2 so that the context view is not forced to coincide with the cell-level evidence, and the prompt asks for architectural cues such as transition fronts, necrosis interfaces, and ductal layout.
P4: multi-scale view description.
Every extracted view is described independently, producing the triple of Sec. 3.2. The prompt constrains the model to observable morphology and to a fixed JSON schema, and it states the specimen organ {organ} of the cohort under evaluation (e.g., breast for WSI-VQA and BCNB) to suppress the organ drift that pathology VLMs exhibit on isolated fields. Diagnostic conclusions, grading terms, and management statements are excluded at the prompt level and again removed by the post-processing rules described below, so that the bank stores observations.
A.2 Inference Prompts
Size-question specialization.
Tumor extent is the one question family that is answered from measurements stored in the global record. A short prompt (P5) decides whether the current question is a measurement question, and if it is, the question is answered from the thumbnail with the scale prior of the stored pyramid geometry (P6). Questions that are not measurement questions are ignored by this branch and follow the ordinary routing path.
Branch answering.
Every retrieved item is evaluated independently by the same VLM with the level-specific prompt of P7 (a-c). The three variants differ only in which memory content accompanies the image: the global branch sees the slide thumbnail alone, the anchor branch additionally receives the stored anchor description, and the patch branch receives the stored view caption. The final answer is from the level-weighted consensus.
Multi-turn context.
For multi-turn evaluation, the bounded history is rendered as a plain block that is prepended to the current question before the branch prompts are built, so that each branch is conditioned on the same discourse context while the visual evidence is retrieved afresh from the persistent bank. The history contains predicted answers, not ground truth, and is truncated to the most recent turns for the same slide.
Appendix B Pathology Concept and Signal Ontology
SlideBank uses a compact pathology ontology to connect a clinical question with relevant visual evidence. A concept describes the clinical intent of a question (e.g., tumor grade or lymphovascular invasion), whereas a signal describes an atomic morphologic finding that can be assessed in a local image region (e.g., high mitotic activity). The resulting reasoning pipeline is
where a question is first assigned to a concept . The concept selects a set of relevant signals , which are then used to retrieve multi-scale evidence for answer generation. This separation makes the retrieval process interpretable: each answer can be traced from the question concept to the supporting morphologic signals and image regions.
B.1 Concept Vocabulary
The concept vocabulary covers common morphology-centered questions in breast pathology, together with metadata-oriented and morphology-based proxy tasks. Table 4 summarizes the concepts and their associated signals. The mapping is many-to-many because a single clinical concept may depend on several morphologic findings, and the same finding may support more than one concept. Concepts without a dedicated signal are answered using global and multi-scale regional evidence.
| ID | Concept | Associated signals |
| M01 | Tumor presence and distribution | S104 |
| M02 | Histological tumor type (coarse) | S103, S111, S113 |
| M03 | Tumor grade | S101, S102, S103, S104 |
| M04 | Necrosis assessment | S112, S132 |
| M05 | Inflammation / tumor-infiltrating lymphocytes | No dedicated signal; multi-scale evidence |
| M06 | In-situ component (DCIS-like) | S111, S112, S113 |
| M07 | Margin status | S141, S142 |
| M08 | Lymphovascular invasion | S121 |
| M09 | Perineural invasion | No dedicated signal; multi-scale evidence |
| M10 | Microcalcifications | S131 |
| M11 | Benign or non-neoplastic background changes | No dedicated signal; multi-scale evidence |
| MD01 | Specimen, procedure, and site metadata | Global evidence |
| MD02 | Tumor size, extent, and stage elements | S141 (weak proxy), with global evidence |
| P01 | Biomarker or outcome proxy prediction | S101–S104, S111, S113, S132 |
B.2 Concept-to-signal Mapping
Signals are affirmative, locally assessable morphologic findings rather than diagnostic labels. Table 5 lists the signal vocabulary used by SlideBank. Each signal is linked to the image regions in which it is observed, allowing the same evidence to be reused across related concepts.
| ID | Signal | Morphologic cue |
| S101 | High nuclear pleomorphism | Marked variation in nuclear size and shape. |
| S102 | High mitotic activity | Frequent mitotic figures indicating increased proliferation. |
| S103 | Low tubule formation | Reduced tubular differentiation with predominantly solid growth. |
| S104 | Support for high tumor grade | Combined visual evidence consistent with a higher-grade appearance. |
| S111 | DCIS present | In-situ ductal proliferation without clear stromal invasion. |
| S112 | DCIS with comedo necrosis | Intraluminal necrotic debris compatible with comedo morphology. |
| S113 | Solid or cribriform DCIS architecture | Solid or cribriform in-situ architecture within ductal units. |
| S121 | Lymphovascular invasion present | Suspicious tumor emboli within endothelial-lined spaces. |
| S131 | Microcalcification present | Fine basophilic calcific deposits in tumor-associated tissue. |
| S132 | Tumor necrosis present | Tumor-associated necrosis with loss of viable architecture. |
| S141 | Tumor close to margin | Tumor located close to a specimen edge in the sampled context. |
| S142 | Tumor involving margin | Tumor involving an edge or margin-like boundary. |
| S151 | Lymph-node metastasis present | Metastatic epithelial cells in a lymphoid background. |
| S152 | Micrometastasis present | A small metastatic focus compatible with micrometastatic burden. |
For each evidence region, a signal is represented as present, unknown, or absent, together with a confidence score and its supporting views. During retrieval, evidence linked to the target signals is prioritized, while uncertain or negative observations may also be retained when they are relevant to the question.
B.3 Multi-scale Evidence Organization
SlideBank organizes evidence at three complementary scales: a global view captures slide-level context, anchor regions localize diagnostically informative areas, and patch views preserve detailed architectural or cellular morphology. Each signal is linked to its supporting anchor and patch views. Consequently, the evidence supplied to the vision-language model retains both its spatial origin and its morphologic interpretation, enabling an answer to be traced back to the corresponding signal and image region.
Appendix C Datasets and Baselines
C.1 Dataset Statistics
WSI-VQA.
The standard evaluation set contains 388 four-option questions associated with 85 WSIs. Questions cover slide-level diagnosis and morphology as well as attributes such as grade, size, receptor status, margins, and staging. We use the multiple-choice conversion of the released annotations and preserve the original slide–question association and option text. Accuracy is computed over questions, so slides with more annotated questions contribute proportionally more examples.
SlideBench-BCNB.
The BCNB split contains 1,058 breast-cancer WSIs. Following the standard SlideBench protocol, we evaluate the three tasks that are answerable from the released multiple-choice annotations: tumor type (1,058 questions), histological grading (926 questions), and molecular subtype (1,058 questions), for 3,042 questions in total. Tumor-type and grading questions have three answer candidates, while molecular-subtype questions have four. We report each task separately and compute the overall BCNB result as the question-count-weighted average
| (14) |
C.2 Baselines
All methods are evaluated on the same questions and answer candidates. Unless otherwise specified, we follow the released inference protocol for each baseline.
General-purpose and domain-specific VLMs. We evaluate Qwen2.5-VL-7B Bai et al. (2025b), Qwen3-VL-8B Bai et al. (2025a), and Qwen3.5-4B Qwen Team (2026) as general-purpose VLMs, together with LLaVA-Med Li et al. (2023), Quilt-LLaVA Seyfioglu et al. (2024), PathGen-LLaVA Sun et al. (2025), and Patho-R1 Zhang et al. (2026) as medical- or pathology-tuned VLMs. Each model receives a single WSI thumbnail and the multiple-choice question, without access to additional high-resolution fields.
WSI-LLaVA Liang et al. (2025). WSI-LLaVA is a multimodal model designed for gigapixel WSI understanding. It uses hierarchical slide representations and a multi-stage alignment and instruction-tuning strategy to connect WSI morphology with language.
TITAN Ding et al. (2025). TITAN is a multimodal whole-slide foundation model pretrained through visual self-supervision and vision–language alignment. It encodes an entire WSI into a slide-level representation that can be transferred to downstream pathology tasks without task-specific fine-tuning.
SlideChat Chen et al. (2025b). SlideChat is a vision–language assistant developed specifically for gigapixel whole-slide pathology. It is instruction-tuned on slide-level captions and VQA pairs to support WSI description and question answering.
MedAgents Tang et al. (2024). MedAgents is a multi-agent medical reasoning framework in which specialized agents analyze a question from complementary perspectives and aggregate their conclusions. We include it to assess whether inference-time collaboration can improve WSI question answering without a dedicated evidence-retrieval module.We adapt the MedAgents multi-agent collaboration framework to the multimodal setting using GPT-4o as the backbone. Each agent receives the WSI thumbnail and the multiple-choice question, and the final answer is obtained through multi-round specialist discussion and consensus.
PathAgent Chen et al. (2025a). PathAgent is a training-free agentic framework for WSI analysis. It iteratively navigates the slide, extracts morphological evidence at selected regions and magnifications, and integrates the observations to produce an interpretable answer.For the backbone setting, we use the official PathAgent implementation.
Appendix D Additional Implementation Details
Fusion weights. We use fixed fusion weights across all datasets and backbones, without dataset-specific tuning.
Multi-turn evaluation. For each WSI-VQA question, Gemini-2.5-Flash generates five semantically equivalent rephrasings, and the answer options are independently shuffled while preserving the correct option text. After retaining complete six-form groups, the evaluation set contains 382 source questions and 2,292 turns from 80 WSIs, with an average of 29 turns per WSI. The evidence bank is constructed once per slide and remains fixed throughout the conversation; each turn uses the immediately preceding predicted question–answer pair as context, without access to ground-truth answers. Rephrase consistency is computed after mapping the shuffled option letters back to their semantic option text.
Answer parsing. We use a deterministic parser that first extracts a valid option letter from a JSON answer field or an explicit answer expression. If this fails, it accepts a standalone option letter or matches the normalized generated text to the provided option text. Outputs that cannot be resolved to a valid candidate are counted as incorrect, without using an additional language model for adjudication.