arXiv is now an independent nonprofit! Learn more
License: CC BY 4.0
arXiv:2609.00342v1 [cs.AI] 31 Aug 2026

1]University of British Columbia 2]Vector Institute 3]Harvard University 4]Rice University 5]University of Chicago 6]BC Cancer Agency \contribution∗ Equal contribution
† Correspondence to Xiaoxiao Li at xiaoxiao.li@ece.ubc.ca

SlideBank: A Persistent Hierarchical Evidence Bank for
Consistent Whole-Slide Reasoning

Beidi Zhao    Gexin Huang    Ciro Zhang    Anqi Li    Yusheng Tan    Chen Zhou    Gang Wang    Zu-hua Gao    Xiaoxiao Li Affiliation: [ Affiliation: [ Affiliation: [ Affiliation: [ Affiliation: [ Affiliation: [
Abstract

Whole-slide images (WSIs) are challenging for vision-language reasoning because diagnostically relevant morphology is sparse, heterogeneous, and distributed across gigapixel-scale images and multiple spatial resolutions. Existing WSI models and pathology agents can aggregate slide features or actively acquire evidence, but the information retained after exploration is often difficult to access semantically while preserving its connection to the original visual evidence. We introduce SlideBank, a training-free framework that represents each WSI as a persistent, concept-indexed, and spatially grounded evidence bank. SlideBank performs question-independent coarse-to-fine exploration to identify informative regions and multi-scale views, converts them into explicit morphological observations, and grounds pathology signals to their supporting patches and WSI coordinates. At inference time, questions are routed to relevant signals and evidence scales, and the linked global, regional, and patch evidence is integrated through confidence-based cross-level consensus. Experiments on WSI-VQA and SlideBench-BCNB show that with Patho-R1, SlideBank reaches 52.77% on WSI-VQA and with Quilt-LLaVA, it reaches 50.92% average accuracy on SlideBench-BCNB, while structured signal-guided retrieval consistently outperforms random evidence sampling. Reusing the same bank across repeated queries further achieves over 99% rephrasing consistency and substantially reduces amortized inference cost through persistent evidence reuse.

1 Introduction

Whole-slide images (WSIs) are among the most challenging visual inputs in computational pathology. Unlike conventional natural or medical images, a single WSI can contain billions of pixels and span tissue structures ranging from large-scale architectural patterns to fine-grained cellular morphology. Diagnostically relevant findings are often sparse, spatially dispersed, and highly heterogeneous: a small focus of atypical cells, an invasive tissue interface, or a mitotic hotspot may occupy only a tiny fraction of the entire slide while being decisive for diagnosis. Consequently, directly compressing a WSI into a thumbnail or uniformly processing a limited set of patches can easily discard critical evidence. Effective WSI reasoning therefore requires not only understanding what is visible, but also determining where diagnostically informative evidence lies, at what scale it should be examined, and how evidence distributed across the slide should be organized for downstream reasoning.

Recent work has approached this challenge from two major directions. WSI-level vision-language models Chen et al. (2025b); Liang et al. (2025); Ding et al. (2025); Alawode et al. (2026) aggregate precomputed patch or slide features into compact latent representations, enabling end-to-end reasoning over gigapixel images. More recently, pathology agents Ghezloo et al. (2025); Chen et al. (2025a); Yang et al. (2026) have treated WSI understanding as an active evidence-acquisition problem, using planning, region selection, and multi-scale navigation to locate diagnostically relevant fields before generating an answer. Several methods further introduce hierarchical representations, cached observations, or slide-specific memory to improve efficiency and reasoning over large visual contexts Alawode et al. (2026); Chen et al. (2025a); Yang et al. (2026); Huang et al. (2025); Li et al. (2026). These approaches have substantially advanced WSI understanding, but the information retained after slide exploration is typically represented either as latent visual features, flat collections of observations, or states optimized for a particular navigation or reasoning trajectory. How to organize visually established evidence from the current slide into an explicit representation that can be semantically accessed by different downstream questions remains comparatively underexplored.

This motivates the central question of this work: how can evidence acquired from a gigapixel WSI be transformed into a persistent representation that supports diverse downstream questions while preserving explicit traceability to the underlying visual observations? This perspective shifts the focus from evidence acquisition alone to how acquired evidence is organized and accessed after exploration. A useful evidence representation should expose clinically meaningful findings in a form that can bridge different question formulations, preserve the multi-scale structure of pathology evidence, and maintain a traceable connection between semantic findings and the fine-grained areas and WSI locations that support them.

Constructing such a representation introduces several challenges. First, exhaustive inspection is infeasible at gigapixel scale, yet aggressive sampling risks missing sparse or heterogeneous abnormalities; the system must therefore identify a compact but diagnostically diverse set of fields across both architectural and cellular scales. Second, the language used in downstream questions does not necessarily match the language of raw visual observations. Questions about nuclear grade, pleomorphism, or cellular atypia, for example, may rely on overlapping morphological signals even when their surface forms differ, making direct retrieval over free-form descriptions unreliable. Third, clinically meaningful evidence is inherently multi-scale and spatially grounded: a semantic finding should remain linked not only to a textual description, but also to the anchor region, magnification, visual field, and WSI coordinates from which it was derived. Finally, different questions may require different combinations of global context, regional architecture, and cellular detail, requiring an evidence interface that remains persistent while supporting query-adaptive composition.

To address these challenges, we propose SlideBank, an agent-driven framework that requires no task-specific training and converts each WSI into a concept-indexed, spatially grounded hierarchical evidence bank. To construct the bank, SlideBank performs question-independent coarse-to-fine exploration, using a global survey and category-aware sampling to identify diagnostically informative anchor regions and complementary architecture- and cell-level views. A pathology generative model converts the selected views into explicit morphological observations, which are subsequently normalized into clinically meaningful signals by a concept-guided pathology signal layer. Each signal is associated with its status, confidence, supporting anchors, multi-scale views, coordinates, and magnifications, establishing a traceable path from pathology signal to supporting evidence, visual field, and ultimately WSI location. At inference time, a signal-aware reader routes each question to a relevant pathology concept, uses the associated signals to retrieve supporting anchors and multi-scale views, and generates candidates independently at the global, anchor, and patch levels. These candidates are combined through confidence-based selection and cross-level consensus. In this way, SlideBank explores each slide once, organizes the observed morphology into a persistent evidence bank, and reuses the same bank to answer different downstream questions.

Our contributions are three-fold:

  • •

    We introduce SlideBank, a structured, slide-specific hierarchical evidence representation that organizes observed morphology through pathology signals while preserving explicit links to supporting regions, multi-scale views, and WSI coordinates.

  • •

    We develop a concept-to-signal retrieval method that couples concept-level question routing with signal-grounded evidence retrieval, providing a consistent semantic interface for indexing and reasoning across diverse question formulations while retaining traceability to the underlying visual evidence.

  • •

    We evaluate SlideBank on two public WSI question-answering benchmarks in terms of accuracy, efficiency, and cross-turn consistency. A controlled comparison using the same evidence bank shows the benefit of concept-to-signal retrieval over random evidence sampling.

Refer to caption
Figure 1: Overview of SlideBank. (a) A pathology agent explores the WSI and organizes multi-scale observations into a spatially grounded evidence bank. (b) At inference time, each question is routed to relevant pathology concepts and linked global-, anchor-, and patch-level evidence, whose predictions are combined through confidence-based selection and weighted cross-level consensus. (c) The same bank is reused across rephrased queries for stable and efficient reasoning.

2 Related Work

2.1 Vision–Language Models and Agentic Reasoning for Whole-Slide Pathology

Early pathology vision–language models mainly learned image–text alignment or instruction following from localized fields. CONCH learns transferable visual–language representations, while Quilt-LLaVA and PathChat support instruction following and open-ended interaction with histology images Lu et al. (2024a); Seyfioglu et al. (2024); Lu et al. (2024b). Because these models cannot directly process a complete WSI at native resolution, subsequent methods aggregate pre-extracted patch features into slide-level representations. WSI-VQA introduced generative question answering over WSIs, TITAN learned a multimodal whole-slide representation, and SlideChat and WSI-LLaVA connected slide features to language models for slide-level dialogue and reasoning Chen et al. (2024); Ding et al. (2025); Chen et al. (2025b); Liang et al. (2025). More recent systems improve hierarchical or large-scale slide understanding: MLLM-HWSI aligns cell-, patch-, region-, and slide-level tokens with pathology language, while ALPaCA and PRISM2 scale slide-level supervision through question answering and clinical dialogue Alawode et al. (2026); Gao et al. (2026); Vorontsov et al. (2026). In parallel, agentic systems make evidence acquisition more adaptive. PathFinder, WSI-Agents, SlideSeek, GIANT, and PathAgent introduce specialized agents, planning, tool use, iterative navigation, or evidence synthesis for gigapixel slides Ghezloo et al. (2025); Lyu et al. (2025); Weishaupt et al. (2025); Buckley et al. (2025); Chen et al. (2025a); HistoSelect performs question-guided coarse-to-fine selection, BEACON acquires patches according to expected information gain, and AdaptivePath learns question-agnostic multiscale navigation Huang et al. (2026); Wong et al. (2026); Chen et al. (2026). EviPathBench further shows that locating diagnostic regions remains substantially harder than reasoning over preselected evidence Liao et al. (2026). PathNavigate is particularly related to our setting because it performs a question-independent scan and maintains an online memory over frozen slide features before question-conditioned search Yang et al. (2026). However, existing methods primarily represent acquired evidence as model-internal features or transient reasoning states. SlideBank instead materializes selected multiscale image fields, morphological descriptions, pathology signals, and WSI coordinates as an independently inspectable bank that can be retrieved across multiple questions about the same slide.

2.2 Hierarchical Evidence and Memory

Hierarchical modeling provides a natural way to represent gigapixel WSIs. HIPT learns nested patch- and region-level representations, while H2-MIL organizes heterogeneous instances into a hierarchy for slide analysis Chen et al. (2022); Hou et al. (2022); however, both primarily structure latent visual features for downstream prediction. Concept bottleneck models expose human-interpretable intermediate variables, and post-hoc variants add such structure to pretrained models Koh et al. (2020); Yuksekgonul et al. (2023). Persistent memory has also been studied in long-horizon agents: Generative Agents store episodic experiences, MemGPT manages context through external memory, and A-MEM and G-Memory organize memories through linked or hierarchical structures Park et al. (2023); Packer et al. (2023); Xu et al. (2025); Zhang et al. (2025). Related ideas have recently appeared in pathology. SurvAgent constructs a multimodal case bank for cross-case survival prediction, whereas PathMem organizes structured diagnostic knowledge as long-term memory and transfers relevant content into working memory Huang et al. (2025); Li et al. (2026). These approaches focus on latent hierarchies, general episodic memory, cross-case experience, or domain knowledge rather than persistent spatial evidence from repeated examination of one WSI. SlideBank combines hierarchical visual evidence with an interpretable routing interface: concepts represent coarse question-side categories, signals represent localized image-derived findings, and a predefined ontology maps concepts to relevant signals. Each signal record retains its status and links to associated anchors, multiscale patch views, descriptions, and WSI coordinates.The resulting slide-specific bank is reused across questions, so every turn is answered from the persistent evidence bank, not only affected by the accumulated conversation history.

3 Method

3.1 Overview

Whole-slide reasoning is challenging since a gigapixel WSI exceeds the accepted maximum image size of current vision-language models (VLMs) and clinically relevant findings may occupy only a small fraction of the slide. SlideBank therefore separates slide exploration from question answering. As illustrated in Fig. 1, the framework consists of three components: Agent-Driven Hierarchical Evidence Bank Construction organizes multi-scale slide observations into a persistent, spatially grounded evidence bank; Hierarchical Evidence-Augmented WSI Reasoning retrieves and integrates question-relevant evidence from the bank; and Persistent Evidence Reuse Across Queries enables the same slide representation to support subsequent questions without reconstructing the underlying evidence.

Let a whole slide image be represented by a resolution pyramid s={L0,…,LH}s=\{L_{0},\ldots,L_{H}\}, where L0L_{0} and LHL_{H} denote the highest- and lowest-resolution views, respectively. In Section 3.2, SlideBank performs question-independent coarse-to-fine exploration to construct a slide-specific evidence bank MsM_{s} that organizes multi-scale observations and links pathology signals to their supporting evidence and WSI locations. In Section 3.3, the online reader maps a question qtq_{t} to relevant signals and evidence scales, retrieves the linked evidence, and combines global, regional, and patch-level predictions through confidence-based consensus. Finally, Section 3.4 reuses the same MsM_{s} across subsequent queries.

3.2 Agent-Driven Hierarchical Evidence Bank Construction

Global Survey and Coordinate Normalization. To capture slide-level contextual cues and facilitate global morphological measurements, such as tumor size and disease extent, we first construct a global thumbnail. For pyramidal WSIs, we use the lowest-resolution level; for non-pyramidal or single-resolution slides, we construct an equivalent bounded thumbnail via chunk-wise subsampling. All spatial coordinates identified on the thumbnail are proportionally mapped back to the level-0 coordinate system of the source WSI. Foreground tissue is separated from the glass background using HSV thresholding, and the valid tissue area is partitioned into a non-overlapping 384×384384\times 384 grid. Each tile is characterized using handcrafted morphological statistics and foreground occupancy, and assigned to one of five coarse tissue classes: tumor, necrosis, adipose, stroma, or other. When physical pixel spacing is available, we additionally estimate the maximum Feret diameter and principal axis of the primary tumor candidate. These slide-level observations, including the natural-language summary, tissue-composition statistics, and geometric diagnostics, are stored in a global record GsG_{s} for slide ss.

Coarse-to-fine Agentic Sampling. Tiles are first ranked according to coarse diagnostic priority, triage confidence, and tissue coverage, and retained under a category-aware budget that allocates more samples to suspicious regions while preserving coverage of the remaining tissue classes. Sampling then proceeds through recursive 4×44\times 4 refinement. For each selected tile, the agent scores all 16 cells of its downsampled representation to propose anchor regions. Within each anchor region, the same procedure is repeated to identify informative 20×20\times and 10×10\times patches. At both stages, scoring is guided by a fixed set of pathology-derived cues provided in the prompt (prompts used throughout SlideBank are provided in Appendix).

Multi-Scale Visual–Textual Observations. The resulting hierarchy separates coarse triage from fine-grained evidence: tiles provide the coarse sampling prior, whereas anchor regions and their associated patches provide the observed morphological evidence. For each patch pp, a pathology-specific VLM generates a concise morphological description dpd_{p}, together with a self-reported confidence score and a quality-control flag, conditioned on the patch image and its magnification. Descriptions with confidence below a predefined threshold or invalid quality-control flags are discarded.

For each retained anchor region aa, the descriptions of its multi-magnification patches are collected as

𝒟a={dp∣p∈𝒫a},\mathcal{D}_{a}=\{\,d_{p}\mid p\in\mathcal{P}_{a}\,\}, (1)

where 𝒫a\mathcal{P}_{a} denotes the patches associated with anchor aa. This anchor-level textual observation summarizes complementary morphology across magnifications while retaining links to the originating patches and their WSI coordinates. We write 𝒜s\mathcal{A}_{s} for the set of anchor regions retained for slide ss, and 𝒫s=⋃a∈𝒜s𝒫a\mathcal{P}_{s}=\bigcup_{a\in\mathcal{A}_{s}}\mathcal{P}_{a} for all multi-scale patch views collected under them; each p∈𝒫sp\in\mathcal{P}_{s} stores its image, magnification, level-0 coordinates, and description dpd_{p}.

Pathology Signal Grounding. Free-form descriptions may express the same finding in different ways. We therefore map 𝒟a\mathcal{D}_{a} to a predefined vocabulary of localized pathology signals, such as marked nuclear pleomorphism, high mitotic activity, low tubule formation, and tumor necrosis. For each applicable signal kk, we store

σa,k=(ya,k,ρa,k,𝒫a,k+),\sigma_{a,k}=\left(y_{a,k},\rho_{a,k},\mathcal{P}^{+}_{a,k}\right), (2)

where ya,k∈{present,absent,unknown}y_{a,k}\in\{\mathrm{present},\mathrm{absent},\mathrm{unknown}\} is the inferred status, ρa,k\rho_{a,k} is its confidence, and 𝒫a,k+⊆𝒫a\mathcal{P}^{+}_{a,k}\subseteq\mathcal{P}_{a} contains the patch views supporting a positive assessment. This standardized signal layer preserves links to the underlying descriptions, images, and WSI locations.

Persistent Hierarchical Storage. The signal records across all anchors are collected as

𝒮s={σa,k∣a∈𝒜s,k∈𝒦a},\mathcal{S}_{s}=\{\,\sigma_{a,k}\mid a\in\mathcal{A}_{s},\ k\in\mathcal{K}_{a}\,\}, (3)

where 𝒦a\mathcal{K}_{a} is the set of signals applicable to anchor aa. We also define an ontology Ω\Omega that maps each coarse pathology concept c∈𝒞c\in\mathcal{C} to a set of relevant signals 𝒦⁡(c)\mathcal{K}(c). The slide bank is then represented as

Ms=(Gs,𝒜s,𝒫s,𝒮s,Ω),M_{s}=\bigl(G_{s},\ \mathcal{A}_{s},\ \mathcal{P}_{s},\ \mathcal{S}_{s},\ \Omega\bigr), (4)

comprising the global record, the retained anchor regions, their multi-scale patch views, the grounded signal records, and the concept-to-signal ontology. The complete pathology concept vocabulary, concept-to-signal mapping and evidence-bank schema are provided in Appendix.

3.3 Hierarchical Evidence-Augmented WSI Reasoning

Question-to-Concept Routing and Signal Retrieval. A pathology concept represents the coarse intent of a question, such as tumor grading, necrosis assessment, or margin status, whereas a pathology signal represents a localized image-derived finding. Given a question qq and its options OO, a deterministic router ψ\psi returns the top-nn pathology concepts

𝒞q=ψ⁡(q,O),|𝒞q|=n,𝒦q=⋃c∈𝒞q𝒦⁡(c).\mathcal{C}_{q}=\psi(q,O),\quad|\mathcal{C}_{q}|=n,\qquad\mathcal{K}_{q}=\bigcup_{c\in\mathcal{C}_{q}}\mathcal{K}(c). (5)

ψ\psi is based on lexical rules over the pathology concept vocabulary 𝒞\mathcal{C} (details in Appendix). We use n=3n=3 by default, sensitivity to n is analyzed in Sec. 4.3. The target signals are then used to retrieve

𝒮⁡(q)={σa,k∈𝒮s∣k∈𝒦q}.\mathcal{S}(q)=\{\,\sigma_{a,k}\in\mathcal{S}_{s}\mid k\in\mathcal{K}_{q}\,\}. (6)

For each target signal, SlideBank selects the anchor regions in which that signal is most strongly supported. Records labeled as present are preferred over uncertain or absent records, and records with the same status are ranked by confidence. The associated anchor images, patch views, and morphological descriptions are then retrieved as the question-specific evidence E⁡(q)=(Gs,𝒜q,𝒫q,𝒮q)E(q)=(G_{s},\mathcal{A}_{q},\mathcal{P}_{q},\mathcal{S}_{q}).

Evidence-Conditioned Branch Predictions. The question-specific evidence E⁡(q)=(Gs,𝒜q,𝒫q,𝒮q)E(q)=(G_{s},\mathcal{A}_{q},\mathcal{P}_{q},\mathcal{S}_{q}) is organized into three complementary branches ℓ∈{G,A,P}\ell\in\{G,A,P\}. The global branch captures slide-wide context, the anchor branch represents regional architecture, and the patch branch provides localized morphology. Anchor and patch inputs include their stored morphological descriptions, whereas the global branch uses the slide thumbnail without a local description.

Each evidence item is evaluated independently by the inference VLM. Let 𝒪\mathcal{O} denote the answer-option set and zℓ,i​(o)z_{\ell,i}(o) the logit assigned to option o∈𝒪o\in\mathcal{O} for item ii from branch ℓ\ell. Its option probability is

πℓ,i​(o)=exp⁡zℓ,i​(o)∑o′∈𝒪exp⁡zℓ,i​(o′).\pi_{\ell,i}(o)=\frac{\exp z_{\ell,i}(o)}{\sum_{o^{\prime}\in\mathcal{O}}\exp z_{\ell,i}(o^{\prime})}. (7)

The corresponding prediction and confidence are

o^ℓ,i=arg⁡maxo∈𝒪​πℓ,i​(o),γℓ,i=maxo∈𝒪⁡πℓ,i​(o).\hat{o}_{\ell,i}=\arg\max_{o\in\mathcal{O}}\pi_{\ell,i}(o),\qquad\gamma_{\ell,i}=\max_{o\in\mathcal{O}}\pi_{\ell,i}(o). (8)

Within the anchor and patch branches, candidates are ranked by γℓ,i\gamma_{\ell,i}, and the top-mm candidates are retained as 𝒯ℓ​(q)\mathcal{T}_{\ell}(q).

Level-Weighted Answer CSonsensus. For each local branch ℓ∈{A,P}\ell\in\{A,P\}, the retained candidates are aggregated by majority voting. The resulting support for option oo is

vℓ(o∣q)=1|𝒯ℓ​(q)|∑i∈𝒯ℓ​(q)𝕀[o^ℓ,i=o].v_{\ell}(o\mid q)=\frac{1}{|\mathcal{T}_{\ell}(q)|}\sum_{i\in\mathcal{T}_{\ell}(q)}\mathbb{I}\!\left[\hat{o}_{\ell,i}=o\right]. (9)

For the global branch, the option probabilities are used directly:

vG​(o∣q)=πG,1​(o).v_{G}(o\mid q)=\pi_{G,1}(o). (10)

The three branches are combined using

F⁡(o∣q)=∑ℓ∈{G,A,P}λℓ​vℓ​(o∣q),∑ℓλℓ=1,F(o\mid q)=\sum_{\ell\in\{G,A,P\}}\lambda_{\ell}v_{\ell}(o\mid q),\qquad\sum_{\ell}\lambda_{\ell}=1, (11)

and the final answer is

o^​(q)=arg⁡maxo∈𝒪⁡F⁡(o∣q).\hat{o}(q)=\arg\max_{o\in\mathcal{O}}F(o\mid q). (12)

We use λG=0.2\lambda_{G}=0.2, λA=0.4\lambda_{A}=0.4, and λP=0.4\lambda_{P}=0.4, placing greater emphasis on localized morphological evidence while retaining regional and global context. The retained candidates, signal records, and spatial evidence links provide a trace from the final answer to its supporting WSI regions.

3.4 Persistent Evidence Reuse Across Queries

SlideBank constructs the evidence bank MsM_{s} once and reuses it across all turns. For each question qtq_{t}, the reader performs question-dependent routing to obtain evidence E⁡(qt)E(q_{t}) from the bank and predicts the answer with the evidence, avoiding repeated WSI exploration and reducing evidence drift. Routing depends only on the current question, so the retrieved evidence for a given question is identical whether it is asked in isolation or within a conversation.

To preserve conversational continuity without accumulating a long history, the reader carries only the previous question–answer pair,

Ht=(qt−1,o^​(qt−1)),H_{t}=\bigl(q_{t-1},\hat{o}(q_{t-1})\bigr), (13)

which is supplied to the inference VLM alongside the retrieved evidence. The visual evidence therefore always comes from the persistent bank, while HtH_{t} provides bounded dialogue context.

4 Experiments

Datasets. We evaluate SlideBank on two public WSI-level benchmarks: WSI-VQA Chen et al. (2024), which contains 85 WSIs with slide-level visual question answering annotations, and SlideBench-BCNB Chen et al. (2025b), which contains 1,058 breast cancer slides covering tumor typing, grading, and molecular subtyping. The two benchmarks provide complementary settings for evaluating general WSI question answering and clinically relevant diagnostic reasoning. For repeated-query evaluation, we additionally construct a multi-turn setting from WSI-VQA by grouping all questions associated with the same slide into one conversation, resulting in an average of 29 questions per WSI. We further generate five semantically equivalent formulations of each question with shuffled answer options using Gemini-2.5-Flash Comanici et al. (2025) to evaluate the stability of predictions under different question formulations.

Baselines. We compare against three families of methods. Zero-shot VLMs include general-purpose and pathology-specific models that receive only a 1024×10241024\times 1024 WSI thumbnail. WSI-trained models include WSI-LLaVA, TITAN, and SlideChat, which are trained or adapted for slide-level understanding. Agentic systems include Med-Agents and PathAgent, which introduce additional inference-time reasoning or WSI exploration. All methods are evaluated on the same question sets and answer spaces. Thumbnail-based baselines receive no additional high-resolution patches. We evaluate SlideBank with Quilt-LLaVA and Patho-R1 to examine backbone transferability.

Metrics. For standard WSI QA, we report accuracy on WSI-VQA and question-count-weighted average over the three BCNB tasks (tumor, histological grading and molecular subtype). For repeated-query evaluation, we additionally report accuracy, the change from single-query accuracy (Δ\DeltaAcc), and rephrase consistency (Consi=𝕀[maxa∑j=05𝕀(a^i(j)=a)≥3],\mathrm{Cons}_{i}=\mathbb{I}\left[\max_{a}\sum_{j=0}^{5}\mathbb{I}\bigl(\hat{a}_{i}^{(j)}=a\bigr)\geq 3\right],) across semantically equivalent question formulations.

Implementation Details. Within each experimental setting, we use the same pathology VLM backbone for both evidence-bank construction and online inference, without additional training. We show the result with Quilt-LLaVA Seyfioglu et al. (2024) and Patho-R1 Zhang et al. (2026) as the backbone. For evidence-bank construction, we partition the slide thumbnail into ×384384\!\times\!384 tiles, retain those with at least 20%20\% tissue coverage, and perform both tile-level anchor localization and within-anchor patch selection on a ×44\!\times\!4 grid. We select three anchors from each tumor-like tile and two from each retained non-tumor tile. For each anchor, we extract an adaptive context crop capped at 4096 level-0 pixels, a 512-pixel fine-detail view corresponding to approximately 20×20\times on slides scanned at 40×40\times, and an additional architectural view at approximately 10×10\times. All views are resized to ×256256\!\times\!256 for VLM processing. During online inference, we use all retrieved anchor-level evidence and sample at most 32 patch-level evidence items with a fixed random seed 42. We retain the top five (mm=5) candidates from the anchor and patch branches according to token-probability confidence and fuse the global, anchor, and patch predictions using (λG,λA,λP)=(0.2,0.4,0.4)(\lambda_{G},\lambda_{A},\lambda_{P})=(0.2,0.4,0.4). All experiments use a single NVIDIA A100 GPU with 80 GB memory, with FP16 inference for Quilt-LLaVA and BF16 inference for Patho-R1. Additional dataset, baseline, multi-turn evaluation, and answer-parsing details are provided in Appendix.

Table 1: Standard WSI question-answering performance. We report standard single-query performance to examine whether this reusable representation preserves sufficient information for downstream WSI question answering. Bold and underline show the best and second best result.
Method Inference Input SlideBench-BCNB Accuracy (%) ↑\uparrow WSI-VQA (%) ↑\uparrow
Tumor Type Grading Subtype Weighted Avg. Accuracy
Vision-Language Models (Zero-shot)
Qwen2.5-VL-7B Bai et al. (2025b) Thumbnail 58.13 27.11 16.07 34.06 33.43
Qwen3-VL-8B Bai et al. (2025a) Thumbnail 55.01 27.65 16.92 33.43 39.92
Qwen3.5-4B Qwen Team (2026) Thumbnail 39.89 38.34 17.30 31.56 44.53
LLaVA-Med Li et al. (2023) Thumbnail 26.56 27.75 22.40 25.48 31.87
Quilt-LLaVA Seyfioglu et al. (2024) Thumbnail 44.33 32.18 25.05 33.93 27.65
PathGen-LLaVA Sun et al. (2025) Thumbnail 45.18 33.05 22.68 33.66 39.90
Patho-R1 Zhang et al. (2026) Thumbnail 61.11 29.93 27.57 39.95 44.28
WSI-trained Models
WSI-LLaVA-7B Liang et al. (2025) Slide 42.53 30.02 32.42 35.21 49.61
TITAN Ding et al. (2025) Slide 83.55 27.00 21.27 44.67 43.12
SlideChat Chen et al. (2025b) Slide 87.99 17.86 20.41 43.14 52.05
Agentic Systems
MedAgents Tang et al. (2024) Thumbnail 89.38 37.37 15.12 47.72 42.75
PathAgent Chen et al. (2025a) Slide 50.76 46.00 27.08 41.08 52.05
SlideBank (w/ Quilt-LLaVA) Thumbnail + Evidence 89.82 33.56 27.22 50.92 43.75
SlideBank (w/ Patho-R1) Thumbnail + Evidence 83.27 40.06 26.47 50.36 52.77

4.1 Standard WSI Question Answering

Table 1 evaluates whether the question-independent evidence bank preserves sufficient slide information for downstream WSI reasoning. Despite requiring no task-specific training, SlideBank achieves strong performance on both benchmarks. On SlideBench-BCNB, SlideBank with Quilt-LLaVA obtains 89.82% accuracy on tumor typing and the highest overall average of 50.92%, substantially improving over the thumbnail-only Quilt-LLaVA backbone. On WSI-VQA, SlideBank with Patho-R1 reaches 52.77%, outperforming its thumbnail-only counterpart by 8.49 percentage points and achieving the best performance among the evaluated methods. These results show that organizing gigapixel WSI observations into a reusable structured evidence bank preserves and can improve standard single-query reasoning performance.

4.2 Effect of Evidence Organization and Access

Multi-scale Hierarchy. We first examine how different levels of the evidence hierarchy contribute to WSI reasoning. As shown in Table 2, local evidence is substantially more informative than the global thumbnail alone. On WSI-VQA, the anchor and patch branches achieve 51.98% and 51.45% accuracy, respectively, compared with 48.55% using only global evidence. A similar trend is observed on SlideBench-BCNB, where anchor-level evidence provides the strongest single-level performance at 50.23%. Combining global, anchor, and patch evidence shows the best overall results on both benchmarks, reaching 52.77% on WSI-VQA and 50.36% on SlideBench-BCNB.

Table 2: Ablation on hierarchical evidence sources and evidence access strategies (w/ Patho-R1).
Variant Evidence WSI-VQA BCNB
Evidence hierarchy
Global only G 48.55 40.57
Anchor only A 51.98 50.23
Patch only P 51.45 48.06
Full hierarchy G+A+P 52.77 50.36
Evidence access
Random G+A+P 50.13 49.62
Signal-guided G+A+P 52.77 50.36

Signal-guided Access. We next show whether the gain arises from structured evidence access rather than simply exposing the model to more evidence. Using the same constructed evidence bank, we replace signal-guided retrieval with uniformly sampling 32 patches at random while keeping other settings unchanged. Signal-guided retrieval improves accuracy from 50.13% to 52.77% on WSI-VQA and from 49.62% to 50.36% on SlideBench-BCNB. The improvement shows the benefit of organizing and accessing slide evidence through pathology signals compared to random sampling.

4.3 Sensitivity and Qualitative Example

Sensitivity of the Number of Concepts.

Refer to caption
Figure 2: Sensitivity to the number of pathology concepts used for concept-routed retrieval (w/ Patho-R1).

We study the sensitivity to the number of pathology concepts used for concept-routed retrieval. As shown in Fig. 2, accuracy increases from 49.9% with a single concept to a peak of 52.8% with three concepts, and then gradually decreases as additional concepts are included. This reflects a trade-off between evidence coverage and retrieval noise: too few concepts may omit relevant pathology signals, whereas overly broad routing introduces less informative evidence. Performance remains above random sampling across a broad range of 2–6 concepts, indicating that the method is not sensitive to a narrowly tuned setting. We therefore use three concepts by default.

Refer to caption
Figure 3: Qualitative evidence trace in SlideBank. Routed pathology signals retrieve supporting multi-scale evidence, which is fused to produce the final prediction.

Qualitative Example. Figure 3 illustrates the complete evidence trace for a representative grading question. The query is first routed to grade-related signals, which retrieve their linked regions and multi-magnification patches. This example demonstrates how SlideBank preserves an explicit path from the final prediction back to the pathology signals and their supporting visual evidence.

Refer to caption
Figure 4: Efficiency and stability under repeated WSI queries. (a) Amortized runtime per query when reusing the evidence bank (w/ Quilt-LLaVA). (b) Runtime and consistency comparison among high-consistency methods.

4.4 Persistent Evidence Reuse Across Queries

Rephrased-QSuery Stability. We evaluate whether a persistent slide-specific evidence bank provides a stable basis for answering repeated questions about the same WSI. We group all WSI-VQA questions associated with each slide into a single sequence, leading to 29 questions per WSI on average, and present each question in a rephrased form with shuffled options. As shown in Table 3, several baselines suffer substantial accuracy degradation compared with independent single-query inference. In contrast, SlideBank shows only modest drops of 1.38 and 3.83 percentage points with Quilt-LLaVA and Patho-R1, respectively, while achieving 99.47% and 99.21% consistency across semantically equivalent question formulations. Since routing depends only on the current question, these small drops arise from question rephrasing and the bounded dialogue context rather than from re-exploring the slide like baselines.

Table 3: Performance comparison on WSI-VQA under the multi-turn setting. Rephrase Cons. measures the prediction consistency across semantically equivalent question formulations.
Method Acc ↑\uparrow Δ\DeltaAcc Rephrase Cons. ↑\uparrow
LLaVA-Med Li et al. (2023) 16.18 -15.69 62.83
Quilt-LLaVA Seyfioglu et al. (2024) 25.22 -2.43 84.25
PathGen-LLaVA Sun et al. (2025) 23.06 -16.84 87.27
Patho-R1 Zhang et al. (2026) 41.93 -2.35 80.94
WSI-LLaVA Liang et al. (2025) 17.06 -32.55 73.21
SlideChat Chen et al. (2025b) 42.50 -9.55 88.69
SlideBank (w/ Quilt-LLaVA) 42.37 -1.38 99.47
SlideBank (w/ Patho-R1) 48.94 -3.83 99.21

Amortized Efficiency. Figure 4 evaluates the computational benefit of reusing the same evidence bank across repeated queries. Although bank construction introduces a one-time cost, this cost is progressively amortized as more questions are asked about the same slide. The average runtime decreases from 94.3 s for a single query to 5.9 s per query over the question sequence, approaching an online reasoning cost of 2.73 s per query. We restrict the comparison to generative VLMs and WSI assistants with comparable conversational inference interfaces; representation-oriented models such as TITAN and agentic methods whose multi-round computation occurs internally within a single query are excluded. As shown in Fig. 4(b), SlideBank has higher consistency and lower runtime than baselines.

5 Conclusion

We introduced SlideBank, a training-free framework that converts each WSI into a persistent, concept-indexed, and spatially grounded evidence bank. SlideBank separates slide exploration from question answering, organizes multi-scale morphology through pathology signals, and retrieves question-relevant evidence with explicit links to supporting regions and patches. Experiments on WSI-VQA and SlideBench-BCNB show competitive performance, improved stability across repeated queries, and efficiency gains from evidence reuse. These results highlight structured evidence organization as a promising direction for WSI reasoning.

References

  • Alawode et al. (2026) B. Alawode, A. Mahmood, M. K. Al Radi, S. Albastaki, A. Khan, M. Bilal, M. A. Abdalla, M. Bennamoun, and S. Javed MLLM-hwsi: a multimodal large language model for hierarchical whole slide image understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13732–13743. Cited by: §1, §2.1.
  • Bai et al. (2025a) S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, et al. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631. Cited by: §C.2, Table 1.
  • Bai et al. (2025b) S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y. Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y. Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin Qwen2.5-vl technical report. External Links: 2502.13923, Link Cited by: §C.2, Table 1.
  • Buckley et al. (2025) T. A. Buckley, K. R. Weihrauch, K. Latham, A. Z. Zhou, P. A. Manrai, and A. K. Manrai Navigating gigapixel pathology images with large multimodal models. arXiv preprint arXiv:2511.19652. Cited by: §2.1.
  • Chen et al. (2025a) J. Chen, L. Cai, Z. Wang, Y. Huang, S. Jiang, S. Huang, H. Wang, and Y. Zhang Pathagent: toward interpretable analysis of whole-slide pathology images via large language model-based agentic reasoning. arXiv preprint arXiv:2511.17052. Cited by: §C.2, §1, §2.1, Table 1.
  • Chen et al. (2026) J. Chen, F. Liu, L. Cai, S. Jiang, S. Huang, H. Wang, L. Yu, and Y. Zhang Agentic visual reasoning in whole-slide pathology images via active perception. arXiv preprint arXiv:2608.08648. External Links: Document Cited by: §2.1.
  • Chen et al. (2024) P. Chen, C. Zhu, S. Zheng, H. Li, and L. Yang Wsi-vqa: interpreting whole slide images by generative visual question answering. In European Conference on Computer Vision, pp. 401–417. Cited by: §2.1, §4.
  • Chen et al. (2022) R. J. Chen, C. Chen, Y. Li, T. Y. Chen, A. D. Trister, R. G. Krishnan, and F. Mahmood Scaling vision transformers to gigapixel images via hierarchical self-supervised learning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 16144–16155. Cited by: §2.2.
  • Chen et al. (2025b) Y. Chen, G. Wang, Y. Ji, Y. Li, J. Ye, T. Li, M. Hu, R. Yu, Y. Qiao, and J. He Slidechat: a large vision-language assistant for whole-slide pathology image understanding. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 5134–5143. Cited by: §C.2, §1, §2.1, Table 1, Table 3, §4.
  • Comanici et al. (2025) G. Comanici, E. Bieber, M. Schaekermann, I. Pasupat, N. Sachdeva, I. Dhillon, M. Blistein, O. Ram, D. Zhang, E. Rosen, et al. Gemini 2.5: pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. arXiv preprint arXiv:2507.06261. Cited by: §4.
  • Ding et al. (2025) T. Ding, S. J. Wagner, A. H. Song, R. J. Chen, M. Y. Lu, A. Zhang, A. J. Vaidya, G. Jaume, M. Shaban, A. Kim, et al. A multimodal whole-slide foundation model for pathology. Nature medicine, pp. 1–13. Cited by: §C.2, §1, §2.1, Table 1.
  • Gao et al. (2026) Z. Gao, K. He, W. Su, X. Pang, I. P. Machado, M. Jimenez-Linan, B. Rous, C. Wang, C. Li, W. McGough, S. Gao, D. Zhang, T. Gong, M. Y. Lu, F. Mahmood, M. Feng, C. Li, and M. Crispin-Ortuzar ALPaCA: adapting llama for pathology context analysis to enable slide-level question answering. Nature Communications. External Links: Document Cited by: §2.1.
  • Ghezloo et al. (2025) F. Ghezloo, M. S. Seyfioglu, R. Soraki, W. O. Ikezogwo, B. Li, T. Vivekanandan, J. G. Elmore, R. Krishna, and L. Shapiro PathFinder: a multi-modal multi-agent system for medical diagnostic decision-making applied to histopathology. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 23431–23441. Cited by: §1, §2.1.
  • Hou et al. (2022) W. Hou, L. Yu, C. Lin, H. Huang, R. Yu, J. Qin, and L. Wang Hˆ 2-mil: exploring hierarchical representation with heterogeneous multiple instance learning for whole slide image analysis. In Proceedings of the AAAI conference on artificial intelligence, Vol. 36, pp. 933–941. Cited by: §2.2.
  • Huang et al. (2025) G. Huang, W. Chen, J. Yang, X. Lyu, X. Luo, S. Yang, X. Xing, and L. Shen SurvAgent: hierarchical cot-enhanced case banking and dichotomy-based multi-agent system for multimodal survival prediction. arXiv preprint arXiv:2511.16635. Cited by: §1, §2.2.
  • Huang et al. (2026) W. Huang, W. Lyu, P. Lou, Q. Hu, X. Hu, S. Abousamra, W. Han, R. Guo, J. Zhou, C. Chen, and C. Wang Act like a pathologist: tissue-aware whole slide image reasoning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6972–6981. Cited by: §2.1.
  • Koh et al. (2020) P. W. Koh, T. Nguyen, Y. S. Tang, S. Mussmann, E. Pierson, B. Kim, and P. Liang Concept bottleneck models. In Proceedings of the 37th International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 119, pp. 5338–5348. Cited by: §2.2.
  • Li et al. (2023) C. Li, C. Wong, S. Zhang, N. Usuyama, H. Liu, J. Yang, T. Naumann, H. Poon, and J. Gao Llava-med: training a large language-and-vision assistant for biomedicine in one day. Advances in Neural Information Processing Systems 36, pp. 28541–28564. Cited by: §C.2, Table 1, Table 3.
  • Li et al. (2026) J. Li, Y. Liang, Q. Li, X. Lyu, J. Qian, H. Chen, K. Wang, Z. Zeng, A. A. Bharath, and Y. Liu PathMem: toward cognition-aligned memory transformation for pathology mllms. arXiv preprint arXiv:2603.09943. Cited by: §1, §2.2.
  • Liang et al. (2025) Y. Liang, X. Lyu, W. Chen, M. Ding, J. Zhang, X. He, S. Wu, X. Xing, S. Yang, X. Wang, and L. Shen WSI-llava: a multimodal large language model for whole slide image. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp. 22718–22727. Cited by: §C.2, §1, §2.1, Table 1, Table 3.
  • Liao et al. (2026) D. Liao, T. Zhang, Y. Wu, X. Zhang, Q. Xue, Z. Liu, D. Zhao, L. Cai, and Y. Jin EviPathBench: benchmarking evidence acquisition and reasoning in vision-language models for whole-slide pathology. arXiv preprint arXiv:2607.19261. External Links: Document Cited by: §2.1.
  • Lu et al. (2024a) M. Y. Lu, B. Chen, D. F. Williamson, R. J. Chen, I. Liang, T. Ding, G. Jaume, I. Odintsov, L. P. Le, G. Gerber, et al. A visual-language foundation model for computational pathology. Nature medicine 30 (3), pp. 863–874. Cited by: §2.1.
  • Lu et al. (2024b) M. Y. Lu, B. Chen, D. F. K. Williamson, R. J. Chen, M. Zhao, et al. A multimodal generative ai copilot for human pathology. Nature 634, pp. 466–473. External Links: Document Cited by: §2.1.
  • Lyu et al. (2025) X. Lyu, Y. Liang, W. Chen, M. Ding, J. Yang, G. Huang, D. Zhang, X. He, and L. Shen Wsi-agents: a collaborative multi-agent system for multi-modal whole slide image analysis. arXiv preprint arXiv:2507.14680. Cited by: §2.1.
  • Packer et al. (2023) C. Packer, S. Wooders, K. Lin, V. Fang, S. G. Patil, I. Stoica, and J. E. Gonzalez Memgpt: towards llms as operating systems. arXiv preprint arXiv:2310.08560. Cited by: §2.2.
  • Park et al. (2023) J. S. Park, J. O’Brien, C. J. Cai, M. R. Morris, P. Liang, and M. S. Bernstein Generative agents: interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology, pp. 1–22. Cited by: §2.2.
  • Qwen Team (2026) Qwen Team Qwen3.5: towards native multimodal agents. External Links: Link Cited by: §C.2, Table 1.
  • Seyfioglu et al. (2024) M. S. Seyfioglu, W. O. Ikezogwo, F. Ghezloo, R. Krishna, and L. Shapiro Quilt-llava: visual instruction tuning by extracting localized narratives from open-source histopathology videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13183–13192. Cited by: §C.2, §2.1, Table 1, Table 3, §4.
  • Sun et al. (2025) Y. Sun, Y. Zhang, Y. Si, C. Zhu, K. Zhang, Z. Shui, J. Li, X. Gong, X. Lyu, T. Lin, et al. Pathgen-1.6 m: 1.6 million pathology image-text pairs generation through multi-agent collaboration. In International Conference on Learning Representations, Vol. 2025, pp. 94611–94653. Cited by: §C.2, Table 1, Table 3.
  • Tang et al. (2024) X. Tang, A. Zou, Z. Zhang, Z. Li, Y. Zhao, X. Zhang, A. Cohan, and M. Gerstein Medagents: large language models as collaborators for zero-shot medical reasoning. In Findings of the Association for Computational Linguistics: ACL 2024, pp. 599–621. Cited by: §C.2, Table 1.
  • Vorontsov et al. (2026) E. Vorontsov, G. Shaikovski, A. Casson, J. Viret, E. Zimmermann, N. Tenenholtz, Y. K. Wang, J. H. Bernhard, R. A. Godrich, J. A. Retamero, J. Shia, M. Gonen, M. R. Weiser, D. S. Klimstra, R. Yousfi, N. Fusi, T. J. Fuchs, K. Severson, and S. Liu End-to-end multimodal pathology foundation model with clinical dialogue. Nature Medicine. External Links: Document Cited by: §2.1.
  • Weishaupt et al. (2025) L. L. Weishaupt, C. Chen, D. F. Williamson, R. J. Chen, G. Jaume, T. Ding, B. Chen, A. Vaidya, L. P. Le, M. Y. Lu, et al. Evidence-based diagnostic reasoning with multi-agent copilot for human pathology. arXiv preprint arXiv:2506.20964. Cited by: §2.1.
  • Wong et al. (2026) B. Wong, X. Xu, H. Fu, N. F. Chen, and M. Y. Yi Beyond relevance: bayesian evidence acquisition for agentic whole-slide image reasoning. arXiv preprint arXiv:2608.05757. External Links: Document Cited by: §2.1.
  • Xu et al. (2025) W. Xu, Z. Liang, K. Mei, H. Gao, J. Tan, and Y. Zhang A-mem: agentic memory for llm agents. arXiv preprint arXiv:2502.12110. Cited by: §2.2.
  • Yang et al. (2026) C. Yang, Q. Liu, W. Zhao, Y. Tang, J. Ge, D. Zhang, J. Liu, L. Wu, J. Lu, N. Zhang, et al. PathNavigate: a training-free pathology agent with surprise-guided scan and shared slide memory for whole-slide image vqa. arXiv preprint arXiv:2605.23559. Cited by: §1, §2.1.
  • Yuksekgonul et al. (2023) M. Yuksekgonul, M. Wang, and J. Zou Post-hoc concept bottleneck models. In International Conference on Learning Representations, Cited by: §2.2.
  • Zhang et al. (2025) G. Zhang, M. Fu, G. Wan, M. Yu, K. Wang, and S. Yan G-memory: tracing hierarchical memory for multi-agent systems. arXiv preprint arXiv:2506.07398. Cited by: §2.2.
  • Zhang et al. (2026) W. Zhang, P. Zhang, J. Guo, T. Cheng, J. Chen, S. Zhang, Z. Zhang, Y. Yi, and H. Bu Patho-r1: a multimodal reinforcement learning-based pathology expert reasoner. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 40, pp. 28418–28426. Cited by: §C.2, Table 1, Table 3, §4.

Appendix A Prompts

A.1 Hierachical Evidence Bank Construction Prompts

This section lists the complete prompts used by the offline bank-construction agent (Sec. 3.2). All prompts are issued to the same frozen pathology VLM backbone that is used for online reasoning. Decoding is greedy with at most 128 new tokens. The VLM infers three times during the construction: anchor localization inside a retained tile (P1), view selection inside an anchor context (P2 and P3), and multi-scale view description (P4).

P1: anchor localization within a tile.

The tile is resized to 256×256256{\times}256 and partitioned into a 4×44{\times}4 grid. The agent scores all 16 cells and returns the top {target_count} cells, which is 3 for tumor-like tiles and 2 otherwise. P1 shows the anchor localization prompt.

P1: Anchor Localization Prompt You are a pathology diagnostic agent performing focused anchor planning. Input image is one low-resolution tile (256x256) from a WSI. Split it into a 4x4 grid (rows 0..3, cols 0..3). Select TOP-{target_count} cells for anchor extraction. Prioritize diagnostically rich structure: tumor-like atypia, hypercellularity, necrosis interface, and gland/duct architecture. Evaluate all 16 cells before choosing. Use ZERO-BASED indices only: row,col must be in [0,1,2,3]. Do NOT default to corner cells (especially R0C0) unless there is clear strongest evidence there. If multiple cells are close in quality, prefer the one closer to the tile center over corner cells. Return EXACTLY one JSON object and nothing else. Output format requirements: 1) Key: selected_cells 2) selected_cells length MUST be exactly {target_count} 3) Each item must include: row (int 0..3), col (int 0..3), confidence (float 0..1), rationale (short text) 4) selected_cells must contain unique (row,col) only Do NOT copy template text from this prompt. Do NOT output placeholder values such as "...", "TBD", or dummy coordinates. Coordinates must come from visual evidence in the image. Rationale must mention at least one concrete visual cue (e.g., pleomorphism, nuclear crowding, necrosis, tubule/gland structure). No tool_trace key. No markdown. No explanation text. tile_id={tile_id}; coarse_label_hint={coarse_label}; tissue_ratio={tissue_ratio}; target_count={target_count}. Keep rationale short.

P2: fine-detail view selection within an anchor.

The anchor region is resized to 256×256256{\times}256 and again partitioned into a 4×44{\times}4 grid. The agent first scores all 16 cells with a quality-control label and morphology tags, and then commits to {target_count} cells that are extracted as 512-pixel level-0 crops. The tag vocabulary is the image-side entry point of the signal ontology. We show the example of the breast used for WSI-VQA and BCNB in P2. The corresponding tag list can be replaced by the ontology of the target cohort for other organs.

P2: Fine-Detail View Selection Prompt (20×20\times) You are a pathology diagnostic agent selecting evidence-driven patches. Input: 4096x4096 level-0 context resized to 256x256. Split into 4x4 grid (row/col 0..3). Select TOP-{target_count} cells for extraction. Priority is morphology evidence, not geometry. Do NOT optimize for diagonals/symmetry. Spatial diversity is ONLY a tie-breaker when evidence is similar. Focus signals (summarized): - High-grade invasive tumor cues: marked nuclear pleomorphism, hypercellular dark hotspots (mitoses proxy), solid growth / low tubules. - DCIS cues: duct-centered intraductal proliferation; solid/cribriform patterns; comedo-type central necrosis. - Necrosis and calcifications: pale granular debris/ghost areas; tiny bright punctate clusters. - LVI / node mets: tumor clusters within vessel-like spaces or in nodal tissue (select only if clearly suggested). - Margin involvement/close: tumor abutting/near tissue edge (only if tumor is visibly near boundary). Return EXACTLY one JSON object. Keys: 1) cell_scores (len 16): each {row,col,qc:"ok"|"blur"|"artifact", score:0..1,tags:[...],note} - tags choose from: pleomorphism, mitosis_hotspot, low_tubules_solid, dcis, dcis_comedo, dcis_cribriform_solid, lvi_suspect, microcalc, necrosis, margin_close, margin_involved, node_met, micro_met - note: <=12 words, visible cue only (no diagnosis). 2) selected_cells (len {target_count}): each {row,col,conf:0..1,primary_tag,rationale} - rationale: <=12 words, include primary_tag + one visible cue. Rules: - unique (row,col). Avoid qc=artifact unless insufficient. - do not select score<0.35 unless not enough cells >=0.35. - rank by score. tie (diff<0.08): prefer qc ok, then cover different rows/cols. - never pick solely for spatial spread. tile_id={tile_id}; anchor_id={anchor_id}; coarse_label_hint={coarse_label}; target_count={target_count}.

P3: architectural context view selection.

Each anchor additionally receives one independently selected low-magnification view (10×10\times). Selection is deliberately decoupled from P2 so that the context view is not forced to coincide with the cell-level evidence, and the prompt asks for architectural cues such as transition fronts, necrosis interfaces, and ductal layout.

P3: Architectural Context View Selection Prompt (10×10\times) You are a pathology diagnostic agent selecting low-magnification context patches. Input: anchor context resized to 256x256 and split into 4x4 grid (row/col 0..3). Select TOP-{target_count} cells for ˜10x context extraction. Goal: - Capture broader tissue architecture and lesion context, independently from high-magnification picks. - Prioritize representative context zones (transition fronts, heterogeneous structure, necrosis interface, ductal layout). - Avoid selecting solely by corners/diagonals/ symmetry. Return EXACTLY one JSON object. Keys: 1) selected_cells (len {target_count}): each {row,col,confidence,rationale} 2) rationale must mention low-mag context cue (architecture/interface/distribution) in <=12 words. Rules: - Use unique (row,col). - Evaluate all 16 cells before selecting. - Do NOT copy text from prompt. - No markdown, no tool_trace, no extra keys. tile_id={tile_id}; anchor_id={anchor_id}; coarse_label_hint={coarse_label}; target_count={target_count}.

P4: multi-scale view description.

Every extracted view is described independently, producing the triple (dv,cv,zv)(d_{v},c_{v},z_{v}) of Sec. 3.2. The prompt constrains the model to observable morphology and to a fixed JSON schema, and it states the specimen organ {organ} of the cohort under evaluation (e.g., breast for WSI-VQA and BCNB) to suppress the organ drift that pathology VLMs exhibit on isolated fields. Diagnostic conclusions, grading terms, and management statements are excluded at the prompt level and again removed by the post-processing rules described below, so that the bank stores observations.

P4: Multi-Scale View Description Prompt You are analyzing ONE {magnification} H&E patch from a human {organ} specimen. Describe only visible morphology; do not infer another organ site, diagnosis, subtype, grade, biomarkers, or treatment. Return exactly one JSON object with this schema and no other keys or text: {"caption":"2-3 morphology-only sentences", "confidence":0.75,"qc_flags":["none"]} caption: - 2-3 sentences of concise pathologist-style morphology description (<=100 words). - Start with concrete visible features: color, cellularity, nuclear features, architecture, stroma, necrosis, calcifications, edge. - Use only observable morphology (e.g., enlarged irregular nuclei, solid nests, duct-like lumen, central pale debris, bright puncta, tumor at edge). - No diagnostic summary. No grading terms. No signal names. No recommendations. - Any high-level inference must be explicitly tied to a stated visible feature. - If minimal tissue or no clear morphology -> brief description + qc_flag=’low_tissue’. confidence: - >=0.8 if multiple clear concrete features. - 0.4-0.7 if limited but interpretable morphology. - <=0.3 if sparse, blurred, or uncertain. qc_flags: [none, blur, out_of_focus, artifact, low_tissue]. tile_id={tile_id}; anchor_id={anchor_id}; view_id={view_id}.

A.2 Inference Prompts

Size-question specialization.

Tumor extent is the one question family that is answered from measurements stored in the global record. A short prompt (P5) decides whether the current question is a measurement question, and if it is, the question is answered from the thumbnail with the scale prior of the stored pyramid geometry (P6). Questions that are not measurement questions are ignored by this branch and follow the ordinary routing path.

P5: Size-Question Router Prompt You are a question router for pathology MCQ. Determine whether this is primarily a tumor size/dimension measurement question. Size-related means it asks about size, dimension, diameter, extent, or values in mm/cm. Return strict JSON only: { "is_size_question": true, "reason": "short reason" } Question: {question} Options: {options}
P6: Size Answering Prompt You are a pathology multiple-choice assistant. This question is size-related. Use the thumbnail image to choose one option. Scale prior: - Original WSI long side is about 50000 px. - Thumbnail long side is about 1000 px (about 50x compression). You may estimate lesion span on thumbnail then multiply by about 50 when needed. Return strict JSON only: { "answer": "A", "reasoning": "short reason" } Question: {question} Options: {options}

Branch answering.

Every retrieved item is evaluated independently by the same VLM with the level-specific prompt of P7 (a-c). The three variants differ only in which memory content accompanies the image: the global branch sees the slide thumbnail alone, the anchor branch additionally receives the stored anchor description, and the patch branch receives the stored view caption. The final answer is from the level-weighted consensus.

P7a: Global Branch Answering Prompt You are a pathology multiple-choice assistant. Given one pathology slide thumbnail image and one multiple-choice question, choose exactly one option. Respond with only one letter: A, B, C, or D. Question: {question} Options: {options} Answer:
P7b: Anchor Branch Answering Prompt You are a pathology multiple-choice assistant. Given one pathology anchor-region image, its stored anchor description, and one multiple-choice question, choose exactly one option. Respond with only one letter: A, B, C, or D. Anchor description: {anchor_description} Question: {question} Options: {options} Answer:
P7c: Patch Branch Answering Prompt You are a pathology multiple-choice assistant. Given one pathology patch image, its stored patch description, and one multiple-choice question, choose exactly one option. Respond with only one letter: A, B, C, or D. Patch description: {patch_description} Question: {question} Options: {options} Answer:

Multi-turn context.

For multi-turn evaluation, the bounded history ℋt\mathcal{H}_{t} is rendered as a plain block that is prepended to the current question before the branch prompts are built, so that each branch is conditioned on the same discourse context while the visual evidence is retrieved afresh from the persistent bank. The history contains predicted answers, not ground truth, and is truncated to the most recent turns for the same slide.

P8: Multi-Turn Context Block
Previous QA context for this slide:
Q1: {question_1}
A1: {predicted_answer_1}
...
Qh: {question_h}
Ah: {predicted_answer_h}

Current question:
{question}

Appendix B Pathology Concept and Signal Ontology

SlideBank uses a compact pathology ontology to connect a clinical question with relevant visual evidence. A concept describes the clinical intent of a question (e.g., tumor grade or lymphovascular invasion), whereas a signal describes an atomic morphologic finding that can be assessed in a local image region (e.g., high mitotic activity). The resulting reasoning pipeline is

q⟶c⁡(q)⟶𝒮c⁡(q)⟶ℰq⟶a^,q\longrightarrow c(q)\longrightarrow\mathcal{S}_{c(q)}\longrightarrow\mathcal{E}_{q}\longrightarrow\hat{a},

where a question qq is first assigned to a concept c⁡(q)c(q). The concept selects a set of relevant signals 𝒮c⁡(q)\mathcal{S}_{c(q)}, which are then used to retrieve multi-scale evidence ℰq\mathcal{E}_{q} for answer generation. This separation makes the retrieval process interpretable: each answer can be traced from the question concept to the supporting morphologic signals and image regions.

B.1 Concept Vocabulary

The concept vocabulary covers common morphology-centered questions in breast pathology, together with metadata-oriented and morphology-based proxy tasks. Table 4 summarizes the concepts and their associated signals. The mapping is many-to-many because a single clinical concept may depend on several morphologic findings, and the same finding may support more than one concept. Concepts without a dedicated signal are answered using global and multi-scale regional evidence.

Table 4: Clinical concepts and their associated morphologic signals.
ID Concept Associated signals
M01 Tumor presence and distribution S104
M02 Histological tumor type (coarse) S103, S111, S113
M03 Tumor grade S101, S102, S103, S104
M04 Necrosis assessment S112, S132
M05 Inflammation / tumor-infiltrating lymphocytes No dedicated signal; multi-scale evidence
M06 In-situ component (DCIS-like) S111, S112, S113
M07 Margin status S141, S142
M08 Lymphovascular invasion S121
M09 Perineural invasion No dedicated signal; multi-scale evidence
M10 Microcalcifications S131
M11 Benign or non-neoplastic background changes No dedicated signal; multi-scale evidence
MD01 Specimen, procedure, and site metadata Global evidence
MD02 Tumor size, extent, and stage elements S141 (weak proxy), with global evidence
P01 Biomarker or outcome proxy prediction S101–S104, S111, S113, S132

B.2 Concept-to-signal Mapping

Signals are affirmative, locally assessable morphologic findings rather than diagnostic labels. Table 5 lists the signal vocabulary used by SlideBank. Each signal is linked to the image regions in which it is observed, allowing the same evidence to be reused across related concepts.

Table 5: Atomic morphologic signal vocabulary.
ID Signal Morphologic cue
S101 High nuclear pleomorphism Marked variation in nuclear size and shape.
S102 High mitotic activity Frequent mitotic figures indicating increased proliferation.
S103 Low tubule formation Reduced tubular differentiation with predominantly solid growth.
S104 Support for high tumor grade Combined visual evidence consistent with a higher-grade appearance.
S111 DCIS present In-situ ductal proliferation without clear stromal invasion.
S112 DCIS with comedo necrosis Intraluminal necrotic debris compatible with comedo morphology.
S113 Solid or cribriform DCIS architecture Solid or cribriform in-situ architecture within ductal units.
S121 Lymphovascular invasion present Suspicious tumor emboli within endothelial-lined spaces.
S131 Microcalcification present Fine basophilic calcific deposits in tumor-associated tissue.
S132 Tumor necrosis present Tumor-associated necrosis with loss of viable architecture.
S141 Tumor close to margin Tumor located close to a specimen edge in the sampled context.
S142 Tumor involving margin Tumor involving an edge or margin-like boundary.
S151 Lymph-node metastasis present Metastatic epithelial cells in a lymphoid background.
S152 Micrometastasis present A small metastatic focus compatible with micrometastatic burden.

For each evidence region, a signal is represented as present, unknown, or absent, together with a confidence score and its supporting views. During retrieval, evidence linked to the target signals is prioritized, while uncertain or negative observations may also be retained when they are relevant to the question.

B.3 Multi-scale Evidence Organization

SlideBank organizes evidence at three complementary scales: a global view captures slide-level context, anchor regions localize diagnostically informative areas, and patch views preserve detailed architectural or cellular morphology. Each signal is linked to its supporting anchor and patch views. Consequently, the evidence supplied to the vision-language model retains both its spatial origin and its morphologic interpretation, enabling an answer to be traced back to the corresponding signal and image region.

Appendix C Datasets and Baselines

C.1 Dataset Statistics

WSI-VQA.

The standard evaluation set contains 388 four-option questions associated with 85 WSIs. Questions cover slide-level diagnosis and morphology as well as attributes such as grade, size, receptor status, margins, and staging. We use the multiple-choice conversion of the released annotations and preserve the original slide–question association and option text. Accuracy is computed over questions, so slides with more annotated questions contribute proportionally more examples.

SlideBench-BCNB.

The BCNB split contains 1,058 breast-cancer WSIs. Following the standard SlideBench protocol, we evaluate the three tasks that are answerable from the released multiple-choice annotations: tumor type (1,058 questions), histological grading (926 questions), and molecular subtype (1,058 questions), for 3,042 questions in total. Tumor-type and grading questions have three answer candidates, while molecular-subtype questions have four. We report each task separately and compute the overall BCNB result as the question-count-weighted average

AccBCNB=13042​(CLOSE\displaystyle\operatorname{Acc}_{\mathrm{BCNB}}=\frac{1}{3042}\big( OPEN1058​Acctumor+926​Accgrade+1058​Accsubtype).\displaystyle 1058\operatorname{Acc}_{\mathrm{tumor}}+926\operatorname{Acc}_{\mathrm{grade}}+1058\operatorname{Acc}_{\mathrm{subtype}}\big). (14)

C.2 Baselines

All methods are evaluated on the same questions and answer candidates. Unless otherwise specified, we follow the released inference protocol for each baseline.

General-purpose and domain-specific VLMs. We evaluate Qwen2.5-VL-7B Bai et al. (2025b), Qwen3-VL-8B Bai et al. (2025a), and Qwen3.5-4B Qwen Team (2026) as general-purpose VLMs, together with LLaVA-Med Li et al. (2023), Quilt-LLaVA Seyfioglu et al. (2024), PathGen-LLaVA Sun et al. (2025), and Patho-R1 Zhang et al. (2026) as medical- or pathology-tuned VLMs. Each model receives a single ×10241024\!\times\!1024 WSI thumbnail and the multiple-choice question, without access to additional high-resolution fields.

WSI-LLaVA Liang et al. (2025). WSI-LLaVA is a multimodal model designed for gigapixel WSI understanding. It uses hierarchical slide representations and a multi-stage alignment and instruction-tuning strategy to connect WSI morphology with language.

TITAN Ding et al. (2025). TITAN is a multimodal whole-slide foundation model pretrained through visual self-supervision and vision–language alignment. It encodes an entire WSI into a slide-level representation that can be transferred to downstream pathology tasks without task-specific fine-tuning.

SlideChat Chen et al. (2025b). SlideChat is a vision–language assistant developed specifically for gigapixel whole-slide pathology. It is instruction-tuned on slide-level captions and VQA pairs to support WSI description and question answering.

MedAgents Tang et al. (2024). MedAgents is a multi-agent medical reasoning framework in which specialized agents analyze a question from complementary perspectives and aggregate their conclusions. We include it to assess whether inference-time collaboration can improve WSI question answering without a dedicated evidence-retrieval module.We adapt the MedAgents multi-agent collaboration framework to the multimodal setting using GPT-4o as the backbone. Each agent receives the WSI thumbnail and the multiple-choice question, and the final answer is obtained through multi-round specialist discussion and consensus.

PathAgent Chen et al. (2025a). PathAgent is a training-free agentic framework for WSI analysis. It iteratively navigates the slide, extracts morphological evidence at selected regions and magnifications, and integrates the observations to produce an interpretable answer.For the backbone setting, we use the official PathAgent implementation.

Appendix D Additional Implementation Details

Fusion weights. We use fixed fusion weights (λG,λA,λP)=(0.2,0.4,0.4)(\lambda_{G},\lambda_{A},\lambda_{P})=(0.2,0.4,0.4) across all datasets and backbones, without dataset-specific tuning.

Multi-turn evaluation. For each WSI-VQA question, Gemini-2.5-Flash generates five semantically equivalent rephrasings, and the answer options are independently shuffled while preserving the correct option text. After retaining complete six-form groups, the evaluation set contains 382 source questions and 2,292 turns from 80 WSIs, with an average of 29 turns per WSI. The evidence bank is constructed once per slide and remains fixed throughout the conversation; each turn uses the immediately preceding predicted question–answer pair as context, without access to ground-truth answers. Rephrase consistency is computed after mapping the shuffled option letters back to their semantic option text.

Answer parsing. We use a deterministic parser that first extracts a valid option letter from a JSON answer field or an explicit answer expression. If this fails, it accepts a standalone option letter or matches the normalized generated text to the provided option text. Outputs that cannot be resolved to a valid candidate are counted as incorrect, without using an additional language model for adjudication.