An MLLM-based retriever that moves KB-VQA retrieval from surface-level visual matching toward
entity-aligned semantic retrieval.
KBMR overview. EVA-CLIP-8B prefilters candidates, an MLLM Semantic Discriminator provides continuous entity-consistency weights, and Qwen2-VL-7B is trained with continuous semantic distillation.
- Overview
- Highlights
- Method
- Results
- Repository Structure
- Installation
- Data and Model Preparation
- Build Training Data
- Training
- Evaluation
- Acknowledgements
- Citation
Knowledge-Based Visual Question Answering (KB-VQA) requires retrieving external evidence for long-tail entities. Existing first-stage retrievers are commonly based on CLIP-style visual similarity. This is fragile when the same entity appears under different viewpoints, periods, or styles, and when different entities look visually similar.
KBMR replaces the first-stage CLIP retriever with an MLLM embedding retriever. Images are prompted with:
<Image> Summary above image in one word:
The final-token hidden state is normalized and used as the image embedding. An MLLM Semantic Discriminator then produces continuous entity-consistency weights for query-candidate pairs. These weights guide hard-negative construction and supervise the retriever through a symmetric KL distillation objective.
Motivation. Surface similarity can retrieve a visually similar but incorrect entity, whereas KBMR targets entity-level semantic relevance.
- The first MLLM-based embedding retriever tailored for KB-VQA, shifting retrieval from surface-level visual similarity toward entity-aligned semantic matching.
- Entity-aware supervision: An MLLM-based Semantic Discriminator generates continuous entity-consistency weights for reliable hard-negative mining and fine-grained soft supervision.
- Continuous Semantic Distillation: KBMR aligns retrieval similarities with an entity-aware semantic prior, improving discrimination among visually similar yet semantically distinct entities.
- Plug-and-play improvements: Replacing the original retriever with KBMR boosts diverse KB-VQA pipelines by up to +9.4 points in end-to-end answer accuracy, without modifying their downstream reasoning or reranking components.
- State-of-the-art performance: KBMR achieves up to +14.7 points in Recall@1 and VQA scores of 54.7, 50.8, and 79.3 on E-VQA, InfoSeek, and OK-VQA, respectively.
EVA-CLIP-8B encodes each query and knowledge-base image. After excluding candidates associated with the same Wikipedia entity, the 50 most similar images form the potential hard-negative set.
potential negatives = Top-50 EVA-CLIP cosine neighbors excluding positives
Qwen2.5-VL-7B receives a query image and candidate image with the following instruction:
You need to determine whether the given Candidate and the Query refer to
the same entity. If they do, answer 'Yes'; otherwise, answer 'No'.
The entity-consistency weight is computed from the Yes and No token logits:
w_i = exp(z_yes / gamma) / (exp(z_yes / gamma) + exp(z_no / gamma))
For positive weight w_pos, candidates with a weight above the margin threshold are removed:
alpha = w_pos - beta
keep candidate i when w_i <= alpha
This release uses beta = 0.01. Remaining candidates are sorted by their discriminator weights, divided into four equal-frequency difficulty strata, and sampled two per stratum. Samples are duplicated when fewer than eight remain. A query is discarded when no candidate survives filtering.
For the target candidate and eight hard negatives, KBMR builds:
- a retriever posterior from cosine similarities;
- a semantic prior from discriminator weights.
Both distributions use temperature-scaled softmax. Training minimizes symmetric KL divergence:
L_CSD = 0.5 * (KL(p_retriever || p_semantic) + KL(p_semantic || p_retriever))
The implementation is in kbmr/loss.py.
| Retriever | E-VQA | InfoSeek | ||||||
|---|---|---|---|---|---|---|---|---|
| R@1 ↑ | R@5 ↑ | R@10 ↑ | R@20 ↑ | R@1 ↑ | R@5 ↑ | R@10 ↑ | R@20 ↑ | |
| Qwen2-VL-7B (zero-shot) | 6.1 | 14.7 | 18.5 | 21.9 | 17.5 | 31.4 | 35.1 | 37.6 |
| EVA-CLIP-8B | 13.3 | 31.3 | 41.0 | 48.8 | 45.6 | 67.1 | 73.0 | 77.9 |
| KBMR (Qwen2-VL-7B) | 22.6 | 43.1 | 50.2 | 53.5 | 55.8 | 71.1 | 77.4 | 83.2 |
Entity retrieval performance (%). Higher is better; the best result in each column is shown in bold.
.
├── assets/
│ ├── method.png # Figure 2: complete KBMR pipeline
│ └── motivation.png # Figure 1: motivation
├── kbmr/
│ ├── data.py # Original split parsing and positive selection
│ ├── eva_clip.py # EVA-CLIP-8B image encoder adapter
│ ├── io.py # Streaming JSON-array I/O
│ ├── judge.py # Qwen2.5-VL Semantic Discriminator
│ ├── loss.py # Symmetric KL distillation objective
│ ├── modeling.py # Attention backend selection
│ └── sampling.py # Filtering and stratified sampling
├── tools/
│ ├── build_eva_index.py # Build the normalized FAISS corpus index
│ └── build_training_data.py # Mine, score, and sample training records
├── scripts/
│ └── train_qwen2vl_7b.sh # Qwen2-VL-7B training entry
├── tests/ # Data, sampling, and loss tests
├── train.py # LoRA + continuous distillation training
├── evaluate.py # Entity-level retrieval evaluation
└── pyproject.toml
git clone <your-repository-url> KBMR
cd KBMR
conda create -n kbmr python=3.10 -y
conda activate kbmr
pip install -e .To install the optional test dependencies, use pip install -e ".[dev]" instead.
Model downloads may require accepting the corresponding Hugging Face licenses.
| Component | Model | Purpose |
|---|---|---|
| Prefilter | EVA-CLIP-8B | Offline Top-50 candidate retrieval |
| Semantic Discriminator | Qwen/Qwen2.5-VL-7B-Instruct |
Entity-consistency weights |
| Trainable retriever | Qwen/Qwen2-VL-7B-Instruct |
Final KBMR retriever |
EVA-CLIP-8B checkpoints use different packaging conventions, so the code accepts a local model path or a compatible model-hub ID instead of assuming one distribution.
The data builder directly consumes the original post-split train.json; no intermediate query schema is introduced. The file is a top-level JSON array. KBMR uses wikipedia_url and related_images from each record.
The following is a field-level excerpt from the first record of the original file:
[
{
"wikipedia_url": "https://en.wikipedia.org/wiki/Heracleum_mantegazzianum",
"related_images": "iNaturalist_2021/train/06513_Plantae_Tracheophyta_Magnoliopsida_Apiales_Apiaceae_Heracleum_mantegazzianum/68bb3bbe-be28-4cdd-90a1-158af2d96b42.jpg"
}
]Other fields, including wikipedia_title, question, answer, retrieval, dataset_name, and unique_id, remain in the source split but are not used during hard-negative construction.
The original preprocessing behavior is preserved:
- Deduplicate globally by
related_images, keeping the first occurrence. - Group images by the exact
wikipedia_urlvalue. - Sort each entity's image paths lexicographically.
- Select the next image with wraparound as the target image.
- Use the same image as the target when an entity has only one image.
The source JSON is streamed with ijson; only the two required fields are retained in memory.
The knowledge-base metadata is another top-level JSON array. Each entry must contain at least:
{
"wikipedia_url": "https://en.wikipedia.org/wiki/Example_entity",
"image_url": "https://upload.wikimedia.org/path/to/image.jpg"
}A separate url2image.json object maps normalized image URLs to local relative paths:
{
"https://upload.wikimedia.org/path/to/image.jpg": "AToMiC-Images-v0.2/images/000/example.jpg"
}python -m tools.build_eva_index \
--metadata-json /path/to/knowledge_base/index.json \
--url-to-image-json /path/to/knowledge_base/url2image.json \
--image-root /path/to/knowledge_base/images \
--model /path/to/EVA-CLIP-8B \
--output-dir artifacts/eva_indexThis creates:
artifacts/eva_index/
├── index.faiss # Normalized EVA-CLIP embeddings with inner-product search
├── corpus.json # Index-aligned image paths and Wikipedia metadata
└── config.json # Encoder and metric metadata
python -m tools.build_training_data \
--train-json /path/to/original_split/train.json \
--query-image-root /path/to/query/images \
--candidate-image-root /path/to/knowledge_base/images \
--index-dir artifacts/eva_index \
--judge-model Qwen/Qwen2.5-VL-7B-Instruct \
--output data/kbmr_train.json \
--top-k 50 \
--beta 0.01 \
--gamma 1.1 \
--num-negatives 8 \
--image-size 336 \
--seed 42The output is a top-level JSON array compatible with the original training loader. The example below is abridged to one negative; every generated record contains exactly eight negatives and eight scores.
[
{
"query_text": "<|image_1|>\nRepresent the given image.\n",
"query_image": "/path/to/query.jpg",
"pos_text": "<|image_1|>\nRepresent the given image.\n",
"pos_image": "/path/to/positive.jpg",
"query_pos_scores": 0.95,
"hard_negatives": [
["<|image_1|>\nRepresent the given image.\n", "/path/to/negative.jpg"]
],
"hard_negatives_scores": [0.20]
}
]Edit the example paths at the top of scripts/train_qwen2vl_7b.sh, then run:
bash scripts/train_qwen2vl_7b.shThe script defaults to the paper's eight-GPU setup and keeps the accumulated query batch size at 1,024. On a four-GPU machine, run:
NUM_PROCESSES=4 bash scripts/train_qwen2vl_7b.shThe training and evaluation entry points use FlashAttention 2 when it is installed and otherwise fall back to PyTorch SDPA.
The evaluator consumes an original benchmark JSON array and the index-aligned corpus.json produced during EVA-CLIP indexing:
python evaluate.py \
--adapter /path/to/output/kbmr-qwen2vl-7b/checkpoint-5000 \
--index-dir artifacts/eva_index \
--queries /path/to/original_split/test.json \
--query-image-root /path/to/query/images \
--candidate-image-root /path/to/knowledge_base/images \
--recall-k 1 5 10 20The evaluator reports entity-level recall by matching normalized Wikipedia URLs. For paper comparisons, use the official E-VQA test split and the complete InfoSeek validation split with the 100K Wikipedia knowledge base. Downstream QA follows the original benchmark protocols: BEM for E-VQA and VQA accuracy for InfoSeek.
This implementation builds on the following open-source projects and models:
- Qwen2-VL and Qwen2.5-VL for the retriever and Semantic Discriminator.
- EVA-CLIP for offline candidate prefiltering.
- FAISS for nearest-neighbor search.
- Hugging Face Transformers, PEFT, Accelerate, and DeepSpeed for model training.
- Encyclopedic-VQA, InfoSeek, and OK-VQA for evaluation and training data.
Please follow the licenses and terms of each upstream model, dataset, and dependency.
If you find KBMR useful, please cite:
@misc{xu2026kbmr,
title={Beyond Visual Similarity: Entity-Aligned Retrieval for Knowledge-Based Visual Question Answering},
author={Hangrui Xu and Zhengxian Wu and Yunyao Yu and Zhuohong Chen and Rui Cong and Xiangwen Deng and Zhifang Liu and Peng Jiao and Haoqian Wang},
year={2026},
eprint={2608.21450},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2608.21450},
}