[ECCV26] Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting
This repository packages the camera-ready artifacts for our ECCV26 work Identifying and Resolving Pitfalls of Knowledge-Based VQA Benchmarks: Auditing, Repairing, and Augmenting.
It contains:
- fixed and unfixed E-VQA / InfoSeek evaluation CSVs;
- augmented multi-entity query CSVs, with composite images hosted in the
HuggingFace dataset repo
VanQianMa/ECCV26-ARA; - reusable audit, repair, and augmentation tools under
fixing/; - one-command evaluation wrappers under
scripts/evaluation/; - method inference wrappers under
scripts/methods/, with code snapshots undermethods/. Feel free to direct to check the issues Qian has opened in original repos and contact with Qian.
Large original images, KB indexes, checkpoints, and full raw method outputs are
not duplicated in Git. They are materialized locally through
configs/paths.env and scripts/setup_local_assets.sh.
For the detailed settings to download images and KB indexes, you may refer to the instructions of EchoSight
We would like to thank the authors of EchoSight, ReflectiVA, CoMEM and Wiki-PRF for their released code and checkpoint.
Use one environment for evaluation and one for audit/repair/augmentation tools:
conda create -n KBVQA_eval python=3.10
conda activate KBVQA_eval
pip install -r requirements-evaluation.txtconda create -n KBVQA_fixing python=3.10
conda activate KBVQA_fixing
pip install -r requirements-fixing.txtMethod inference uses method-specific environments. Follow the corresponding original repositories or code snapshot READMEs for those dependencies:
| Method | Default wrapper environment |
|---|---|
| EchoSight | echosight |
| IBA | echosight |
| ReflectiVA | reflectiva |
| CoMEM | CoMEM |
| Wiki-PRF | echosight |
This only checks that the committed sample outputs can be parsed and scored. It does not reproduce the full paper tables.
conda activate KBVQA_eval
bash scripts/evaluation/run_smoke.shThe summary is written to
results/evaluation/smoke/infoseek_fixed/summary.json.
Full fixed/unfixed/augmented evaluation expects preserved raw method outputs
under outputs/raw_methods/, plus local KB/image/checkpoint paths. Configure
those paths once:
cp configs/paths.example.env configs/paths.env
$EDITOR configs/paths.env
bash scripts/setup_local_assets.shThen run all evaluation splits:
conda activate KBVQA_eval
bash scripts/evaluation/run_all_evaluations.shOutputs are written under results/evaluation/.
Useful variants:
MAX_SAMPLES=100 bash scripts/evaluation/run_all_evaluations.shMAX_SAMPLES is useful for a fast dry run. To reproduce the reported EVQA
scores, the BEM model must be available and loaded by the EVQA scorer.
These commands test method inference wrappers on one example and immediately score the generated output. They require the method-specific environments and local assets described above.
bash scripts/run_comem_infoseek_sample.shThis writes a one-example CoMEM prediction to
outputs/generated_methods/CoMEM/infoseek/fixed/ and scores it through
rag_evaluation/infoseek/score_fixed_infoseek_methods.py.
Run one real EchoSight retrieval/reranking sample:
bash scripts/run_echosight_evqa_reranker_sample.shThis writes outputs/generated_methods/EchoSight/evqa/fixed/sample_reranker_k3.json.
Set FORCE=1 to rerun after the sample output already exists.
Run one real IBA sample against the local Qwen vLLM server on port 8001, then score the generated answer:
bash scripts/run_iba_infoseek_sample.shThis writes sample metadata and answers under
outputs/generated_methods/IBA/infoseek/fixed/, then scores
sample_answers.jsonl directly.
Run one real Wiki-PRF sample, then score the generated details file:
bash scripts/run_wikiprf_infoseek_sample.shThis writes Wiki-PRF outputs under
outputs/generated_methods/Wiki-PRF/infoseek_fixed/, then scores
sample_results_infoseek_fixed_step600_topk3.json_0_generation_details.jsonl
directly. The default uses CUDA_VISIBLE_DEVICES=0,1; override
MODEL_GPU_ID, FILTER_GPU_ID, and RETRIEVER_GPU_ID if needed.
See data/README.md for the full layout.
Main files:
data/ground_truth/evqa_unfixed_test_with_id.csvdata/ground_truth/evqa_fixed.csvdata/ground_truth/infoseek_unfixed_subset.csvdata/ground_truth/infoseek_fixed.csvdata/augmented/evqa/evqa_challenging_queries_full_seed3185_with_images.csvdata/augmented/infoseek/infoseek_challenging_queries_full_seed3185_with_images.csv
Augmented composite images are intentionally not stored in Git. They are hosted
in the HuggingFace dataset repo VanQianMa/ECCV26-ARA under
images/augmented/, matching the relative paths in the augmented CSVs.
The reusable tools are under fixing/:
fixing/evqa/: E-VQA audit, question repair, evidence checking, final check, challenging-query generation, and composite-image construction.fixing/infoseek/: InfoSeek KB extraction, question/evidence/QA repair, final recheck, challenging-query generation, and composite-image construction.
Entry summaries:
sed -n '1,220p' fixing/README.md
sed -n '1,220p' fixing/evqa/README.md
sed -n '1,220p' fixing/infoseek/InfoSeek_Check_Plan.mdMost audit and repair scripts use standard CLI arguments and support small
--limit runs. LLM-based scripts need an OpenAI-compatible API key/base URL.
Use these wrappers to run the method code snapshots. Each supports --dry-run
to print the exact command before loading models. The environments should be
created from the corresponding original method repo or snapshot README; this
repository only standardizes the entrypoints and data paths.
| Method | Environment | Wrapper |
|---|---|---|
| EchoSight | echosight |
scripts/methods/run_echosight.sh --dataset evqa --split fixed --stage all |
| IBA | echosight |
scripts/methods/run_iba.sh --dataset evqa --split fixed --stage all |
| ReflectiVA | reflectiva |
scripts/methods/run_reflectiva.sh --dataset evqa --split fixed |
| CoMEM | CoMEM |
scripts/methods/run_comem.sh --dataset evqa --split fixed |
| Wiki-PRF | echosight |
scripts/methods/run_wikiprf.sh --dataset evqa --split fixed |
Set SKIP_CONDA_ACTIVATE=1 if the target environment is already active. Common
overrides such as CUDA_VISIBLE_DEVICES, OUTPUT_DIR, CHECKPOINT_PATH, and
model server URLs are supported by the underlying method scripts.
For IBA, the default OpenAI-compatible Qwen endpoint is
http://127.0.0.1:8001/v1, matching the local Qwen2.5-VL-7B-Instruct vLLM
server used for the runtime sample.
The paper evaluation covers five methods: EchoSight, IBA, CoMEM, ReflectiVA, and Wiki-PRF.
Individual scoring commands:
bash scripts/evaluation/run_evqa_fixed.sh
bash scripts/evaluation/run_evqa_unfixed.sh
bash scripts/evaluation/run_evqa_augmented.sh
bash scripts/evaluation/run_infoseek_fixed.sh
bash scripts/evaluation/run_infoseek_unfixed.sh
bash scripts/evaluation/run_infoseek_augmented.shEach scorer accepts --max-samples N and path overrides for ground truth or
method outputs. InfoSeek scorers use the answer-reward style rules in
rag_evaluation/infoseek/answer_reward_utils.py. EVQA scores reported in the
paper use BEM, so BEM_MODEL_DIR should be configured in configs/paths.env
before running scripts/setup_local_assets.sh.
For development only, EVQA scorers accept --allow-exact-match-fallback to
debug parsing and file wiring without loading BEM. Do not use this fallback to
reproduce paper results.
Augmented scoring evaluates the anchor, intra-category/method1, and
inter-category/method2 variants for IBA, EchoSight, and Wiki-PRF. Their raw
prediction files are expected under outputs/raw_methods/*/augmented/; use
scripts/setup_local_assets.sh to create local symlinks to the preserved full
outputs.
configs/paths.example.env is the publishable template. Copy it to
configs/paths.env, replace placeholders with machine-local roots, and run
scripts/setup_local_assets.sh. The local configs/paths.env file is ignored
by Git.
CoMEM and Wiki-PRF use their own local roots (COMEM_ROOT, WIKIPRF_ROOT) and
checkpoint directories. IBA preserved raw outputs are configured as explicit
files through IBA_EVQA_FIXED_OUTPUT, IBA_EVQA_UNFIXED_OUTPUT,
IBA_INFOSEEK_FIXED_OUTPUT, and IBA_INFOSEEK_UNFIXED_OUTPUT.
Readers do not need this for normal use. Before pushing a release package, we use:
bash scripts/check_camera_ready.shThe check verifies that required public files exist and that large local assets remain outside Git.