Code for the paper SimLoss: Single-Pass Fine-Grained Image Captioning via Embedding Space Distillation.
Modern vision-language models produce fluent high-level captions but routinely miss the attributes, counts, textures, materials, and spatial relations that make an image visually specific. Multi-stage systems (generate → decompose → verify → rewrite) recover some of these details at a substantial cost in inference latency. SimLoss is a reference-free embedding-space objective for single-pass fine-grained captioning: it trains a VLM to align its projected hidden-state representation with a frozen image embedding through an InfoNCE contrastive loss — no human-written fine-grained captions and no pseudo-captions from a multi-stage pipeline.
- SimLoss FFT backpropagates through a locally available embedding model (Qwen3-VL-Embedding-2B).
- SimLoss GRPO treats the embedding model (Gemini Embedding 2) as a black-box reward.
On IIW-400 with a Qwen2.5-VL-7B backbone, SimLoss FFT attains the highest precision of any method evaluated (0.849 vs. 0.788 for the unadapted backbone) at an F1 indistinguishable from the multi-stage CapMAS pipeline, while running roughly 20× faster.
| Folder | Method | What it is |
|---|---|---|
simloss/ |
SimLoss FFT | InfoNCE alignment of the captioner's projected, mean-pooled hidden states with frozen Qwen3-VL-Embedding image embeddings |
simloss/grpo_gemini/ |
SimLoss GRPO | GRPO with a Gemini-embedding cosine-similarity reward (black-box variant) |
capmas/ |
CapMAS baseline | Multi-agent captioner (5 captions → merge → decompose → fact-check → rewrite) + the Factuality / Coverage / CLAIR / DCScore evaluation metrics |
dcscore_rl/ |
FeedQuill-style GRPO/PPO baseline | RL fine-tuning with a DCScore (VLM-judge) reward |
papo/ |
PAPO, PAPO-YOLO baselines | Perception-aware RL; the YOLO variant masks detected object regions |
latency/ |
benchmark | Per-image inference latency for every method on IIW-400 |
All captioning policies are Qwen2.5-VL-7B-Instruct.
├── common/ shared code: config.py (single source of truth), utils, rewards,
│ judge_worker, dcscore_reward, latency_utils, prompts/
├── config/env.sh shell counterpart of common/config.py (sourced by every run_*.sh)
├── data/ small eval data committed; large data documented in data/README.md
├── models/ checkpoints — not committed; see models/README.md
├── capmas/ dcscore_rl/ papo/ simloss/ one folder per method
└── latency/ cross-method latency benchmark
Every method folder has the same shape: train and/or eval entrypoints (capmas is a training-free pipeline), run_*.sh SLURM scripts that source config/env.sh, and a README.md.
./setup.sh # builds .venv from requirements.txt (one pinned env for everything)Run scripts source config/env.sh, which sets PYTHONPATH, the base model, and data paths. The #SBATCH headers reflect the cluster the experiments ran on — adapt partition/GPU-constraint names to yours. Submit each script from its own directory after creating logs/ there (the SLURM log paths are submit-dir-relative):
(cd simloss && mkdir -p logs && sbatch run_train_encoder.sh)
(cd papo && mkdir -p logs && sbatch run_papo.sh)
(cd latency && mkdir -p logs && ./submit_all.sh)Credentials are read from the environment or git-ignored files:
export HF_TOKEN=... # gated HF models (meta-llama/*)
# CapMAS GPT-4o metrics: cp capmas/conf/gpt4o.example capmas/conf/gpt4o (add keys)
# SimLoss GRPO: cp simloss/grpo_gemini/.env.example simloss/grpo_gemini/.envEvaluation decoding is greedy (T = 0) for every method, with the per-method token budgets reported in the paper: plain baseline 512, CapMAS 512 per stage (10-token verification calls), PAPO variants 256, SimLoss FFT 768, SimLoss GRPO 500.
data/README.md lists what is committed (COVERAGE_TEST_VQA, IIW-400 ground-truth captions) and how to obtain the rest (IIW-400 images, COCO train2017, generated training captions, Gemini image embeddings).
Code under capmas/ is derived from Adobe Research's CapMAS and remains under the Adobe Research License (non-commercial research use); see capmas/UPSTREAM_README.md for the upstream citation.