Skip to content

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

SimLoss: Single-Pass Fine-Grained Image Captioning via Embedding Space Distillation

Code for the paper SimLoss: Single-Pass Fine-Grained Image Captioning via Embedding Space Distillation.

Modern vision-language models produce fluent high-level captions but routinely miss the attributes, counts, textures, materials, and spatial relations that make an image visually specific. Multi-stage systems (generate → decompose → verify → rewrite) recover some of these details at a substantial cost in inference latency. SimLoss is a reference-free embedding-space objective for single-pass fine-grained captioning: it trains a VLM to align its projected hidden-state representation with a frozen image embedding through an InfoNCE contrastive loss — no human-written fine-grained captions and no pseudo-captions from a multi-stage pipeline.

  • SimLoss FFT backpropagates through a locally available embedding model (Qwen3-VL-Embedding-2B).
  • SimLoss GRPO treats the embedding model (Gemini Embedding 2) as a black-box reward.

On IIW-400 with a Qwen2.5-VL-7B backbone, SimLoss FFT attains the highest precision of any method evaluated (0.849 vs. 0.788 for the unadapted backbone) at an F1 indistinguishable from the multi-stage CapMAS pipeline, while running roughly 20× faster.

Methods

Folder Method What it is
simloss/ SimLoss FFT InfoNCE alignment of the captioner's projected, mean-pooled hidden states with frozen Qwen3-VL-Embedding image embeddings
simloss/grpo_gemini/ SimLoss GRPO GRPO with a Gemini-embedding cosine-similarity reward (black-box variant)
capmas/ CapMAS baseline Multi-agent captioner (5 captions → merge → decompose → fact-check → rewrite) + the Factuality / Coverage / CLAIR / DCScore evaluation metrics
dcscore_rl/ FeedQuill-style GRPO/PPO baseline RL fine-tuning with a DCScore (VLM-judge) reward
papo/ PAPO, PAPO-YOLO baselines Perception-aware RL; the YOLO variant masks detected object regions
latency/ benchmark Per-image inference latency for every method on IIW-400

All captioning policies are Qwen2.5-VL-7B-Instruct.

Layout

├── common/        shared code: config.py (single source of truth), utils, rewards,
│                  judge_worker, dcscore_reward, latency_utils, prompts/
├── config/env.sh  shell counterpart of common/config.py (sourced by every run_*.sh)
├── data/          small eval data committed; large data documented in data/README.md
├── models/        checkpoints — not committed; see models/README.md
├── capmas/  dcscore_rl/  papo/  simloss/   one folder per method
└── latency/       cross-method latency benchmark

Every method folder has the same shape: train and/or eval entrypoints (capmas is a training-free pipeline), run_*.sh SLURM scripts that source config/env.sh, and a README.md.

Setup

./setup.sh            # builds .venv from requirements.txt (one pinned env for everything)

Run scripts source config/env.sh, which sets PYTHONPATH, the base model, and data paths. The #SBATCH headers reflect the cluster the experiments ran on — adapt partition/GPU-constraint names to yours. Submit each script from its own directory after creating logs/ there (the SLURM log paths are submit-dir-relative):

(cd simloss && mkdir -p logs && sbatch run_train_encoder.sh)
(cd papo    && mkdir -p logs && sbatch run_papo.sh)
(cd latency && mkdir -p logs && ./submit_all.sh)

Credentials are read from the environment or git-ignored files:

export HF_TOKEN=...                    # gated HF models (meta-llama/*)
# CapMAS GPT-4o metrics: cp capmas/conf/gpt4o.example capmas/conf/gpt4o  (add keys)
# SimLoss GRPO:          cp simloss/grpo_gemini/.env.example simloss/grpo_gemini/.env

Decoding configuration

Evaluation decoding is greedy (T = 0) for every method, with the per-method token budgets reported in the paper: plain baseline 512, CapMAS 512 per stage (10-token verification calls), PAPO variants 256, SimLoss FFT 768, SimLoss GRPO 500.

Data

data/README.md lists what is committed (COVERAGE_TEST_VQA, IIW-400 ground-truth captions) and how to obtain the rest (IIW-400 images, COCO train2017, generated training captions, Gemini image embeddings).

License

Code under capmas/ is derived from Adobe Research's CapMAS and remains under the Adobe Research License (non-commercial research use); see capmas/UPSTREAM_README.md for the upstream citation.

About

No description, website, or topics provided.

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages