Skip to content

Repository files navigation

DBT-Bleed: Dual-Branch Temporal Modeling with Key-Frame Selection for Surgical Bleeding Detection

MICCAI 2026

Sudhanshu Mishra1*, Jialang Xu3*, Jensen Ang4, Evangelos B. Mazomenos3, Beng Ti Christopher Ang4, Yueming Jin1,2†

1 Department of Electrical and Computer Engineering, National University of Singapore, Singapore
2 Department of Biomedical Engineering, National University of Singapore, Singapore
3 UCL Hawkes Institute, Department of Medical Physics and Biomedical Engineering, University College London, London, United Kingdom
4 Department of Neurosurgery, National Neuroscience Institute, Singapore

* Equal contribution  ·  † Corresponding author (ymjin@nus.edu.sg)

Abstract. Intraoperative Adverse Events (IAEs) detection is critical for improving surgical safety, with bleeding being among the most frequent events. Existing methods struggle to distinguish bleeding IAE from visually similar residual blood due to limited temporal reasoning, and modeling long surgical videos while preserving fine-grained temporal dynamics is computationally challenging. We propose DBT-Bleed, a dual-branch multi-scale temporal modeling framework that disentangles bleeding and normal representations using layer-wise temporal adapters for short- and long-term bleeding progression. To efficiently process long surgical videos, we introduce HiRED, a Hierarchical Entropy-Driven frame selection strategy that retains temporally informative segments while removing redundancy. On the MultiBypass dataset DBT-Bleed improves F1 by 6.53%, Recall by 5.62% and MCC by 9%, and demonstrates robust zero-shot cross-procedure transfer on EndoPit-IAE, a newly curated Endonasal Pituitary Surgery dataset — the first IAE-annotated dataset in neurosurgery.


DBT-Bleed framework: HiRED key-frame selection, dual-branch model, and multi-scale temporal adapter

Method

DBT-Bleed is a CLIP-based dual-branch framework with multi-scale temporal modeling. Two components are central:

  • HiRED key-frame selection — a Hierarchical Entropy-Driven strategy that scores frames by red-channel Shannon entropy and iteratively prunes redundant segments, keeping the top K informative frames from each N-frame window.
  • Multi-scale Temporal Adapter (MTA) — lightweight temporal transformers inserted at CLIP adapter depths {6, 12, 18, 24} to model both short- and long-range bleeding progression.

Datasets

  • MultiBypass (public) — laparoscopic bypass surgery videos with frame-level IAE annotations: CAMMA-public/MultiBypass140. Used for training and supervised evaluation.
  • EndoPit-IAE (in-house) — an Endonasal Pituitary Surgery dataset annotated for IAEs, used as an external zero-shot cross-procedure benchmark.

Pre-processing. Each surgical video is segmented into fixed-length clips of 300 frames with an overlap of 100 frames, and a clip is labelled positive if any of its frames contains bleeding.

Installation

Requires Linux with a CUDA-capable NVIDIA GPU (experiments used a single 24GB RTX A5000) and conda.

conda env create -f environment.yml
conda activate dbt_bleed

Configuration

Set your dataset/output paths by editing the --csv_dir, --output_dir, --exp_name (and --checkpoint for evaluation) flags inside scripts/train.sh and scripts/test.sh. No code changes are needed.

Pretrained weights

  1. CLIP backbone — download ViT-L-14-336px.pt and place it under CLIP/ckpt/.
  2. DBT-Bleed checkpoint — can be downloaded here and placed under checkpoints/.

Training

bash scripts/train.sh

Edit the flags in scripts/train.sh (see Configuration). Only the lightweight adapters and learnable text prompts are optimised; the CLIP backbone stays frozen.

Evaluation

bash scripts/test.sh

Edit the --checkpoint, --csv_dir, and --output_dir flags in scripts/test.sh (see Configuration); the model/sampling flags must match training.

Citation

If you find this work useful, please cite:

@misc{mishra2026dbtbleeddualbranchtemporalmodeling,
      title={DBT-Bleed: Dual-Branch Temporal Modeling with Key-Frame Selection for Surgical Bleeding Detection},
      author={Sudhanshu Mishra and Jialang Xu and Jensen Ang and Evangelos B. Mazomenos and Beng Ti Ang and Yueming Jin},
      year={2026},
      eprint={2606.22829},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2606.22829},
}

Acknowledgements

This repository builds on MadCLIP, MVFA-AD, and CoOp.

We also thank the authors of the baselines we compare against for releasing their code: SEDMamba, VadCLIP, ActionCLIP, and MadCLIP.

License

See LICENSE.

About

No description, website, or topics provided.

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages