Source code of paper "Exploring In-Image Machine Translation with Real-World Background" (ACL 2025 Findings) .
Download the IIMT30k dataset from
The DebackX model is composed of three main modules:
- Text-Image Background Separation Model
- Image Translation Model
- Text-Image Background Fusion Model
Follow the instructions below to train each module.
- Complete the configuration file:
./config/config-separate.json - Start training:
cd scripts
sh train-separate.sh- Complete the configuration:
./config/config-codebook.json - Start training:
cd scripts
sh train-codebook.sh- After codebook training, decode images to code sequences:
cd scripts
# Before decoding, make sure to:
# - Set the correct config and checkpoint paths
# - Set input_textimg_dir and output paths
sh decode-codebook.sh- Fill in the configuration file:
./config/config-translation.json - Train the translation model:
cd scripts
sh train-translation.sh📚 Optional: Pre-training for Translation
We provide code to construct synthetic text-images for pre-training.
-
Edit
build_text_img.py:- Replace font paths and parallel text paths.
-
Tokenize texts using SentencePiece:
spm_encode --model=./scripts/multi30k.model --output_format=piece --extra_options=bos:eos < path/to/texts > path/to/tokenized/texts/subtitle.tok.txt
spm_encode --model=./scripts/multi30k.model --output_format=id --extra_options=bos:eos < path/to/texts > path/to/tokenized/texts/subtitle.tok.id.txt-
Train on the synthetic data as in Step II.
-
Finetune on IIMT30k. In
config-translation.json, set"load_pretrain"to the pre-trained model path.
- Complete the configuration:
./config/config-fuse.json - Start training:
cd scripts
sh train-fuse.shAfter training all three models, generate the translated results:
cd scripts
# Before running, ensure all config and checkpoint paths are correct
# Set appropriate input/output directories
sh decode-separate.sh
sh decode-translation.sh
sh decode-fuse.shEvaluate the generated images using OCR (EasyOCR):
cd scripts
# Make sure to update `img_dir` and `result_file` in ocr.py
python ocr.py
