MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities
Paper β’ 2607.25948 β’ Published β’ 14
MODUS for any-to-any generation (15 aligned modalities) as presented in MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities.
Project page: https://modus-multimodal.epfl.ch/ Code: https://github.com/EPFL-VILAB/Modus
Inference config (important):
conf/modalities/instruction_16mod_stage2.yamlmodel.safetensors β trained weights (bf16)ae.safetensors β VAE (image decode)config.json / llm_config.json / vit_config.json β architecture configvocab.json / merges.txt / tokenizer_config.json β tokenizer