Adding image generation makes your LLM dumber (MoE/MoT collapse)📉. We fixed it!
🪨Rosetta (HKUST 🤝 Tencent Hunyuan Foundation Model Team) eliminates gradient conflicts to keep your LLM smart.
⚡️Strictly ZERO extra VRAM overhead.
💻Single-GPU pretraining reproducible in 90 mins!

