Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published 6 days ago • 28
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published 12 days ago • 75
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published 12 days ago • 75
StarVLA-$α$: Reducing Complexity in Vision-Language-Action Systems Paper • 2604.11757 • Published Apr 13
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published 12 days ago • 75
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published 12 days ago • 75
SkillOpt-Lite: Better and Faster Agent Self-evolution via One Line of Vibe Paper • 2607.03451 • Published about 1 month ago • 34
ParaVT: Taming the Tool Prior Paradox for Parallel Tool Use in Agentic Video Reinforcement Learning Paper • 2605.20342 • Published May 19 • 34
Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Paper • 2604.28185 • Published Apr 30 • 92
UniG2U-Bench: Do Unified Models Advance Multimodal Understanding? Paper • 2603.03241 • Published Mar 3 • 88
LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling Paper • 2511.20785 • Published Nov 25, 2025 • 188
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning Paper • 2510.13515 • Published Oct 15, 2025 • 12
OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe Paper • 2511.16334 • Published Nov 20, 2025 • 96
Complete Dictionary Learning via $\ell_p$-norm Maximization Paper • 2002.10043 • Published Feb 24, 2020
ImageBrush: Learning Visual In-Context Instructions for Exemplar-Based Image Manipulation Paper • 2308.00906 • Published Aug 2, 2023 • 15
Sparse Mixture-of-Experts are Domain Generalizable Learners Paper • 2206.04046 • Published Jun 8, 2022 • 1
CLIP-Mamba: CLIP Pretrained Mamba Models with OOD and Hessian Evaluation Paper • 2404.19394 • Published Apr 30, 2024
ALPINE: Unveiling the Planning Capability of Autoregressive Learning in Language Models Paper • 2405.09220 • Published May 15, 2024 • 27
Large Multi-modal Models Can Interpret Features in Large Multi-modal Models Paper • 2411.14982 • Published Nov 22, 2024 • 19