TL;DR: FlowLong is a training-free, model-agnostic inference-time method that extends pretrained flow-based video diffusion models beyond their native generation horizon — works uniformly for text-to-video, audio-video joint, and text-to-3D scene generation.
Jangho Park*, Geon Yeong Park*, Gihyun Kwon†, Jong Chul Ye†.
KAIST, Amazon
- [2026.05.21] Our paper is now available on arXiv!
30-second audio-video joint generation from LTX-2. Click ▶ to play with sound.
🔊 Click the unmute icon to hear sound.
audiovideo_drummer_70.mp4
30-second video generation from HunyuanVideo-1.5.
video_kangaroo_70.mp4
Long text-to-3DGS generation by extending VIST3A (Wan 2.1-14B + AnySplat).
