SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue Paper • 2605.30993 • Published May 29 • 65
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published Jul 20 • 61
DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation Paper • 2606.26058 • Published Jun 24 • 67
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding Paper • 2411.17762 • Published Nov 26, 2024