All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation Paper • 2609.27901 • Published 11 days ago • 22
APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants Paper • 2609.37559 • Published 5 days ago • 41
Omni-IO Skills: Harnessing Your Agent Omni-Native Paper • 2609.31847 • Published 9 days ago • 285
YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality Paper • 2609.33757 • Published 7 days ago • 243
CoWindow Attention: Full Causal Coverage Is a Collective Property Paper • 2609.32704 • Published 8 days ago • 66
Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching Paper • 2608.09444 • Published 9 days ago • 12
DataoceanAI/Chinese_Female_Speech_Synthesis_Corpus_Live_Streaming_for_Sales Updated Jan 10, 2025 • 64 • 7
PierrunoYT/higgs-audio-v2-generation-3B-base Text-to-Speech • 6B • Updated Jul 28, 2025 • 49 • 4
ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds Paper • 2609.30199 • Published 10 days ago • 29
VladS159/common_voice_16_1_romanian_speech_synthesis Viewer • Updated Feb 26, 2024 • 39.1k • 104 • 5
Just-in-Time Memory: Learning to Curate Task-Adaptive Memory for LLM Agents Paper • 2609.27334 • Published 11 days ago • 54
szhengac25/higgs-audio-v2-generation-3B-base Text-to-Speech • 6B • Updated Aug 29, 2025 • 54 • 6
RULER: Instance-aware Rubric Rewards for SVG Generation Paper • 2609.25270 • Published 13 days ago • 102