OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Paper ⢠2607.23855 ⢠Published 3 days ago ⢠22
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Paper ⢠2607.23855 ⢠Published 3 days ago ⢠22
OpenMOSS-Team/MOSS-Transcribe-Diarize Audio-Text-to-Text ⢠0.9B ⢠Updated 5 days ago ⢠163k ⢠339
Running on Zero MCP 1 MOSS-VL-Instruct-0708 š§ 1 Image and video understanding with MOSS-VL multimodal model
Running on Zero Agents 3 MOSS-Music-8B-Instruct šµ 3 Analyze uploaded music and get detailed textual insights
Running on Zero Agents 3 MOSS-Music-8B-Instruct šµ 3 Analyze uploaded music and get detailed textual insights
OpenMOSS-Team/MOSS-VL-Instruct-0408 Video-Text-to-Text ⢠11B ⢠Updated 14 days ago ⢠14.6k ⢠101