meta-models/Muse-Glimmer-30B Image-Text-to-Text • 30B • Updated 6 days ago • 334k • • 1.65k
nvidia/nemotron-speech-streaming-en-0.6b Automatic Speech Recognition • 0.6B • Updated 12 days ago • 134k • 608
Running on Zero Agents Featured 175 AudioX 👀 175 Generate audio from text, video, or audio prompts
Running on Zero Agents Featured 5.08k MusicGen 🎵 5.08k Generate music from a text description and optional melody
view article Article ColFlor: Towards BERT-Size Vision-Language Document Retrieval Models ahmed-masry • Oct 18, 2024 • 22
Qwen/Qwen2-VL-2B-Instruct-GPTQ-Int4 Image-Text-to-Text • 2B • Updated Sep 21, 2024 • 755 • 28