Buckets:
35.8 GB
15 files
Updated about 1 month ago
Ctrl+K
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| .gitattributes | 1.52 kB xet | 818ba6de | |
| README.md | 3.74 kB xet | 95abf896 | |
| campplus.onnx | 28.3 MB xet | db9e751b | |
| config.json | 6.37 kB xet | 6cf13bfc | |
| model-00001-of-00008.safetensors | 5 GB xet | 377aebf8 | |
| model-00002-of-00008.safetensors | 5 GB xet | a28aec1f | |
| model-00003-of-00008.safetensors | 5 GB xet | caa9521e | |
| model-00004-of-00008.safetensors | 5 GB xet | 0a06019a | |
| model-00005-of-00008.safetensors | 5 GB xet | 059f9091 | |
| model-00006-of-00008.safetensors | 5 GB xet | 1db2421c | |
| model-00007-of-00008.safetensors | 4.99 GB xet | 8d3c9598 | |
| model-00008-of-00008.safetensors | 820 MB xet | 2a63f83e | |
| model.safetensors.index.json | 635 kB xet | 494a0be5 | |
| tokenizer.json | 9.75 MB xet | ad147adb | |
| tokenizer_config.json | 54.9 kB xet | 0454394c |
๐Project Page ๏ฝ๐ค Hugging Face๏ฝ ๐ค ModelScope | ๐ฎ Gradio Demo-zh | ๐ฎ Gradio Demo-en | ๐ฌ DingTalk(้้)
Introduction
Ming-omni-tts is a high-performance unified audio generation model that achieves precise control over speech attributes and enables single-channel synthesis of speech, environmental sounds, and music. Powered by a custom 12.5Hz continuous tokenizer and Patch-by-Patch compression, it delivers competitive inference efficiency (3.1Hz). Additionally, the model features robust text normalization capabilities for the accurate and natural narration of complex mathematical and chemical expressions.
๐ Core Capabilities
- ๐ Fine-grained Vocal Control: The model supports precise control over speech rate, pitch, volume, emotion, and dialect through simple commands. Notably, its accuracy for Cantonese dialect control is as high as 93%, and its emotion control accuracy reaches 46.7%, surpassing CosyVoice3.
- ๐ Intelligent Voice Design: Features 100+ premium built-in voices and supports zero-shot voice design through natural language descriptions. Its performance on the Instruct-TTS-Eval-zh benchmark is on par with Qwen3-TTS.
- ๐ถ Immersive Unified Generation: The industryโs first autoregressive model to jointly generate speech, ambient sound, and music in a single channel. Built on a custom 12.5Hz continuous tokenizer and a DiT head architecture, it delivers a seamless, "in-the-scene" auditory experience.
- โก High-efficiency Inference: Introduces a "Patch-by-Patch" compression strategy that reduces the LLM inference frame rate to 3.1Hz. This significantly cuts latency and enables podcast-style audio generation while preserving naturalness and audio detail.
- ๐งช Professional Text Normalization: The model accurately parses and narrates complex formats, including mathematical expressions and chemical equations, ensuring natural-sounding output for specialized applications.
Evaluation
- Reconstruction: The 12Hz tokenizer supports high-quality reconstruction across speech, music, and sound. Its performance is comparable to existing state-of-the-art methods across key fidelity metrics.
- Dialect Generation: Achieves 96% accuracy on WSYue-TTS-Eval and 86% WSC-TTS-Eval, outperforming CosyVoice3.
- Emotional Expressiveness: Delivers an average accuracy of 76.7% on CV3-Eval emotional sets and 46.7% on neutral emotion sets, significantly surpassing CosyVoice3-Base (40%) to reach SOTA levels.
- Instruction-based Voice Design: Scores 76.20% on InstructTTS-Eval-ZH. Its instruction-following capability is on par with Qwen3-TTS-VoiceDesign.
- Zero-shot Voice Clone: Exhibits exceptional stability on Seed-tts-eval (Chinese) with a WER of 0.83%, outperforming SeedTTS and GLM-TTS.
- Text Normalization (TN): On internal technical testsets, the model achieves a CER of 1.97% in normalized regions, delivering performance comparable to Gemini-2.5 Pro.
Example Usage
git clone https://github.com/inclusionAI/Ming-omni-tts.git
cd Ming-omni-tts
python3 cookbooks/test.py
- Total size
- 35.8 GB
- Files
- 15
- Last updated
- Jun 16
- Pre-warmed CDN
- US EU US EU