{"repo":"OpenMOSS/MOSS-TTS","free":true,"listed":false,"github":"https://github.com/OpenMOSS/MOSS-TTS","clone":"git clone https://github.com/OpenMOSS/MOSS-TTS.git","description":"MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.","language":"Python","stars":4005,"topics":["audio","audio-tokenizer","llm","multimodal","text-to-speech","voice-cloning"],"license":"Apache-2.0","category":"media-processing","readme_excerpt":"MOSS-TTS Family &nbsp;&nbsp;&nbsp;&nbsp; English 简体中文 MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity , high‑expressiveness , and complex real‑world scenarios , covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS. News 2026.6.18: 🚀 MOSS-TTS-Local-Transformer-v1.5 receives Day-0 support in SGLang-Omni — the first inference backend to support the MossTTSLocal architecture, with an OpenAI-compatible /v1/audio/speech endpoint, streaming, and voice cloning. See the cookbooks: moss tts local , moss tts . 2026.6.18: 🚀 Released MOSS-TTS-Local-Transformer-v1.5, a 4B MossTTSLocal checkpoint that inherits all v1.5 features (language tags, stable cloning, explicit pause control, etc.), scales the backbone from Qwen3-1.7B to Qwen3-4B, and uses MOSS-Audio-Tokenizer-v2 for native 48 kHz stereo output. 2026.6.7: 🚀 Released MOSS-Audio-Tokenizer-v2, natively supporting 48 kHz stereo input and output. Check out the MOSS-Audio-Tokenizer repository for more details! 2026.6.2: 🚀 vLLM-Omni now supports the full MOSS-TTS series ( MossTTSDelay , MossTTSRealtime , and MossTTSNano architectures), including MOSS-TTS-v1.5, MOSS-TTS, MOSS-TTSD, MOSS-SoundEffect, MOSS-VoiceGenerator, MOSS-TTS-Realtime, and MOSS-TTS-Nano. See the recipe and examples. 2026.5.26: 🚀 Released MOSS-SoundEffect-v2.0, a new text-to-audio model us","default_branch":null,"files":null,"tree":[],"storefront":"/r/OpenMOSS","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/OpenMOSS/MOSS-TTS/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}