{"repo":"wildminder/ComfyUI-VibeVoice","free":true,"listed":false,"github":"https://github.com/wildminder/ComfyUI-VibeVoice","clone":"git clone https://github.com/wildminder/ComfyUI-VibeVoice.git","description":"ComfyUI custom node for the VibeVoice TTS. Expressive, long-form, multi-speaker conversational audio","language":"Python","stars":595,"topics":["audio","t2s","tts","ai-voice","text-to-speech","voice-generation","comfyui-custom-nodes-text-to-speech","comfyui-nodes","vibevoice","vibevoice-microsoft"],"license":"MIT","category":"media-processing","readme_excerpt":"ComfyUI-VibeVoice A custom node for ComfyUI that integrates Microsoft's VibeVoice, a frontier model for generating expressive, long-form, multi-speaker conversational audio. Report Bug · Request Feature [![Stargazers][stars-shield]][stars-url] [![Issues][issues-shield]][issues-url] [![Contributors][contributors-shield]][contributors-url] [![Forks][forks-shield]][forks-url] About The Project VibeVoice is a novel framework by Microsoft for generating expressive, long-form, multi-speaker conversational audio. It excels at creating natural-sounding dialogue, podcasts, and more, with consistent voices for up to 4 speakers. The custom node handles everything from model downloading and memory management to audio processing, allowing you to generate high-quality speech directly from a text script and reference audio files. ✨ Key Features: Multi-Speaker TTS: Generate conversations with up to 4 distinct voices in a single audio output. High-Fidelity Voice Cloning: Use any audio file ( .wav , .mp3 ) as a reference for a speaker's voice. Hybrid Generation Mode: Mix and match cloned voices with high-quality, zero-shot generated voices in the same script. Flexible Scripting: Use simple [1] tags or the classic Speaker 1: format to write your dialogue. Advanced Attention Mechanisms: Choose between eager , sdpa , flash attention 2 , and the high-performance sage attention for fine-tuned control over speed and compatibility. Robust 4-Bit Quantization: Run the large language model component in ","default_branch":null,"files":null,"tree":[],"storefront":"/r/wildminder","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/wildminder/ComfyUI-VibeVoice/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}