Advanced text-to-speech nodes for ComfyUI, FL CosyVoice3 leverages the CosyVoice3 model family to offer features like voice cloning without prior training, multilingual synthesis, and voice transformation capabilities.
- Zero-shot voice cloning enables users to replicate a voice using just a short audio sample.
- The tool supports multiple languages, allowing seamless speech generation while maintaining the original voice's nuances.
- It integrates with Whisper for automatic transcription and offers adjustable speech speeds for enhanced control over output.
Context
FL CosyVoice3 is a specialized extension for ComfyUI designed to enhance text-to-speech functionality. Its primary aim is to provide advanced voice synthesis capabilities, making it easier for users to create realistic speech outputs from text inputs.
Key Features & Benefits
This tool includes several practical features such as zero-shot voice cloning, which allows users to replicate voices from brief audio clips, and cross-lingual synthesis, enabling speech generation in various languages while keeping the original voice's characteristics intact. Additionally, it offers voice conversion, allowing users to change one voice to sound like another, thereby expanding creative possibilities in audio production.
Advanced Functionalities
FL CosyVoice3 includes nodes for advanced operations, such as the Instruct2 node, which allows users to clone voices based on specific instructions. The Save Speaker node enables users to save voice presets for future use, while the dialog synthesis feature can generate multi-speaker conversations with up to four different voices, enhancing the interactivity of audio outputs.
Practical Benefits
This tool significantly streamlines workflows within ComfyUI by providing high-quality voice synthesis options that are easy to implement. Users can efficiently create diverse audio outputs, maintain control over speech characteristics, and improve overall efficiency in audio production tasks.
Credits/Acknowledgments
The FL CosyVoice3 extension is based on the CosyVoice model family and is made available under the Apache 2.0 license. The original contributions come from the developers at FunAudioLLM and other collaborators involved in the project.




