Advanced voice cloning tool designed for ComfyUI, featuring over 80 emotional presets and the ability to mix multiple emotions for nuanced voice outputs. This tool enhances voice synthesis by allowing users to create emotionally rich audio that can convey a wide range of feelings and attitudes.
- Supports 80+ emotion presets, enabling users to select from a vast emotional palette for voice cloning.
- Offers multi-emotion mixing, allowing for complex emotional layers in generated speech.
- Features two voice cloning modes for either quick generation or high-quality output, catering to different user needs.
Context
This tool, known as the Qwen3-TTS Emotional Voice Clone, is an advanced extension for ComfyUI that specializes in voice cloning with a significant emotional range. Its primary purpose is to provide users with the ability to create voice outputs that are not only clear but also infused with various emotions, making it ideal for applications requiring expressive speech synthesis.
Key Features & Benefits
The tool includes over 80 emotion presets that cover a wide spectrum of feelings, from basic emotions like happiness and sadness to more complex states such as sarcasm or confidence. This extensive range allows for precise emotional expression, which is crucial for applications in storytelling, gaming, and interactive media. Additionally, the multi-emotion mixing feature enables users to blend emotions, thereby enhancing the depth and character of the generated voice outputs.
Advanced Functionalities
The tool provides two distinct voice cloning modes: a Fast Mode for rapid generation that focuses solely on speaker embedding, and an Accurate Mode that produces high-quality voice clones using a reference transcript. This flexibility allows users to choose between speed and fidelity based on their specific requirements. Moreover, users can control the intensity of emotions, adjusting the emotional strength from subtle to extreme, which adds another layer of customization to the voice outputs.
Practical Benefits
By integrating this tool into their workflow, users can significantly improve the expressiveness and emotional depth of their voice synthesis. This capability allows for more engaging and relatable audio outputs, enhancing user experience in applications such as virtual assistants, animated characters, and interactive storytelling. The ability to mix emotions also enables more realistic and dynamic character voices, which can lead to better audience engagement.
Credits/Acknowledgments
The Emotional Voice Clone Node was developed with contributions from Claude (Anthropic) and is based on the Qwen3-TTS framework created by the Alibaba Qwen Team. Integration into ComfyUI was facilitated by the contributions of flybirdxx, while the ComfyUI project itself is maintained by comfyanonymous. The tool is distributed under the GPL-3.0 License, aligning with the licensing of the base Qwen3-TTS package.




