ComfyUI-FishAudioS2 is a specialized extension for ComfyUI that integrates the Fish Audio S2 Pro text-to-speech (TTS) model, known for its advanced capabilities in generating natural and emotive speech. This tool allows users to leverage state-of-the-art TTS technology with features like voice cloning and inline emotion control, enhancing the creative possibilities within ComfyUI.
- Zero-shot voice cloning enables users to replicate any voice using just a short reference audio clip.
- Inline emotion and prosody control allows for nuanced speech synthesis by using simple tag syntax for emotions and vocal styles.
- Multi-speaker synthesis facilitates the generation of conversations with multiple distinct voices in a single processing pass.
Context
The ComfyUI-FishAudioS2 extension serves as a bridge between the powerful Fish Audio S2 Pro TTS model and the ComfyUI framework. Its main purpose is to provide users with an easy-to-use, node-based interface for integrating advanced text-to-speech functionalities into their AI workflows, allowing for more expressive and varied audio generation.
Key Features & Benefits
This extension is packed with practical features that significantly enhance the TTS experience. The zero-shot voice cloning capability allows users to create unique voice profiles from brief audio samples, which can be particularly useful for applications requiring personalized audio outputs. Additionally, the extensive library of over 1500 emotive tags provides granular control over voice modulation, enabling users to convey specific emotions or vocal styles seamlessly. The multi-speaker synthesis feature allows for the creation of dialogues with distinct characters, making it ideal for interactive applications and storytelling.
Advanced Functionalities
ComfyUI-FishAudioS2 stands out with its advanced functionalities such as automatic language detection and support for 83 languages without the need for phoneme preprocessing. This broad language support, combined with the ability to synthesize multi-speaker conversations, makes it a robust tool for developers and content creators looking to produce diverse audio outputs. Furthermore, its optimized performance settings, including support for various precision types (bf16, fp16, fp32) and advanced attention mechanisms, ensure efficient processing on compatible hardware.
Practical Benefits
Incorporating ComfyUI-FishAudioS2 into a workflow enhances efficiency by streamlining the process of generating high-quality speech outputs. The ability to control emotions and voice characteristics inline allows creators to experiment with audio in a more intuitive manner. Moreover, the auto-download feature for model weights simplifies the setup process, ensuring users can quickly start generating audio without extensive configuration.
Credits/Acknowledgments
This project is developed by contributors from the Fish Audio team and is licensed under the Fish Audio Research License, permitting research and non-commercial use. For commercial applications, a separate license must be obtained from Fish Audio. The underlying model weights are sourced from the Hugging Face repository, ensuring access to state-of-the-art TTS technology.




