IndexTTS2 · Voice Cloning With Emotion Control
Clone any voice and control emotional delivery with IndexTTS2, the open-weight TTS model. Upload a voice sample, type the text, and hit run. MIT license.
Audio
indextts2
text to speech
voice cloning
819
ABOUT THE WORKFLOW
Clone a Voice and Speak With Emotion Upload a short audio clip of the voice you want to clone. Type the text to be spoken. The model reproduces the voice and delivers the line with emotion inferred from the text. The result is saved as a WAV file. No training, no fine-tuning, no dataset collection.
Model
IndexTTS2 by the Index team (Bilibili). Paper published June 2025, open-sourced September 2025. A GPT-based autoregressive text-to-speech model trained on 55,000 hours of multilingual audio. Zero-shot voice cloning from a short reference clip (3 to 10 seconds) with emotion-identity decoupling: the model separates who speaks from how they speak. 92 percent or higher voice similarity on standard benchmarks. MIT license.
HOW IT WORKS
Step 1. Upload a voice sample A short audio clip (3 to 10 seconds) of the voice you want to clone. The model reads its tone, timbre, accent, and speaking style. Works great with: recordings · voiceovers · dialogue clips · podcast samples
Step 2. Type the text Write what you want spoken. Punctuation shapes the delivery: exclamation marks add energy, question marks add rising intonation, commas add pauses.
Step 3. Hit run and download The model clones the voice, infers emotion from the text, and saves the result as a WAV under tts2. Ready for: Premiere · DaVinci Resolve · Audacity · any audio editor
First time? Leave every setting as-is. The defaults (emotion weight 1 · WAV PCM16 · normalize off) are the right starting point for almost everyone.
RECOMMENDED SETTINGS
Quick-start guide. Find the goal that matches yours and copy the settings.
Standard voice clone (most people) — Emotion weight 1 · WAV PCM16 · normalize off. The right starting point for almost everyone.
The delivery is too flat — Add punctuation. Exclamation marks, question marks, ellipses, and dashes all steer how the model reads the line. "I can't believe it!" sounds different from "I cannot believe it."
The delivery is too intense — Lower the emotion weight below 1. At 0.5 the model delivers the line with less emotional variation. At 0 it reads the text neutrally.
Want to steer emotion from a separate clip — Connect an audio file to the emotion audio input. The model reads the emotional tone from that clip instead of inferring it from the text.
Want to use an emotion vector — Connect a precomputed emotion vector to the emotion vector input. This overrides both text-based and audio-based emotion inference.
The clone does not sound like the reference — Use a cleaner, longer reference clip. Background noise, music, or multiple speakers weaken the clone. 5 to 10 seconds of clean speech is the target.
Want MP3 instead of WAV — Switch the output format in the save node. MP3 at 320k is available.
Text: Write how you would speak. Punctuation is the emotion controller. "Wait... are you serious? That's incredible!" carries surprise, hesitation, and excitement. "That is incredible" carries nothing. The model reads punctuation as performance direction.
LEARN
📹 Videos
ComfyUI 101 Free Course ft. Sebastian Kamph
Floyo 101 for Team Collaboration
✨ Quick links
USE CASES
🎬 Film and Video Dubbing Clone a voice from existing footage and generate new lines in the same voice with matched emotion, without re-recording.
📖 Audiobook and Podcast Production Turn text into narrated audio in a cloned voice with natural pacing and emotional delivery.
🌍 Multilingual Voiceover Clone a speaker's voice and generate speech in a different language while preserving their timbre and style.
📢 Marketing and Product Narration Generate consistent brand voiceover across campaigns from a single reference clip.
WHAT WORKS BEST / WHAT TO AVOID
✅ Works great
Clean reference clips of 5 to 10 seconds with one speaker
Text with expressive punctuation (exclamation, questions, pauses)
Short to medium passages (one to three sentences)
Emotion weight at or near 1
⚠️ May produce softer results
Reference clips with background noise, music, or multiple speakers
Long passages where prosody can flatten over time
Text with no punctuation or emotional cues
Extremely short reference clips under 3 seconds
FAQ
What is IndexTTS2? IndexTTS2 is a GPT-based autoregressive text-to-speech model by the Index team at Bilibili, open-sourced in September 2025 under the MIT license. It clones any voice from a short reference clip (3 to 10 seconds) with no training or fine-tuning required. Its key innovation is emotion-identity decoupling: it separates who speaks from how they speak, so you can control the emotional delivery independently from the voice identity. It achieves 92 percent or higher voice similarity on standard benchmarks.
How does the emotion control work? By default the model infers emotion from the text: punctuation, phrasing, and word choice steer the delivery. You can also provide a separate emotion audio clip or a precomputed emotion vector to override the text-inferred emotion. The emotion weight slider (default 1) controls how strongly the emotional delivery is applied.
What languages does IndexTTS2 support? The model was trained on 55,000 hours of multilingual audio covering Chinese, English, and Japanese. It supports mixed character input and Pinyin pronunciation control for Chinese.
How long does the reference clip need to be? 3 to 10 seconds of clean speech from a single speaker. Longer clips give the model more to work with, but diminishing returns set in past 10 seconds. The quality of the clip matters more than the length: background noise, music, or overlapping voices weaken the clone.
Is IndexTTS2 free for commercial use? The sticky notes and community documentation list the MIT license, which allows commercial use, modification, and redistribution. However, the GitHub repository for the IndexTTS project references the Bilibili Model Use License Agreement, which may apply to certain versions. Verify the licence on the specific checkpoint before using it in commercial work.
How does IndexTTS2 compare to VibeVoice? Both are zero-shot voice cloning models. VibeVoice (Microsoft, MIT license) is a general-purpose TTS with strong naturalness. IndexTTS2 adds explicit emotion-identity decoupling with a controllable emotion weight, emotion audio input, and emotion vector input, giving finer control over the emotional delivery. Pick VibeVoice for straightforward voice cloning and IndexTTS2 when you need to steer how the line is delivered.
How to run IndexTTS2 online? You can run IndexTTS2 online through Floyo. No installation, no setup, no model downloads. Open the workflow in your browser, upload a voice sample, type the text, and hit run. Free to try.
WHY FLOYO?
Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.
A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.
For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.
Ready to try it? Upload a voice sample, type the text, and run it. The cloned voice speaks with emotion.
Questions? Watch the free course or check the FAQ above.
Read more






