VibeVoice · Text to Speech
Clone any voice from a short clip and read your text in it using VibeVoice Large by Microsoft. Upload a voice sample, type your script, hit run. MIT license.
Audio
text to speech
TTS
VibeVoice
voice cloning
7
1.3k
Nodes & Models
LoadTextFromFileNode
VibeVoiceSingleSpeakerNode
LoadAudio
PreviewAudio
ABOUT THE WORKFLOW
Read Text in a Cloned Voice Upload a short audio clip of the voice you want, type the text you want spoken, and hit run. VibeVoice clones the voice from your sample and reads your script in it. The result plays back in the workflow as a preview.
Model
VibeVoice Large by Microsoft. Open-sourced August 2025 under the MIT License. A text-to-speech model with voice cloning that reads from a short reference sample and generates speech matching its tone, pacing, and character. Built on a next-token diffusion framework with continuous speech tokenizers at 7.5 Hz. Supports English and Mandarin Chinese.
HOW IT WORKS
Step 1. Upload a voice sample A short audio clip of the voice you want cloned, around 10 seconds. Clear speech with minimal background noise produces the closest match. Works great with: voice memos · interview clips · podcast samples · narration recordings
Step 2. Type your script Write the text you want spoken into the node. Punctuation controls pacing. Line breaks add pauses.
Step 3. Hit run and listen The model clones the voice from your sample and reads your script in it. The result plays in the audio preview.
Step 4. Download Right-click the audio preview to save the file. There is no automatic save step. Ready for: Premiere · DaVinci Resolve · Audacity · any audio editor
First time? Leave every setting as-is. The defaults (VibeVoice-Large · 20 diffusion steps · CFG 1.3 · fixed seed) are the right starting point for almost everyone.
RECOMMENDED SETTINGS
Quick-start guide. Find the goal that matches yours and copy the settings.
Standard voice clone (most people) — VibeVoice-Large · 20 diffusion steps · CFG 1.3 · fixed seed. The right starting point for almost everyone.
Voice sounds robotic or flat — Try a cleaner reference clip. Background noise, music, or multiple speakers in the sample dilute the clone.
Voice drifts from the reference — Raise CFG toward 2. Higher values hold the output closer to the reference voice, but push too far and the speech sounds forced.
Pacing is too fast or too slow — Adjust voice speed. Lower than 1 slows it down, higher speeds it up.
Long text has audible seams — Text over 250 words is split into chunks. Seams between chunks can be audible. Shorten the text or place natural pauses (periods, line breaks) where the split might land.
Smoother output — Raise diffusion steps above 20. More passes refine the audio at the cost of generation time.
Want a different read — Change the seed number. The default runs on fixed, so the same sample, text, and seed return the same output.
Text from a file — A LoadTextFromFile node sits muted by default. Enable it and point it at a text file instead of typing into the node.
Script: Punctuation is your pacing tool. Periods create full stops. Commas create short pauses. Line breaks add longer pauses. "Hello there. (pause) Let's see how this sounds." reads differently from "Hello there, let's see how this sounds." Write the script the way you want it heard.
LEARN
📹 Videos
ComfyUI 101 Free Course ft. Sebastian Kamph
Floyo 101 for Team Collaboration
✨ Quick links
USE CASES
🎙️ Narration and Voiceover Clone a narrator's voice from a short sample and generate reads for explainer videos, courses, or audiobooks.
🎧 Podcast and Dialogue Read scripted dialogue in a specific voice for podcast pilots, demo episodes, or prototyping.
📢 Ads and Product Audio Generate a branded voice read for a product spot without booking a studio session.
🌍 Multilingual Content Read the same script in English and Mandarin Chinese in the same cloned voice for bilingual content.
WHAT WORKS BEST / WHAT TO AVOID
✅ Works great
Clean reference clips of 10 to 30 seconds with one speaker
Scripts with deliberate punctuation and line breaks for pacing
English and Mandarin Chinese text
Narration, voiceover, and dialogue reads
⚠️ May produce softer results
Reference clips with background music, noise, or multiple speakers
Scripts longer than 250 words in a single run (chunk seams appear)
Languages other than English and Mandarin
Expecting pixel-perfect voice matching on extreme vocal ranges
FAQ
What is VibeVoice? VibeVoice is a family of open-source voice AI models from Microsoft. VibeVoice Large, used in this workflow, is a text-to-speech model with voice cloning that reads from a short reference audio sample and generates speech matching its tone, pacing, and character. It was open-sourced in August 2025 and accepted as an Oral at ICLR 2026. It uses continuous speech tokenizers at 7.5 Hz and a next-token diffusion framework.
How long does the voice sample need to be? About 10 seconds of clear speech from a single speaker. Longer samples give the model more to work with, but diminishing returns set in past 30 seconds. Quality matters more than length: a clean 10-second clip outperforms a noisy 60-second one.
What languages does VibeVoice support? English and Mandarin Chinese, with cross-lingual generation that allows natural language switching within the same output. Other languages are not officially supported and produce lower quality results.
Is VibeVoice free for commercial use? The weights were released under the MIT License, which permits commercial use, modification, and redistribution. Microsoft's project page adds that VibeVoice is "intended for research and development purposes" and recommends further testing for commercial applications. The MIT licence is permissive, but read the responsible use guidance before deploying in production.
Why is there no save step? The workflow outputs to an audio preview node with no SaveAudio node wired after it. Right-click the preview to download the file. Adding a save node would make downloads automatic.
How long can the output be? VibeVoice Large can generate extended speech, but this workflow splits text into chunks of 250 words. Seams between chunks can be audible. For the cleanest output, keep each run under 250 words and join the clips in an editor.
How to run VibeVoice online? You can run VibeVoice online through Floyo. No installation, no setup, no model downloads. Open the workflow in your browser, upload a voice sample, type your script, and hit run. Free to try.
WHY FLOYO?
Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.
A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.
For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.
Ready to try it? Upload a voice sample, type your script, and run it. The settings are already set.
Questions? Watch the free course or check the FAQ above.
Read more







