VoxCPM2 for Voice Cloning
Upload a short voice sample and type what you want it to say. VoxCPM2 clones the voice and generates new speech in 30 languages at 48 kHz studio quality.
audio
multilingual
open source
text to speech
tts
voice clone
voxcpm2
0
63
Nodes & Models
LoadAudio
PreviewAudio
FloyoStickyNote
VoxCPM2_Clone
ABOUT THE WORKFLOW
Clone a Voice from Audio Upload a short voice clip and type the words you want spoken. VoxCPM2 reads the voice from your clip and generates new speech that sounds like the same speaker. Add a voice description to steer emotion or pace on top of the clone.
Open source, so you only pay for generation time. All models pre-loaded.
Model
VoxCPM2 by OpenBMB (Tsinghua University / ModelBest). A 2B-parameter tokenizer-free TTS model that clones voices from short reference clips across 30 languages, outputting 48 kHz audio. Apache 2.0 license.
HOW IT WORKS
Step 1. Upload a voice sample A short clip of the voice you want to clone. Clean recordings with a single speaker work best. Three to ten seconds is enough. Works great with: podcast clips · voiceovers · interview recordings · voice memos
Step 2. Type your text Write the words you want spoken in the cloned voice. Any language VoxCPM2 supports works here.
Step 3. Add a voice description (optional) A short phrase to steer the style, emotion, or pace on top of the cloned voice. Example: "warm and calm, slow pace."
Step 4. Add the reference transcript (optional) Type out what the speaker says in your reference clip. Filling this in activates a stronger cloning mode and improves voice match.
Step 5. Hit run and preview VoxCPM2 generates the cloned speech. Preview it on canvas, then download the audio file. Ready for: podcast editors · video timelines · game engines · e-learning platforms
First time? Leave every setting as-is. The defaults (CFG 2 · 10 timesteps · 1024 max tokens) are the right starting point for almost everyone.
RECOMMENDED SETTINGS
Quick-start guide. Find the goal that matches yours and copy the settings.
Standard voice clone (most people) — CFG 2 · 10 timesteps · 1024 max tokens · denoiser off. The right starting point for almost everyone.
Longer speech output — Raise max tokens above 1024. At the default you get roughly 15 to 20 seconds of speech. Higher values produce longer clips, but quality may drift past 30 seconds.
Stronger voice match — Fill in the prompt text field with a transcript of your reference clip. This activates a tighter cloning mode that picks up more of the speaker's tone and rhythm.
Cleaner output from a noisy reference — Turn enable denoiser on. It cleans background noise in the generated audio.
Different take on the same text — Change the seed to any new number. Keep the seed locked to reproduce an output you liked.
Higher quality, slower generation — Raise inference timesteps above 10. More passes add detail to the speech at the cost of speed.
The clone sounds flat or off — Use a cleaner reference clip with minimal background noise and a single speaker. Clips with music, reverb, or overlapping voices degrade the match.
Prompt: Write the text you want spoken in full sentences, with natural punctuation. For the voice description field, keep it to a short phrase describing the tone and pace you want, like "confident, upbeat, medium pace" rather than a long paragraph.
LEARN
📹 Videos
ComfyUI 101 Free Course ft. Sebastian Kamph
Floyo 101 for Team Collaboration
✨ Quick links
USE CASES
🎙️ Podcast & Narration Clone a narrator's voice and generate additional lines without booking another recording session. Useful for pickups, corrections, or extending a read.
🌍 Multilingual Voiceover Take a speaker's English voice and generate speech in other languages while keeping the same vocal identity. VoxCPM2 supports 30 languages.
🎮 Game & Interactive Media Generate character dialogue in a consistent voice from a single reference clip. Prototype voice lines before committing to full studio sessions.
📚 E-Learning & Accessibility Produce narrated course content in a consistent instructor voice across lessons, or convert written materials to spoken audio for accessibility.
WHAT WORKS BEST / WHAT TO AVOID
✅ Works great
Clean, single-speaker reference clips
3 to 10 second reference recordings
Well-punctuated input text
Short to medium outputs (under 30 seconds)
⚠️ May produce softer results
References with background music or reverb
Multiple speakers in the reference clip
Long outputs past 30 seconds (voice may drift)
Low-quality or heavily compressed reference audio
FAQ
What is VoxCPM2 and who made it? VoxCPM2 is a 2-billion-parameter text-to-speech model made by OpenBMB, a research lab backed by Tsinghua University and ModelBest. It was released in April 2026. The model generates speech from text, clones voices from short audio clips, and supports 30 languages with 48 kHz output. It is open source under the Apache 2.0 license.
How long does the reference audio clip need to be for voice cloning? Three to ten seconds of clean, single-speaker audio is the target range. Shorter clips give the model less to work with, and longer clips do not add much benefit. What matters more than length is recording quality: a clear clip with no background noise or music produces a better match than a longer, noisy one.
Does VoxCPM2 support languages other than English? Yes. VoxCPM2 supports 30 languages. You can clone a voice from a clip in one language and have it speak in another, or generate speech natively in any of the supported languages. The model was trained on over 2 million hours of multilingual speech data.
What is the difference between voice cloning and voice design in VoxCPM2? Voice cloning starts from a reference audio clip and replicates that speaker's voice. Voice design creates a new voice from a text description alone, with no reference audio needed. This workflow uses the cloning mode: you supply a voice sample, and the model speaks your text in that voice.
Can I use VoxCPM2 voice clones for commercial projects? Yes. VoxCPM2 is licensed under Apache 2.0, which permits commercial use. Make sure you have the right to use the voice you are cloning, and label AI-generated speech as synthetic where required. The model should not be used for impersonation, fraud, or disinformation.
How does VoxCPM2 compare to ElevenLabs for voice cloning? VoxCPM2 scored 85.4% on English voice similarity in the Minimax-MLS benchmark, compared to 61.3% for ElevenLabs. The tradeoff: ElevenLabs is a managed service with uptime guarantees, enterprise support, and a polished API. VoxCPM2 is free, open source, and self-hostable, but you manage infrastructure yourself or run it through a platform like this one.
How to run VoxCPM2 voice cloning online? You can run VoxCPM2 voice cloning online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, upload a voice clip, type your text, and hit run. Free to try.
WHY FLOYO?
Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.
A producer runs a voice clone and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.
For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.
Ready to try it? Upload a voice clip, type the text you want spoken, and hit run.
Questions? Watch the free course or check the FAQ above.
Read more







