VoxCPM2 for Text to Speech
Turn text into spoken audio with VoxCPM2. Describe the voice you want in plain language, type what it should say, and hit run. 30 languages, 48 kHz output.
audio
multilingual
text to speech
tts
voice cloning
voice design
voxcpm2
0
55
ABOUT THE WORKFLOW
Generate Speech from Text Type what you want said and describe the voice you want. VoxCPM2 generates a spoken audio clip that matches your description. Add a voice LoRA to clone a specific speaker. That's it.
Model
VoxCPM2 by OpenBMB (Tsinghua University / ModelBest). A 2B-parameter text-to-speech model trained on over 2 million hours of speech. Supports 30 languages, voice design from text descriptions, and 48 kHz studio-quality output. Apache 2.0 license.
HOW IT WORKS
Step 1. Write what you want said Type the text you want the voice to speak. Any language works. The model detects the language on its own, no tags needed. Works great with: dialogue · narration · voiceovers · announcements
Step 2. Describe the voice Optional. Write a short description of the voice you want: age, gender, tone, pace, emotion. Example: "A young woman with a calm, warm voice, speaking slowly."
Step 3. Hit run and preview VoxCPM2 generates the audio and plays it back in the preview node. Download the clip when you like what you hear. Ready for: video editors · podcasts · games · e-learning · apps
First time? Leave every setting as-is. The defaults (CFG 2 · 10 timesteps · 1024 max tokens) are the right starting point for almost everyone.
RECOMMENDED SETTINGS
Quick-start guide. Find the goal that matches yours and copy the settings.
Standard voiceover (most people) — CFG 2 · 10 timesteps · max tokens 1024 · fixed seed. The right starting point for almost everyone.
Want a specific voice style — Write a detailed voice description with age, gender, speed, emotion, and accent. Example: "An older British man with a deep, slow, authoritative voice." The more specific you are, the closer the result lands.
Need a longer speech clip — Raise max tokens above 1024. This caps the output length. For paragraphs of text, set it higher so the model does not cut off early.
Want higher quality audio — Raise inference timesteps to 15 or 20. More steps improve clarity and naturalness at the cost of generation time.
Cloning a specific speaker — Select a trained voice LoRA from the dropdown. This overrides the voice description and reproduces that speaker's sound.
The voice does not match the description — Try generating 2 to 3 times. Voice design results vary between runs. Change the seed each time for a different take.
Getting cut-off audio — Raise max tokens. The default of 1024 works for short to medium text. Longer passages need more headroom.
Prompt: The voice description steers style, not identity. Use concrete traits: "young, female, fast, excited" lands better than "a nice voice." For a specific person's voice, skip the description and use a LoRA instead.
LEARN
📹 Videos
ComfyUI 101 Free Course ft. Sebastian Kamph
Floyo 101 for Team Collaboration
✨ Quick links
USE CASES
🎬 Video Creators & Filmmakers Generate voiceovers for trailers, explainers, or short films without booking a voice actor. Describe the voice you need, type the script, and get a 48 kHz clip ready for your timeline.
🎮 Game Developers Prototype NPC dialogue and in-game narration during development. Test different character voices by changing the description, then swap in final recordings later or ship the generated audio directly.
📚 E-Learning & Audiobook Producers Narrate courses, tutorials, or book chapters in any of 30 supported languages. Consistent pacing and tone across long scripts, without scheduling recording sessions.
🌍 Multilingual Content Teams Produce the same script in multiple languages from a single workflow. No language tags needed. Type the text in the target language and VoxCPM2 detects it on its own.
WHAT WORKS BEST / WHAT TO AVOID
✅ Works great
Short to medium text (a few sentences to a paragraph)
Specific voice descriptions with age, gender, tone, and pace
Single-language passages
Clean, punctuated text
⚠️ May produce softer results
Long unbroken text (may hit the token limit)
Vague voice descriptions like "a nice voice"
Mixed-language text in the same passage
Rapid-fire dialogue with multiple speakers in one run
FAQ
What is VoxCPM2 and who made it? VoxCPM2 is an open-source text-to-speech model by OpenBMB, a research group backed by Tsinghua University and ModelBest. It has 2 billion parameters, supports 30 languages, and outputs 48 kHz audio. It is the successor to ChatTTS and VoxCPM 1.5, released under an Apache 2.0 license.
How does VoxCPM2 voice design work? You describe the voice you want in plain language: age, gender, tone, emotion, pace, accent. The model generates speech that matches that description without needing any reference audio. Results vary between runs, so generating a few times with different seeds helps you find the right fit.
What languages does VoxCPM2 support? VoxCPM2 supports 30 languages and 9 Chinese dialects. You do not need to tag or specify the language. Type your text in the target language and the model detects it on its own.
Can I clone a specific voice with VoxCPM2? Yes, through voice LoRAs. Select a trained voice LoRA from the dropdown in the workflow, and VoxCPM2 reproduces that speaker's sound. Without a LoRA, the model generates a voice based on your text description instead.
Is VoxCPM2 output good enough for production use? VoxCPM2 outputs at 48 kHz, which is broadcast and studio quality. For voiceovers, narration, and game dialogue, the quality holds up in production. For the closest match to a professional recording, raise inference timesteps and write a detailed voice description.
Can I use VoxCPM2 audio commercially? Yes. VoxCPM2 is released under the Apache 2.0 license, which allows free commercial use. You can use the generated audio in shipped products, client work, ads, courses, and published content without additional licensing.
How to run VoxCPM2 text to speech online? You can run VoxCPM2 text to speech online through Floyo. No installation, no setup, no API key to wire up. Open the workflow in your browser, type your text, describe the voice, and hit run. Free to try.
WHY FLOYO?
Floyo is the only platform with team collaboration for ComfyUI in the browser. You run workflows with no install. You share run history, assets, and models across your team. You pay only when you generate. Floyo supports open-source and closed-source models.
A designer runs an edit and likes the result. A teammate opens that exact run from shared history and keeps going. No file handoffs. No version confusion.
For studios and enterprise teams, Floyo adds private workspaces, pooled resources, and a team usage dashboard. Other ComfyUI cloud tools run for one person at a time. Floyo runs for the whole team, with transparent per-generation costs.
Ready to try it? Type what you want said, describe the voice, and run it. The settings are already set.
Questions? Watch the free course or check the FAQ above.
Read more







