CosyVoice Voice Clone
Clone a voice from audio and generate speech
Upload a short reference sample and synthesize any script in the cloned voice using Alibaba CosyVoice v3 Plus.
Model Overview
CosyVoice Voice Clone
CosyVoice clones timbre from a short audio sample, then synthesizes your script with natural prosody. Voice enrollment is free; synthesis is billed per character.
Sample-based voice cloning
Alibaba CosyVoice v3 Plus synthesis
Volume, rate, and pitch controls
Multi-language support
API-only, no GPU queue
Why use it on Moosky AI?
Model-specific controls with no subscription, clean outputs, and a workflow built for fast creative iteration.
No subscription required
Buy credits when you need them and pay only for what you generate.
Natural speech
Generate clear narration, dialogue, and voiceover audio from text.
Voice control
Pick built-in voices or reuse a custom voice design ID.
Commercial workflows
Use results in professional projects subject to platform and provider terms.
Example outputs
Featured public generations from the Moosky community using this composer.
How it works
Three focused steps from setup to finished output.
Upload a voice sample
Provide 5–20 seconds of clear reference speech.
Enter your script
Type the text you want the cloned voice to speak.
Download your audio
Receive synthesized speech in the cloned timbre.
Pricing
Credit-based generation
Uses Moosky credits with cost shown before you generate. No subscription required.
Clone a voiceFrequently Asked Questions
What audio sample do I need?
A clip with at least 5 seconds of clear speech. CosyVoice recommends 10–20 seconds (max 60 seconds, 10 MB). Longer video or audio uploads are trimmed automatically before enrollment.
How is this different from Vibe Voice?
Vibe Voice runs on Moosky GPU infrastructure. CosyVoice uses Alibaba Cloud Model Studio APIs for enrollment and synthesis.
Is voice enrollment billed?
Creating the cloned voice is free. You are billed for speech synthesis based on the character count of your script.