Upload a sample → clone the voice
Paste any voice sample, write a script, and generate speech that follows those words in a voice matched to your sample. Preview the clone, then make a talking video.
1. Upload photo
Full body or portrait works best. Face should be clear.
Drop your photo here
or click to browse · JPEG, PNG, WebP · max 12MB
Clone a voice → speak your script
Upload a sample of the voice you want. Paste the lines to speak. We generate speech that follows your script in that voice.
1. Voice sample
Drop or click to upload sample audio
MP3 / WAV / M4A · one speaker, quiet room, natural talking
2. Script to speak (cloned voice will say this)
Optional: what the sample says (improves clone quality)
Tip: 10–30 seconds of clean, single-speaker audio clones best. Cloud models are used when free GPU quota allows; otherwise a local adaptive clone runs on this server.
Or pick a celebrity AI voice (optional — disable clone above)
2. Celebrity AI voice (speaks your script)
Pick a voice style — the video will speak whatever you type in the script box.
2b. Emotion & delivery
Controls feeling in the voice — pauses, intensity, and tone.
3. Motion mood
Mood for the clip — speech + lip-sync still drive the mouth motion.
4. Spoken script (required for celebrity voices)
Type the exact words the AI celebrity voice should say.