The voice model is starting up (about 1–2 minutes after the VM boots). You can already record; trying out your clone works once it's ready.
✅ Your voice is ready
Reference sample
Record your voice
- Quiet room, no music. Background noise gets cloned too.
- Read the text below naturally and expressively, like you're talking to a friend. The clone copies your tone, pace and microphone.
- 15–20 seconds is ideal. Speak English (the voice model speaks English and Chinese).
Read this out loud
0:00
…or of just you talking (10–60 s; WAV, MP3, M4A, …).
Check your sample
I trimmed the silence and kept the best s. Listen to it, then make sure the transcript matches exactly what you said: the clone depends on it.
connecting…
Tap to turn on the mic