How to make a great voice clone
Cloning takes about a minute, and the result is almost entirely decided by the audio you feed it. A clean sample gives you something uncanny. A noisy one gives you something that sounds broken — same bot, same model, different recording.
The short version
One person, talking normally, close to a decent mic, in a quiet room, for about 30–60 seconds, in the language you plan to use. No music, no echo, no second voice. That's most of it — the rest of this page is why.
What makes a good sample
One voice, nothing else
The model copies everything it hears. Background chatter, a second person, game audio, music under the speech — all of it gets baked into the clone. A quiet room with one person talking beats a studio mic in a noisy Discord call.
Give it real speech, not padding
Aim for 20–90 seconds of someone actually talking. Trim the silence at the start and end before you upload — dead air teaches the model nothing, and a 30-second file that's 24 seconds of room tone is really a 6-second sample.
Talk normally
Read a few full sentences in the tone you'll actually use. Flat, robotic reading produces a flat, robotic clone. So does shouting, whispering, or singing — the clone will copy that too and sound strange saying ordinary messages.
Loud and clean, not maxed out
Quiet, distant recordings make thin and unstable clones. So does audio pushed so hard it distorts. Speak close to the mic at a normal volume — if the recording is obviously crackling or clipping, re-record rather than upload it.
Keep every sample consistent
If you upload several files (up to 10, each under 10MB), record them the same way — same person, same mic, same room, same energy. Mixing a phone voice memo with a headset recording gives the model two different voices to average together.
Cover a variety of sounds
A sample that repeats one short phrase gives the model a narrow slice of the voice. A minute of varied sentences — questions, statements, different words — covers far more of how the person actually sounds.
What ruins a clone
Every one of these produces a clone that technically works and sounds wrong.
Music or a game in the background
Gets cloned along with the voice and bleeds into every message.
Two or more people talking
The clone lands somewhere between them and sounds like neither.
Phone speaker or laptop mic across the room
Thin, echoey source audio makes a thin, echoey clone.
Heavy reverb, filters, or autotune
Whatever effect is on the recording becomes part of the voice.
Clips ripped from a compressed video
Re-compressed, low-bitrate audio loses the detail the model needs.
A few seconds repeated to hit the minimum
Length without variety does not help — the model needs range.
Cloning in other languages
Mimiq's default voice catalog leans English. If your server speaks something else, cloning is how you get a voice that sounds right — but the sample matters even more.
Record in the language you'll speak
This is the big one. A clone built from English audio carries an English accent into every language it speaks — Portuguese comes out sounding like an English speaker reading Portuguese. Record the sample in Portuguese, and it sounds Portuguese.
Use a native speaker if you can
The clone inherits pronunciation habits from whoever is talking. A native speaker's sample gives you natural vowels and rhythm; a learner's sample clones the learner's accent, faithfully.
Cover the sounds your language actually uses
Read varied sentences rather than one repeated phrase, so the sample includes the distinctive sounds — nasal vowels in Portuguese, rolled r's in Spanish, umlauts in German. Anything the model never hears, it has to guess at later.
Pairing a clone with translation
/tts translate_to translates your message and then
speaks it in the voice you picked. Results are best when the clone's own language
matches what you're translating into — clone a Portuguese voice for Portuguese
output, not an English one.
Once it's made
Test it with a few different messages
Try a short one, a long one, and a question. Cloned voices vary between generations, so one odd-sounding message isn't a verdict — if the same clone sounds wrong across several tries, the sample is what needs fixing.
Re-cloning is free
If a clone isn't good enough, record a better sample and make a new one. Delete the
old voice with /deletevoice or from the dashboard
to free up a slot. There's no cost and no limit on trying again.
Questions
How long should a voice clone sample be?
Between 20 and 90 seconds in total, across up to 10 files. Somewhere around 30–60 seconds of clean, varied speech is the sweet spot — past that you get diminishing returns, and near the 20-second floor the clone has less to work with and sounds less stable.
My clone sounds wrong — words are clipped or it sounds nothing like the person. What now?
First, just try the message again. Clone output varies from one generation to the next, and an occasional bad take is not the same as a bad clone. If it sounds wrong consistently, the sample is usually the problem — re-record with the tips on this page and create a fresh voice. Deleting and re-cloning is free and takes a minute.
Can I clone a voice in a language other than English?
Yes, and cloning is the best way to get a natural voice in a language the default catalog doesn't cover. Record the sample in the language you'll actually be speaking — a clone built from English audio will carry an English accent into every other language.
Can I clone someone else’s voice?
Only with their explicit permission. Mimiq asks you to confirm this before every clone. Impersonating people, deceiving, or harassing with a cloned voice will get you banned from the bot, and server admins can delete any clone at any time.
Who can use a cloned voice once it exists?
Anyone in that server — clones are shared server-wide, not personal. Members can set one as their own default with /setvoice, and admins can delete clones from the dashboard or with /deletevoice.
Does cloning cost anything?
No. Voice cloning is free and unlimited, like everything else in Mimiq.