Filmtools
Filmmakers go-to destination for pre-production, production & post production equipment!
Shop Now
After the popular Transcription Shootout: DaVinci Resolve, Descript, FCP, FlexClip, Premiere Pro which I recently published (which is speech-to-text), today I bring you Cloned Voice Shootout: FlexClip, ElevenLabs, Descript, DaVinci Resolve Studio. This new article is text-to-my cloned voice (or yours), which is the opposite of the first one about transcription. In this new one about cloned voices, the differences between the results are much greater among the competing services. Most of them don’t even sound like replicas of my human voice. I am including DaVinci Resolve Studio here even though it requires an extra external step to accomplish the goal, as I’ll cover ahead. As in the prior article, I am only covering the US-English results in ProVideo Coalition. You can find the Castilian (aka «Spanish») results and conclusions in Competencia de voces clonadas in Escuchalibros.
How I cloned my voice with each service
My intention was to send the exact same sample recording (one in each language) to each service. However, that became impossible since ElevenLabs recommends a minimum of 30 minutes although prefers to receive 2-3 hours of clean, consistent speech for ideal results. Even though I sent a 56-minute sample to ElevenLabs, the others wouldn’t accept such a long duration, so I had to send just the first part of that full sample to the others. Although FlexClip specifically states that it accepts a maximum of 90 seconds (so I uploaded the first 90 seconds I had sent to ElevenLabs), Descript was not so specific, so I wrote to their Support to see if I could upload 56 minutes. This is the response I received from Descript’s Support Department:
Currently, there is no option to upload an extended sample to an existing AI Speaker (voice clone) to upgrade it to a Professional Voice Clone (PVC) or to add more training data after the initial creation. Here’s how the process works and what’s possible: When creating a custom AI Speaker, you provide a voice sample by recording the required consent statement directly in Descript, or by uploading a recording of the consent statement if you’re acting on behalf of someone else. The voice model is trained only on this initial sample. There is no feature to upload additional or extended samples to further train or enhance an existing AI Speaker. If you want a different style or improved quality, you can create a new AI Speaker by recording a new sample with your preferred tone and delivery. Each new recording creates a separate AI Speaker; you cannot add to or modify an existing one. If you’re looking for the ability to upload longer or additional samples for a more advanced or “professional” voice clone, this feature is not currently supported. You’re encouraged to leave feedback or upvote this request on our feature request board: https://descript.canny.io/feature-requests
Fortunately, DaVinci Resolve Studio accepted the full sample, even though it took the longest (almost 24 hours on a MacBook Pro M4) to create the cloned voice in each language. That’s partially because DaVinci Resolve Studio does the clone locally, while the others do it in their respective cloud. Fortunately, it is necessary to do that long (nearly 24 hour) process only once per cloned voice with DaVinci Resolve Studio.
FlexClip
Above is my US-English cloned voice rendered from the same text I had uploaded to all of the services. If you have heard my human voice in the recent Transcription Shootout: DaVinci Resolve, Descript, FCP, FlexClip, Premiere Pro, you’ll probably agree that my cloned voice from FlexClip is not even recognizable as being the same person. FlexClip did a much better job with transcription than with cloning a voice.
ElevenLabs
Above is the my US-English cloned voice from ElevenLabs, reading the same text I sent to all of the services. If you have heard my human voice in the recent Transcription Shootout: DaVinci Resolve, Descript, FCP, FlexClip, Premiere Pro, you’ll probably agree that my cloned voice from ElevenLabs is quite similar to my human voice, and is quite recognizable as being a reasonable facsimile of my human voice.
Descript
Above is my US-English cloned voice from Descript, reading the same text I sent to all of the services. If you have heard my human voice in the recent Transcription Shootout: DaVinci Resolve, Descript, FCP, FlexClip, Premiere Pro, you’ll probably agree that my cloned voice from Descript is not even recognizable as being the same person, although it does sound more human than FlexClip’s cloned voice. Descript did a much better job with transcription than with cloning a voice.
DaVinci Resolve Studio
Above is my US-English cloned voice from DaVinci Resolve Studio, reading the same text I sent to all of the other services. If you have heard my human voice in the recent Transcription Shootout: DaVinci Resolve, Descript, FCP, FlexClip, Premiere Pro, you’ll probably agree that my cloned voice from DaVinci Resolve Studio is not even recognizable as being the same person, although it does sound more human than FlexClip’s cloned voice. DaVinci Resolve Studio did a much better job with transcription than with cloning a voice.
As stated in the introductory paragraph, as of the publication date of this article, even though DaVinci Resolve Studio is completely capable of cloning voices, it cannot (yet) render one of its cloned voices directly from text. Currently, it is designed to render its cloned voice from audio, even if read by a completely different person. So the current workaround I used is:
- Record the audio of the desired text using any other AI voice (even the one built into macOS), for the sole purpose of feeding it to DaVinci Resolve Studio so it can do its magic.
- Import the audio file into DaVinci Resolve Studio.
- Move the audio file onto the timeline and select it.
- Right-click the selected clip and the timeline and select Voice Convert…
- I chose New Track to keep the original audio.
- In the Voice Model dropdown, select the custom model you trained earlier. I selected Allan-English for the ProVideo Coalition test and Allan-castellano for the one I did in Castilian for Escuchalibros.
- I deselected the option Tight Matching to Source since there is no video in this test.
- Click Render, and DaVinci Resolve Studio will generate the new audio onto a new track. I muted the original track.
Conclusions from the US-English test
To my ear:
- The best quality is from ElevenLabs. It sounds like the most convincing clone of my human voice. It is too bad that currently, in order to have two of the best-quality cloned voices in two different languages, we must have two separate ElevenLabs accounts, each one for US$22 per month using two different email addresses. But it’s the price for that quality. I wish they’d allow us to have a single account for bilingual use from the same individual at the highest quality.
- I would say that DaVinci Resolve Studio and Descript are tied for second place: Although neither is recognizable as my human voice, they are both usable voices for certain projects. However (as indicated in the article) extra steps are currently required to accomplish this task with DaVinci Resolve Studio.
- FlexClip’s cloned voice is the most robotic-sounding compared with the others. FlexClip did a much better job with transcription than with voice cloning in its current version.
For voice cloning in the Castilian language (aka «Spanish»), I have different rendered results and observations in that version of the test in Escuchalibros, as linked below.
Lee y escucha este estudio en buen castellano en Escuchalibros
Competencia de voces clonadas: FlexClip, ElevenLabs, Descript, DaVinci Resolve Studio
(Re-)Subscribe for upcoming articles, reviews, radio shows, books and seminars/webinars
Stand by for upcoming articles, reviews, books and courses by subscribing to my bulletins.
In English:
- Email bulletins, bulletins.AllanTepper.com
- In Telegram, t.me/TecnoTurBulletins
- Twitter (bilingual), AllanLTepper
En castellano:
- Boletines por correo electrónico, boletines.AllanTepper.com
- En Telegram, t.me/boletinesdeAllan
- Twitter (bilingüe), AllanLTepper
Most of my current books are at books.AllanTepper.com, and also visit AllanTepper.com and radio.AllanTepper.com.
FTC disclosure
None of the listed individuals or companies mentioned above has paid for this article. Blackmagic and FlexClip have sent Allan Tépper equipment or NFR software for evaluation and reviews. Some of the manufacturers listed above have contracted Tépper and/or TecnoTur LLC to carry out consulting and/or translations/localizations/transcreations. So far, none of the manufacturers listed above is/are sponsors of the TecnoTur, BeyondPodcasting, CapicúaFM or TuSaludSecreta programs, although they are welcome to do so, and some are, may be (or may have been) sponsors of ProVideo Coalition magazine. Some links to third parties listed in this article and/or on this web page may indirectly benefit TecnoTur LLC via affiliate programs. Allan Tépper’s opinions are his own. Allan Tépper is not liable for misuse or misunderstanding of information he shares.
