Categories:
Tools
elevenlabs voice-ai dubbing voice-cloning content-creation

One Recording, 90 Languages: The ElevenLabs v4 Creator Playbook

Feature image for One Recording, 90 Languages: The ElevenLabs v4 Creator Playbook

Ten seconds of recorded speech. That is what ElevenLabs says it now needs to build a usable clone of your voice, and it is the headline feature of Eleven v4 and Eleven v4 Turbo, the two speech models the company released on September 28. Language support grew in the same jump, from roughly 70 to more than 90.

TechCrunch covered the launch in detail, but the part creators should stare at is the production math. One recording of you can now address markets you have never visited, in languages you do not speak, in a voice that still sounds like you. The dubbing studio, the casting budget, the month of scheduling — that entire layer of cost just got compressed into an API call.

What is actually new in v4

Two models shipped, and they aim at different problems.

Eleven v4 is the quality tier. Its main trick is context-aware delivery: instead of applying one emotional setting to an entire script, the model reads who is speaking and what happened earlier, then decides how a line should land. A question reads like a question. A confession lands softer than a sales pitch. For years the gap between synthetic and human speech lived exactly here, in the flatness of delivery across a long script, and v4 attacks that gap directly.

The expression system got more granular too. v3 introduced inline tags for steering delivery, stage directions for speech. v4 lets you stack several tags and have the model follow them in order. If you produce multi-speaker audio, ads, or dubbed content, this moves you from setting a mood to directing a performance.

Eleven v4 Turbo is the latency tier, built for real-time agents. It starts playing audio while the underlying language model is still generating the rest of the response. That sounds like an engineering footnote until you remember what kills conversational AI in practice: the half-second of dead air before every reply. Turbo removes it. If you are wiring a voice agent into anything customer-facing, Turbo is the one to benchmark first, and v4 is the one for anything your audience downloads and replays.

Why the enterprise number matters

More than half of ElevenLabs’ revenue now comes from large enterprises. That single fact tells you where this market went. Synthetic voice stopped being a creator novelty and became infrastructure for call centers, support lines, and sales systems.

For independent creators this cuts both ways. The tools keep improving because enterprise contracts pay for the development. The competition also gets louder, because every other creator has access to the same models. The edge moves to whoever has something worth saying and the discipline to localize it properly.

A dubbing workflow you can run this week

Here is the pattern. Exact menu names will shift as the product updates, but the workflow holds.

  1. Record a clean sample. Ten seconds clears the technical bar. A minute or two of quiet, close-mic audio makes the clone noticeably better, so give it that.
  2. Clone your voice once, verify it hard. Test the clone on a sentence with numbers, a question, and an emotional shift. Listen for the cadence, not just the timbre.
  3. Script once, translate deliberately. Machine-translate your script, then have a native speaker fix it before synthesis. Voice cloning makes bad translations sound confident, which is worse than sounding broken.
  4. Stack expression tags where it counts. Reserve them for lines that carry emotion. Tagging every sentence produces audio that feels micromanaged.
  5. Review each language output. You will not understand most of them. A native-speaking friend checking five minutes per language is cheap insurance against embarrassing errors.

Pick two languages first. Not twenty. The teams getting real traction from dubbing treat each market as a real audience with its own hooks and titles, and that mindset shows up in their retention numbers.

Ten-second cloning is a useful feature wrapped around a dangerous one. The same mechanism that clones your voice clones anyone’s voice, and the industry’s answer to consent still trails the deployment.

Rules worth holding yourself to, regardless of what the platforms enforce this quarter:

  • Only clone voices you own or have written permission to use. Spouse, friend, celebrity impression: permission or nothing.
  • Disclose synthetic audio where disclosure is expected. Dubbed creator content is increasingly normal, and audiences punish deception far more than they punish dubbing.
  • Keep a record of your consent agreements if you work with other people’s voices. Future disputes get much easier when the paperwork predates them.

There is also a self-interested reason to be careful. Voice is becoming an identity layer for agents and payments. The creators who build reputations as careful, disclosed users of synthetic voice are the ones who will keep platform access when policy tightens. And it will tighten.

What to do now

If you publish video or audio content, run one experiment this week. Take your best-performing piece, clone your voice, dub it into the single largest non-English market in your audience analytics, and publish it with clear dubbing disclosure. One market, one video, measured like everything else you ship.

Eleven v4 makes the technical side almost free now. The parts that still cost real effort — translation quality, market research, honest consent practices — are exactly the parts your competition will skip. That is the opportunity.

Related Articles