Skip to main content
The Human Bit
← AI tools

Voice & audio · ElevenLabs

ElevenLabs

The leading AI voice platform: turn text into natural speech, clone a voice, dub video into other languages, transcribe audio, and build voice agents.

Visit ElevenLabs

What it does

Text to speech

Turns written text into expressive speech across 70-plus languages. Eleven v3 is the most expressive model, Multilingual v2 is steady and lifelike, and Flash runs at roughly 75 milliseconds for live conversation.

For work Narration, voiceover, audiobooks, phone systems and a consistent voice for a product or app.

Voice cloning

Recreate a voice from a sample, or design a new synthetic one from a description. Instant cloning needs about a minute of audio; professional cloning uses more for higher fidelity.

For work A single, consistent brand or creator voice used at scale across many pieces of content.

Dubbing

Translates a video or recording into other languages while re-synthesising the original speaker's own voice, matching their cadence, rather than replacing it with a generic one.

For work Localising courses, marketing videos and social content into new languages.

Speech to text (Scribe)

Transcription with speaker labelling (who said what) and word-level timestamps, across many languages.

For work Subtitles, captions, and meeting or interview transcripts.

Voice agents

Agents that listen and reply in real time across phone, chat and messaging apps in many languages, with analytics and workflow logic.

For work Support lines, booking and qualification calls handled by voice.

Music and sound effects

Generates music tracks and custom sound effects from a prompt. ElevenLabs states its music is trained on licensed data and suitable for commercial use.

For work Background music, jingles and sound design for video, podcasts and games.

Best for

Narration, voiceover, dubbing and prototyping a voice agent, where you want studio-quality speech without a recording booth.

Not for

Anything where using a real person's voice without their clear, recorded consent would be wrong, or where a synthetic voice must not be disclosed.

Real work with it

A small-business owner

Multilingual voiceover for a small business

  1. Write the script, or take an existing video you want in another language.
  2. Pick a library voice or clone your own, then generate the speech, or run Dubbing on the video.
  3. Export the audio and drop it into your video or phone system.

Human checkpoint Listen end to end for mispronounced product and brand names, and confirm a translation actually says what you mean before it goes out.

A teacher or course creator

Narrated lessons with captions

  1. Draft the lesson script.
  2. Generate narration in one consistent voice.
  3. Run Scribe over the audio to produce synced captions and a transcript.

Human checkpoint Proofread the auto-transcript for technical terms and names, and decide whether learners should be told the narration is AI.

Bit, the Human Bit guide

The part that stays yours

A voice is someone's identity. Whose voice you may clone, getting consent on the record, disclosing that a voice is synthetic, and approving the script and tone are all human calls the tool does not make for you.

Plans and pricing

Free

US$0

Trying it out. No commercial licence.

Starter

US$6/mo

First paid step: commercial licence, instant cloning, dubbing.

Creator

US$22/mo

Serious creators: professional voice cloning, extra credits.

Pro

US$99/mo

High volume and API, higher audio quality.

Scale

US$299/mo

Small teams: seats and collaboration.

Business

US$990/mo

Larger teams: low-latency speech, more clones and seats.

Enterprise

Custom

Contracts, SSO, higher concurrency and support.

Free tier, then paid plans from about US$6 (Starter) through Creator, Pro, Scale and Business to custom Enterprise, priced by monthly credits. Only paid plans carry a commercial licence.

When to pick something else

Descript

Your real job is editing a whole podcast or video, not just generating voice. Descript edits audio by editing text.

Murf

You want a friendlier studio interface and template-driven team workflows, and can accept slightly more 'perfect'-sounding voices.

OpenAI or Google TTS

You just need cheap, developer-simple speech inside a stack you already use, without cloning or dubbing depth.

Watch-outs

  • Cloning someone else's voice requires their recorded consent, and cloning public figures is blocked. Keep the consent on file.
  • The prohibited-use policy requires telling people when they are talking to an AI voice, which matters for agents.
  • Generated audio carries an inaudible watermark traceable to the creator, so misuse can be traced back.
  • Credits burn much faster for transcription and dubbing than for plain speech; heavy dubbing can drain a plan quickly.
  • Free-tier output is not licensed for commercial use; you need a paid plan to publish.
Current version
Eleven v3 (text to speech), general availability February 2026
Category
Voice & audio
Checked
31 August 2026
Review due
14 September 2026
Owner
The Human Bit editorial