How to Use ElevenLabs: A Beginner’s Guide (2026)
Introduction
Creating a professional voiceover normally requires a microphone, a quiet recording space, repeated takes, and audio-editing software. ElevenLabs simplifies the process by allowing users to turn written text into natural-sounding speech directly from a browser.
Beginners can select a voice, paste or type a script, choose an AI model, and generate downloadable audio within minutes. No coding or traditional audio-production experience is required for ElevenLabs’ browser-based creative tools.
ElevenLabs includes more than basic Text to Speech. Users can browse thousands of voices, design a synthetic voice, clone an authorized voice, transform recorded speech, create sound effects and music, translate existing audio or video, and organize longer projects inside ElevenCreative Studio.
The platform currently offers several speech models for different purposes. Eleven Multilingual v2 focuses on stable long-form narration, Eleven v3 provides more expressive delivery and multi-speaker dialogue, and Flash v2.5 prioritizes faster generation for real-time applications.
This beginner’s guide will explain how to:
- Create an ElevenLabs account and understand the dashboard
- Generate a first voiceover with Text to Speech
- Choose a voice and speech model
- Format scripts for more natural narration
- Adjust stability, similarity, style, and speaking speed
- Correct pronunciation and pacing problems
- Browse and save voices from the Voice Library
- Create a voice with Voice Design
- Clone your own voice with permission
- Transform a recording with Voice Changer
- Create multi-speaker dialogue
- Build longer projects inside ElevenCreative Studio
- Generate sound effects and music
- Transcribe, isolate, and dub existing recordings
- Download, organize, and reuse generated audio
- Manage credits, licensing, and plan limits
ElevenLabs provides a free account for learning the basic tools, but free-plan content is intended for personal, noncommercial use. Commercial usage rights require an eligible paid subscription.
Beginners should start with a short script and one ready-made voice before experimenting with cloning, dubbing, dialogue, or long-form production. Learning the basic Text to Speech workflow first makes the platform’s more advanced audio tools easier to understand.
Create an ElevenLabs Account
ElevenLabs allows users to create a free account before choosing a paid subscription. New accounts are automatically placed on the Free plan, so payment information is not required to begin testing the platform.
Sign Up for ElevenLabs
- Open the ElevenLabs website.
- Select Get Started Free.
- Sign up with an email address or an available single-sign-on option.
- Enter the requested account information.
- Accept the applicable terms.
- Complete the email-verification step when required.
- Enter the ElevenLabs workspace.
ElevenLabs currently supports traditional email registration along with Google and Apple sign-in. GitHub and Facebook are no longer offered as new sign-in options.
Choose the Sign-In Method Carefully
Using an email address and password provides a standard login that can be managed directly through ElevenLabs.
When using Google sign-in, the ElevenLabs account remains connected to that Google email address. ElevenLabs states that users who register through Google cannot later change the account to a different email address.
Use an email address that you expect to keep, especially when the account may eventually contain custom voices, long-form projects, downloaded audio, or a paid subscription.
Verify the Email Address
Users who register with a traditional email address must verify it before generating audio.
- Open the verification message from ElevenLabs.
- Click the verification link.
- Return to the ElevenLabs website.
- Sign in when requested.
Google and Apple sign-ins are verified automatically. When an email verification message does not appear, check the spam or junk folder before requesting another one.
After verification, ElevenLabs normally directs the user to the Speech Synthesis or Text to Speech area, where a first voiceover can be generated immediately.
Understand the Free Account
The Free plan currently includes 10,000 credits per month. For standard Text to Speech use, this represents approximately ten minutes of narration, although the actual amount depends on the selected model and any other ElevenLabs tools that use the same credit balance.
The free account provides access to tools such as:
- Text to Speech
- Voice Changer
- Sound Effects
- Voice Design
- The Voice Library
- The creative Playground
ElevenLabs’ current quick-start documentation also lists access to more than 3,000 searchable voices for new creative-platform users.
The Free plan is enough for learning the interface and testing short scripts. It is not intended to support a large production schedule.
Do Not Use All the Credits Immediately
Every time a user confirms a generation, ElevenLabs deducts credits from the account. Credits are charged for the generation request rather than for downloading the finished audio.
Begin with a short test such as:
“Welcome to this ElevenLabs tutorial. In this guide, you’ll learn how to create a natural AI voiceover.”
A brief sample makes it possible to test the voice, model, pronunciation, and settings without spending credits on a complete script that may need to be rewritten.
Understand the Shared Credit Balance
ElevenLabs products draw from one shared pool of credits. Using credits for music, transcription, dubbing, sound effects, Voice Changer, or another tool leaves fewer credits for Text to Speech.
This means 10,000 free credits do not guarantee ten minutes of completed narration when the account is also being used for other features.
For example, a beginner who generates several sound effects and repeatedly tests different voices may have fewer credits available for the final voiceover.
Check the Remaining Credits
To review the available balance:
- Select the profile icon.
- Open Subscription.
- Review the current plan and remaining credits.
- Check when the next billing cycle begins.
The Subscription page also displays the available paid plans and their included allowances.
Check the balance before creating a long voiceover, dubbing a video, or testing several expensive generation tools.
Understand Credit Renewal
Free-plan credits reset at the beginning of each billing cycle. Unused Free-plan credits do not roll over into later months. Credit rollover is limited to active paid subscriptions that remain on an eligible plan.
There is therefore no benefit to saving free credits indefinitely, but they should still be used carefully enough to complete meaningful tests.
Use Only One Free Account
ElevenLabs currently allows one free account per user and IP address. Creating several free accounts may trigger the platform’s abuse-prevention system and disable free-tier usage.
A VPN, shared workplace network, school network, mobile network, or household connection may also occasionally trigger an unusual-activity warning because several users appear under the same IP address.
Use the original account whenever possible rather than creating another one when login or credit issues occur.
Review the Account Settings
Open the profile menu and select Settings to review the account’s email address and other available controls. This is also useful when saved voices or projects appear to be missing, because the user may have accidentally signed into a different account.
Before creating important content, confirm that:
- The correct email address is connected.
- The account is verified.
- The expected subscription is active.
- The credit balance is visible.
- The correct workspace is open.
Start With the Free Plan
Most beginners should remain on the Free plan until they understand how Text to Speech, voice selection, models, settings, and credits work.
Create several short samples before upgrading. Pay attention to whether the platform provides the required voice quality, languages, accents, and workflow.
Once the account is ready, the next step is learning how to navigate the ElevenLabs dashboard and locate its main creative tools.
Understand the ElevenLabs Dashboard
The ElevenLabs dashboard brings its voice, music, sound-effect, dubbing, transcription, and production tools into one workspace. The exact sidebar may change as ElevenLabs updates the platform, but the main sections are organized around Playground, Products, Voices, Assets, and account management.
Open the Creative Platform
After signing in, confirm that you are inside ElevenCreative rather than ElevenAgents or the developer documentation area.
ElevenLabs separates its platform into three main environments:
- ElevenCreative for generating and editing audio, music, video, and other creative content
- ElevenAgents for building interactive voice assistants
- ElevenAPI for connecting ElevenLabs features to external software
Beginners creating voiceovers should remain inside ElevenCreative.
Use the Playground
The Playground is the easiest place to experiment with individual AI audio tools. It includes workflows for Text to Speech, Voice Changer, and Sound Effects.
Use the Playground when you want to:
- Test a short script
- Compare several voices
- Experiment with voice settings
- Transform a recording
- Generate an individual sound effect
- Learn how credits are deducted
The Playground is generally better for short tests and individual audio clips than for organizing a complete audiobook, course, or long-form production.
Find Text to Speech
Text to Speech is the main tool beginners will use.
Open the speech-generation area and look for:
- The script box
- The selected voice
- The speech model
- The voice settings
- The generation button
- The playback and download controls
- The generation history
Enter or paste a short script, select a voice, and generate a sample. Do not adjust every setting immediately. The default controls provide a useful starting point for learning how the selected voice behaves.
Open the Voice Library
The Voices section contains the voice options available for speech generation. ElevenLabs currently provides more than 10,000 community voices in addition to default, cloned, and artificially designed voices.
Voice categories may include:
- Default voices created by ElevenLabs
- Community voices shared through the Voice Library
- Cloned voices created with authorized recordings
- Voice Design voices generated from written descriptions
From the Voice Library, users can search according to language, accent, age, use case, and other available characteristics. A voice can then be added to the user’s personal collection for easier access during future generations.
Review My Voices
The personal voice area contains voices that have been saved, designed, or cloned under the account.
Use this section to:
- Reopen saved community voices
- Manage Instant Voice Clones
- Access Professional Voice Clones
- Find voices created through Voice Design
- Rename or remove voices
- Review voice information and availability
Organize voices carefully when testing several options. Names such as Warm Narrator, Training Voice, or Spanish Product Voice are more useful than keeping several similar automatic names.
Open ElevenCreative Studio
ElevenCreative Studio is the main production workspace for longer audio and video projects. Studio 3.0 uses a timeline containing tracks for narration, video, captions, music, and sound effects.
Studio can be used to create:
- Audiobooks
- Podcasts
- Narrated articles
- Courses
- Video voiceovers
- Faceless videos
- Multispeaker projects
- Audio stories
- Productions containing music and sound effects
Beginners should learn basic Text to Speech before moving into Studio. The timeline provides more control, but it also introduces additional tracks, chapters, timing tools, and export settings.
Start a Studio Project
The Studio page provides several starting options.
Users can:
- Upload an existing file
- Create a blank audio or video project
- Create a faceless video
- Add captions to a video
- Dub a video
- Add a voiceover
- Generate a soundtrack
- Narrate an article
- Generate a script
- Narrate an audiobook
Each guided option opens a workflow designed for that particular task. A blank project provides more control, while a prepared workflow reduces the amount of setup required.
For a first Studio test, create a short audio project containing only narration. Add video, captions, music, and sound effects after the basic timeline feels familiar.
Understand the Studio Timeline
The timeline displays each part of the project on a separate track. Users can move, trim, split, duplicate, and align audio or video clips.
Common tracks include:
- Narration
- Video
- Captions
- Music
- Sound effects
The scene or chapter continues for as long as its active content requires. A music clip or video that extends beyond the narration may therefore make the exported project longer than expected.
Studio also allows users to regenerate individual paragraphs, change voices for selected sections, review earlier generations, and lock approved narration so it is not changed accidentally.
Open Dubbing
The Dubbing section translates existing audio or video into another language while attempting to preserve each speaker’s voice, timing, tone, and emotional delivery.
To begin, users can upload a file or paste a supported online video URL, choose one or more target languages, review the displayed cost, and generate the dub. ElevenLabs currently supports dubbing across more than 90 languages.
Automatic Dubbing provides a faster workflow. Dubbing Studio provides more control over transcripts, speakers, translations, and regenerated sections.
Do not begin with dubbing while still learning basic speech generation. It uses more credits and requires careful review of language, timing, and speaker assignments.
Open Music
The Music section generates complete songs or instrumental tracks from written prompts.
Users can describe:
- Genre
- Mood
- Instruments
- Vocal style
- Song structure
- Approximate duration
- Intended use
Eleven Music also supports editing song sections, working with lyrics, and using short audio references to guide the style of supported generations.
Music generation uses the account’s shared credits. Test a short and specific prompt before creating several complete tracks.
Find Sound Effects
The Sound Effects tool creates individual audio effects from written descriptions. It is useful for ambience, Foley sounds, transitions, interface sounds, cinematic effects, and background environments.
Sound effects can be created separately in the Playground or generated directly inside ElevenCreative Studio and placed on the timeline. Studio allows users to trim, duplicate, layer, and reposition the effects.
Use descriptive prompts such as:
“Short digital confirmation sound for a mobile payment app.”
This is more useful than entering only “notification sound.”
Use Assets
The Assets area provides centralized management for creative files and resources used throughout ElevenCreative.
Depending on the account and plan, this area may include:
- Saved voices
- Uploaded recordings
- Music
- Sound effects
- Generated media
- Shared workspace assets
- Files connected to Studio projects
Clear filenames make assets easier to find later. Rename uploads before several similar recordings accumulate in the workspace.
Review Generation History
Generation History stores earlier outputs from supported ElevenLabs tools. Use it to reopen and download an approved version instead of spending credits to generate the same content again.
When comparing several attempts:
- Listen to each version.
- Download the strongest result.
- Rename the file clearly.
- Record which voice, model, and settings were used.
- Avoid deleting an approved version until the project is complete.
AI output can vary between generations, so an earlier version may be difficult to reproduce exactly.
Check the Subscription and Credits
Open the profile or account menu to review the active subscription, remaining credits, billing information, and workspace settings.
Check the credit balance before using:
- Long Text to Speech scripts
- Dubbing
- Music generation
- Voice Changer
- Voice Isolation
- Large Studio exports
- Several voice tests
All of these tools may draw from the same shared balance, so using one feature reduces the amount available for others.
Confirm the Active Workspace
Users who belong to a shared ElevenLabs workspace should confirm that the correct workspace is active before creating voices or projects.
Workspaces can contain shared credits, members, assets, and permissions. Full-seat members can access ElevenCreative, ElevenAgents, and ElevenAPI according to the workspace’s available features and credit pool.
Creating content in the wrong workspace can make it difficult to locate later or may use credits belonging to another team.
Ignore Advanced Tools at First
ElevenLabs includes many features, but beginners do not need to learn them all immediately.
Start with:
- Text to Speech
- Voice selection
- Model selection
- Basic voice settings
- Generation History
- Downloading audio
After those steps feel comfortable, move into Voice Design, cloning, Voice Changer, Studio, dubbing, music, and sound effects.
Learning one workflow at a time makes the dashboard easier to understand and prevents unnecessary credit usage. Once the main navigation feels familiar, the next step is generating a first voiceover with Text to Speech.
Generate Your First Voiceover With Text to Speech
Text to Speech is the simplest ElevenLabs workflow for beginners. Users enter a script, select a voice and speech model, adjust any optional settings, and generate an audio file directly in the browser.
Open Text to Speech
- Sign in to ElevenLabs.
- Open Text to Speech from the sidebar or Playground.
- Confirm that the script input box is visible.
- Check the selected voice and model.
- Leave the advanced settings at their defaults for the first test.
The selected voice has the greatest effect on the result, followed by the speech model and then the remaining voice settings.
Enter a Short Test Script
Type or paste the words you want ElevenLabs to read.
For a first test, use a short passage such as:
“Welcome to this beginner’s guide. Today, you’ll learn how to create a natural AI voiceover with ElevenLabs.”
A short script lets you evaluate the voice, pronunciation, pacing, and model without using credits on a complete project that may need substantial editing.
ElevenLabs currently limits individual website generations to 2,500 characters on the Free plan and 5,000 characters on paid plans. Longer content is better organized inside ElevenCreative Studio.
Select a Voice
Open the voice menu near the lower-left portion of the Text to Speech interface and choose a narrator.
For the first generation, use one of ElevenLabs’ ready-made voices rather than creating or cloning a new one. Listen to its preview and consider whether it matches the project’s audience and tone.
A calm, clear voice may suit a tutorial, while a more energetic narrator may work better for promotional or social-media content. The exact voice-selection process will be covered in greater detail in the next section.
Choose a Speech Model
Select the model that best fits the test.
ElevenLabs currently recommends different models for different priorities:
- Eleven Multilingual v2 for stable, natural long-form speech
- Eleven v3 for expressive or dramatic delivery
- Eleven Flash v2.5 for fast, lower-cost generation and real-time applications
Eleven v3 supports more than 70 languages and expressive audio directions, Multilingual v2 supports longer and more stable narration, and Flash v2.5 prioritizes speed.
Beginners creating a standard narrated video can start with Multilingual v2. Use Eleven v3 later when the script requires stronger emotion, character performance, or multi-speaker dialogue.
Let ElevenLabs Detect the Language
ElevenLabs determines the spoken language from the written script instead of requiring users to select the language separately in the standard Text to Speech workflow. The selected voice still influences the accent and pronunciation.
Enter the full test sentence in one language. Mixing several languages in a very short sample can make it harder to judge whether the voice is suitable.
For the most natural result, select a voice that already matches the language and regional accent you need.
Generate the Speech
Once the script, voice, and model are selected:
- Review the displayed credit cost.
- Click Generate or Generate Speech.
- Wait for the audio to process.
- Press play to hear the result.
- Listen to the complete recording before changing any settings.
ElevenLabs creates the audio based on the exact script, selected voice, model, and current settings. Generative speech is not completely deterministic, so repeating the same generation may produce small differences in pacing or emphasis.
Review the First Result
Listen for:
- Incorrect pronunciation
- Awkward pauses
- Unnatural emphasis
- Speech that is too fast or slow
- A voice that does not fit the content
- Sentences that sound written rather than spoken
- Sudden changes in tone or volume
Do not immediately adjust every slider. First decide whether the main problem comes from the script, voice, or model.
For example, rewrite a complicated sentence before changing several voice settings. A clearer script often produces a more natural result without additional customization.
Improve the Script
When the narration sounds unnatural, make small changes to the wording and punctuation.
For example:
Original:
“ElevenLabs is an AI-powered platform which provides its users with numerous different audio-generation capabilities.”
Improved:
“ElevenLabs is an AI audio platform. It can create voiceovers, music, sound effects, and translated audio.”
The improved version uses shorter sentences and clearer pauses, making it easier for the AI voice to deliver naturally.
Commas, periods, question marks, and paragraph breaks can influence the narration’s rhythm and emphasis. ElevenLabs models adapt their delivery to textual cues, including context and punctuation.
Generate Another Version
After making a correction:
- Confirm that the revised script is accurate.
- Generate the audio again.
- Compare it with the previous version.
- Keep the stronger result.
- Avoid repeatedly generating the entire passage for one weak sentence.
When the wording is already correct, another generation may still provide a better performance because the output can vary slightly each time.
For longer scripts, divide the content into manageable sections. This makes it easier to correct one weak passage without regenerating several minutes of approved narration.
Download the Audio
After selecting the preferred version, click the download button near the generated audio.
ElevenLabs supports immediate downloads and allows earlier Text to Speech generations to be retrieved from the History panel. Standard download choices include MP3 and WAV, while advanced options currently include higher-bitrate MP3, M4A, and FLAC.
Use:
- MP3 for smaller files, websites, social content, and general video editing
- WAV for higher-quality editing and professional production
- M4A when that format better suits the intended application
- FLAC for lossless audio with a smaller file size than an uncompressed WAV
Reopen a Previous Generation
To find earlier audio:
- Open Text to Speech.
- Select the History tab or history icon.
- Find the correct generation.
- Play it to confirm the content.
- Click its download icon.
- Choose the preferred file format.
Previously generated Text to Speech files remain accessible through History, allowing users to download an earlier result instead of spending credits to recreate it.
Name the Download Clearly
Rename the file immediately after downloading it.
For example:
ElevenLabs-Tutorial-Introduction-Final.wav
When several versions exist, include the section, voice, and revision:
ElevenLabs-Tutorial-Introduction-Warm-Voice-V2.wav
Clear filenames prevent an unfinished test from being confused with the approved recording.
Complete a First-Generation Check
Before moving on, confirm that:
- The correct script was generated.
- The narrator fits the project.
- The language and accent sound appropriate.
- Important words are pronounced correctly.
- The pacing remains comfortable.
- The strongest version has been downloaded.
- The file has a recognizable name.
- The remaining credit balance is sufficient.
The first generation does not need to be perfect. Its purpose is to show how the voice, model, and script work together. Once this basic workflow feels familiar, the next step is learning how to choose the right voice and speech model for a complete project.
Choose the Right Voice and Speech Model
The selected voice and speech model work together to determine how the finished narration sounds. The voice controls the speaker’s basic identity, accent, tone, and delivery style, while the model affects qualities such as expressiveness, consistency, language support, generation speed, and input length.
Beginners should select the voice first and then choose the model that best matches the project.
Choose a Voice for the Project
Open the voice selector inside Text to Speech and preview several options using part of the real script.
Consider:
- The audience
- The language and regional accent
- The speaker’s apparent age
- The required energy level
- Whether the project is educational, conversational, promotional, or dramatic
- How long viewers will listen to the voice
A calm narrator may work well for tutorials, audiobooks, and training. A more energetic voice may suit advertisements, introductions, and short social-media videos.
Do not judge a voice only by its default preview. Generate a short passage containing the same terminology, sentence length, and tone used in the actual project.
Use a Voice That Matches the Language
ElevenLabs automatically detects the language from the script when speech is generated through its website. However, the voice itself determines much of the accent.
A voice that was not trained in the target language may retain its original accent or drift between accents. ElevenLabs recommends choosing a voice trained in the intended language and regional accent for the most natural pronunciation and intonation.
For example, a Spanish voice intended for Spain may sound different from one trained with a Mexican or American Spanish accent.
Avoid placing several languages in one short prompt while testing. Automatic language detection may become less reliable when the context is unclear.
Search the Voice Library
To find more voices:
- Select Voices from the sidebar.
- Open Explore.
- Enter a keyword or voice name.
- Select the intended language.
- Apply an accent when available.
- Filter by age, gender, category, or other relevant characteristics.
- Play the preview for each promising voice.
- Select Use Voice to open it directly in Text to Speech.
The Voice Library contains community-shared Professional Voice Clones. It can be searched by name, keyword, or voice ID and filtered by language, accent, gender, age, category, notice period, moderation status, and custom credit rate.
Users can also upload a clean speech recording to search for the original voice or similar options.
Save a Voice to My Voices
When you find a suitable option, click the plus button to add it to My Voices. Saved voices become available in the voice menus across ElevenLabs.
A saved voice can also be opened directly in Text to Speech by selecting its T button from My Voices.
Give important voices a clear purpose in your production notes, such as:
- Main YouTube narrator
- Calm training voice
- Spanish tutorial voice
- Audiobook character
- Product advertisement voice
This helps maintain consistency across related projects.
Check the Voice’s Credit Rate
Some Voice Library voices have custom credit multipliers. A voice with a 2x multiplier consumes twice as many credits as a standard-rate voice for the same generation.
Voices with credit multipliers may also be unavailable on the Free plan. ElevenLabs displays the restriction when a free user attempts to select one.
Check the rate before using a community voice for a long audiobook, course, or recurring video series.
Check the Notice Period
Community voices are controlled by their owners and may eventually be removed from the Voice Library. Some have a notice period that allows existing users to continue accessing them temporarily after removal.
When a voice is used, its current notice period is saved for the account.
For a long-term project, consider using:
- An ElevenLabs default voice
- A Voice Design voice you created
- An authorized Instant Voice Clone
- A verified Professional Voice Clone
These choices reduce the risk of depending on a community voice that later becomes unavailable.
Choose Eleven Multilingual v2 for Reliable Narration
Eleven Multilingual v2 is a strong default for professional content, audiobooks, training, tutorials, and video narration.
It supports 29 languages, accepts up to 10,000 characters per request through supported workflows, and is ElevenLabs’ most stable option for longer speech generations. It prioritizes lifelike output, consistent voice quality, and emotionally nuanced delivery.
Use Multilingual v2 when:
- Audio quality matters more than speed
- The narration is several minutes long
- The project contains numbers, dates, or detailed information
- The speaker should remain consistent
- The content is educational or professional
- A multilingual voiceover needs polished delivery
This is usually the safest starting model for a beginner creating a normal prerecorded voiceover.
Choose Eleven v3 for Expressive Speech
Eleven v3 is ElevenLabs’ most expressive Text to Speech model. It supports more than 70 languages, multi-speaker dialogue, emotional performance, and a character limit of approximately 5,000 per request.
Use Eleven v3 for:
- Dramatic narration
- Fictional characters
- Audiobook dialogue
- Podcast-style conversations
- Emotional storytelling
- Advertisements
- Game dialogue
It can respond to punctuation, text structure, and supported audio tags that guide effects such as whispering, laughter, excitement, or hesitation.
However, Eleven v3 can produce more variable results than the v2 and v2.5 models, especially with some Voice Library voices. Voice selection is particularly important when using this model.
Do not choose v3 merely because it is newer. A stable tutorial may sound better with Multilingual v2.
Choose Flash v2.5 for Speed
Flash v2.5 prioritizes low latency and lower-cost API generation. It supports 32 languages, offers a character limit of up to 40,000 in supported API requests, and has an estimated model inference latency of approximately 75 milliseconds before network and application delays are included.
Use Flash v2.5 for:
- Voice agents
- Chatbots
- Interactive applications
- Games requiring immediate responses
- Bulk speech generation
- Fast draft previews
- Projects where speed matters more than maximum polish
Flash v2.5 provides a useful balance between speed and quality, but Multilingual v2 remains the stronger recommendation when the highest-quality prerecorded narration is the priority.
Be Careful With Numbers in Flash v2.5
Flash v2.5 does not normalize numbers, dates, phone numbers, and currencies as reliably by default because that additional processing would increase latency.
ElevenLabs recommends Multilingual v2 when number normalization matters or rewriting the script in the exact form that should be spoken.
For example, instead of entering:
“Call 610-555-0184 on 7/28.”
Write:
“Call six one zero, five five five, zero one eight four, on July twenty-eighth.”
Always preview financial figures, dates, addresses, measurements, and phone numbers before generating the full project.
Match the Model to the Use Case
A simple selection guide is:
- Choose Multilingual v2 for polished narration, audiobooks, professional videos, and long-form content.
- Choose Eleven v3 for emotional performance, character voices, and dialogue.
- Choose Flash v2.5 for speed, interactive applications, and lower-cost API generation.
ElevenLabs officially recommends Multilingual v2 when quality is the main requirement, Flash models for low-latency applications, and either Multilingual v2 or Flash v2.5 for supported multilingual use.
Test the Voice and Model Together
A strong voice may sound different across models. Do not assume that a voice that performs well with Multilingual v2 will deliver the same pacing or emotion with Eleven v3.
Create the same short passage with two likely models and compare:
- Naturalness
- Pronunciation
- Emotional delivery
- Speaking speed
- Consistency
- Background artifacts
- Credit cost
Use a passage containing the most difficult part of the script rather than a simple introduction.
Avoid Changing Voices Mid-Project
Once a voice has been approved for a long project, use the same voice, model, and general settings throughout related sections.
Changing the narrator without a clear reason can make an audiobook, course, or video feel inconsistent. Even the same voice may perform differently after switching models.
Record the selected combination somewhere outside the platform:
- Voice name or ID
- Speech model
- Stability setting
- Similarity setting
- Style setting
- Speaking speed
This makes it easier to produce future updates that sound similar.
Complete a Voice and Model Check
Before generating the full script, confirm that:
- The voice matches the audience and subject.
- Its accent fits the target language and region.
- The model supports the required language.
- The voice handles important names and terminology.
- The emotional delivery fits the project.
- The credit multiplier is acceptable.
- The voice’s availability is suitable for long-term use.
- The selected combination sounds consistent across several sentences.
Choosing the correct voice and model at the beginning reduces the need to regenerate large sections later. Once the combination is approved, the next step is formatting the script so the narration sounds more natural.
Format Scripts for More Natural Narration
ElevenLabs uses more than the individual words in a script. Sentence structure, punctuation, capitalization, paragraph breaks, and surrounding context can all influence pacing, emphasis, emotion, and pronunciation. A script written for speech will usually sound more natural than text copied directly from an article or report.
Write for Listening
Use language that sounds natural when spoken aloud. Short sentences are generally easier for both the narrator and listener to follow.
For example:
Written style:
“ElevenLabs provides users with several audio-production capabilities that can be used across a variety of personal and professional projects.”
Spoken style:
“ElevenLabs includes several AI audio tools. You can use them for personal projects, professional videos, audiobooks, and more.”
The second version creates clearer pauses and reduces the chance of unnatural emphasis.
Read the Script Aloud
Before generating the narration, read the complete script yourself. Rewrite any sentence that feels difficult to say in one breath or sounds overly formal.
Pay attention to:
- Sentences containing several separate ideas
- Repeated words
- Complicated transitions
- Unnecessary introductions
- Technical language the audience may not understand
- Wording that looks correct but sounds unnatural
Reading aloud also helps identify where the script needs another sentence break or a more natural pause.
Use Standard Punctuation
Periods, commas, question marks, and exclamation points help ElevenLabs interpret the intended rhythm.
Use:
- A period for a complete stop
- A comma for a brief pause
- A question mark for a question
- An exclamation point for stronger energy
- A dash for a sudden break or interruption
- An ellipsis for hesitation or a weighted pause
Punctuation has a particularly strong effect on Eleven v3. ElevenLabs notes that ellipses add pauses and weight, while capitalization can increase emphasis.
For example:
Flat:
“You finished the first step now open your account settings.”
Improved:
“You finished the first step. Now, open your account settings.”
Use Ellipses Carefully
An ellipsis can create hesitation, uncertainty, or a dramatic pause.
For example:
“I thought the process would be difficult… but it only took a few minutes.”
Do not use ellipses simply because you want ordinary spacing between every sentence. ElevenLabs warns that they can introduce a hesitant or nervous delivery that may not fit professional narration.
Add Precise Pauses With Break Tags
Supported models and workflows can use an SSML break tag when the script needs a more controlled pause.
For example:
The first step is complete. <break time="1.5s" /> Now, open the dashboard.
Break durations can be set up to three seconds. However, using too many break tags in one generation may cause faster speech, extra noise, or other audio instability.
Eleven v3 does not use SSML break tags in the same way. For v3, control pauses with punctuation, script structure, or expressive instructions such as [pause], [short pause], and [long pause].
Divide the Script Into Paragraphs
Begin a new paragraph when:
- The topic changes
- Another step begins
- A different speaker starts talking
- The emotional direction changes
- The narration moves from an introduction to an explanation
- The script needs a noticeable pause
Paragraph breaks give the model clearer context and make the script easier to review. They also simplify corrections because each passage can be generated separately.
Keep Each Generation Focused
Avoid pasting an entire long article into one generation simply because the character limit allows it. Generate the narration in logical sections such as the introduction, main points, examples, and conclusion.
Shorter sections make it easier to:
- Correct one weak sentence
- Compare different voices
- Fix pronunciation
- Match pacing between passages
- Replace a line without regenerating the complete recording
- Organize approved versions
Keep enough surrounding context for the voice to understand the intended emotion and delivery. A single isolated word may not provide enough information for a natural performance.
Write Numbers the Way They Should Sound
Numbers, dates, prices, phone numbers, measurements, and symbols may have more than one possible pronunciation. ElevenLabs enables text normalization by default in its website workflow, but the result can still depend on the language, model, and context.
When accuracy matters, write the exact spoken version.
For example:
2026→twenty twenty-six$1,250→one thousand two hundred fifty dollars7/28→July twenty-eighth610-555-0184→six one zero, five five five, zero one eight four5 km→five kilometers
Listen carefully to financial figures, addresses, medication quantities, statistics, and other information where a pronunciation mistake could change the meaning.
Write Acronyms Clearly
An acronym may be pronounced as a word or as separate letters.
For example:
NASAis normally spoken as one word.AIis normally spoken as separate letters.FAQmay be spoken as “F.A.Q.” or “frequently asked questions.”
Add periods or spaces when ElevenLabs combines letters that should be read individually:
A.I.U.S.A.C.E.O.
Alternatively, replace the acronym with its full spoken form when that sounds clearer. ElevenLabs also supports pronunciation dictionaries for consistently replacing acronyms and other terms in supported workflows.
Rewrite URLs and Symbols
Do not assume the narrator will pronounce a URL, email address, hashtag, mathematical symbol, or special character correctly.
Instead of:
Visit elevenlabs.io/docs.
Write:
Visit eleven labs dot I O slash docs.
ElevenLabs specifically recommends rewriting URLs according to how they should be spoken. Symbols and non-text characters may also reduce speech quality when the intended pronunciation is unclear.
Add Emphasis With Capitalization
Eleven v3 can interpret capitalization as stronger emphasis.
For example:
“This is the MOST important setting.”
Use capitalization selectively. Writing complete sentences in capital letters may cause exaggerated or unnatural delivery.
For normal informational narration, clear sentence structure usually works better than repeatedly forcing emphasis.
Use Audio Tags With Eleven v3
Eleven v3 supports audio tags that guide emotional delivery and audible reactions.
Examples include:
[whispers][laughs][sighs][excited][sarcastic][short pause]
Place the tag before or after the section it should influence.
For example:
[excited] Your first voiceover is ready!
or:
I thought I had lost the recording. [relieved sigh] But it was still in my history.
Tags should match the voice’s natural character. A serious narrator may not respond convincingly to playful instructions, while an energetic voice may struggle to sound calm or secretive. Tag performance can also vary between voices, so test the result before applying the same style throughout a project.
Do Not Overuse Emotional Instructions
Audio tags are most useful when a line genuinely requires a change in emotion or delivery. Adding tags to every sentence can make narration sound exaggerated and inconsistent.
For a normal tutorial, a clean script with natural punctuation may be enough. Reserve expressive tags for introductions, reactions, dialogue, important warnings, and storytelling moments.
Format Dialogue Clearly
When creating a conversation, place each speaker’s line separately and identify the correct voice.
For example:
Speaker 1:
“Did you finish the voiceover?”
Speaker 2:
“Almost. I’m correcting one pronunciation first.”
Keep dialogue conversational. Short responses, interruptions, questions, and natural transitions generally sound more believable than dividing a formal essay between two speakers.
Preserve Context Around Emotional Lines
A model may struggle to interpret a short line such as:
“Fine.”
That word could sound angry, relieved, disappointed, or satisfied.
Provide enough context through the surrounding dialogue, punctuation, or an audio tag:
[frustrated] Fine. We’ll try it your way.
ElevenLabs recommends natural narrative structure and clear emotional context when prompting expressive speech.
Remove Text That Should Not Be Spoken
Delete headings, production notes, citations, formatting instructions, and visual directions unless the narrator should read them aloud.
For example, do not leave this inside a normal Text to Speech script:
Show a screenshot of the dashboard here.
Store visual instructions separately or place only supported audio tags inside the narration.
Avoid Unusual Formatting
Decorative brackets, code fragments, unsupported symbols, and copied webpage formatting may produce unexpected speech. Clean the text before generating it.
Convert:
- Bulleted fragments into complete spoken sentences
- Tables into a clear explanation
- Footnotes into normal wording or remove them
- Abbreviations into their spoken forms
- Links into understandable verbal directions
- Visual headings into natural transitions
Test Difficult Sections First
Before generating the complete script, test the passage containing the most difficult names, numbers, abbreviations, foreign terms, or emotional delivery.
This helps determine whether the selected voice and model can handle the project before a larger amount of credits is used.
Complete a Script-Formatting Check
Before generating the final narration, confirm that:
- The script sounds natural when read aloud.
- Long sentences have been divided.
- Punctuation creates the intended rhythm.
- Paragraph breaks separate major ideas.
- Numbers and dates are written clearly.
- Acronyms are pronounced correctly.
- URLs and symbols have been rewritten for speech.
- Emotional tags are used only when needed.
- Production notes have been removed.
- Difficult sections have been tested separately.
A well-formatted script reduces the amount of adjustment required later. Once the wording and structure sound natural, the next step is learning how to adjust stability, similarity, style, and speaking speed.
Adjust Stability, Similarity, Style, and Speaking Speed
ElevenLabs’ voice settings control how consistent, expressive, similar, and fast the generated narration sounds. These settings can improve a voice, but extreme adjustments may also introduce unnatural pacing, distortion, mispronunciations, or unexpected sounds.
Beginners should start with the default settings and change only one control at a time. ElevenLabs identifies approximately 50 stability, 75 similarity, and 0 style as a common starting point for supported voices and models.
Open the Voice Settings
- Open Text to Speech.
- Select the voice you want to use.
- Open Voice Settings.
- Review the available controls.
- Generate a short test before making changes.
- Adjust one setting.
- Generate the same passage again.
- Compare the two versions.
The controls displayed can vary according to the selected model and voice. Settings such as similarity, style exaggeration, and Speaker Boost may not be available in every workflow.
Adjust Stability
Stability controls how consistent or variable the voice sounds between generations.
A lower stability setting allows more emotional variation and creative delivery. This may help with storytelling, advertisements, character dialogue, and expressive narration. However, very low stability can make the voice unpredictable, overly dramatic, unusually fast, or inconsistent.
A higher stability setting produces more controlled and repeatable narration. This may work better for tutorials, training videos, news-style content, and serious explanations. Setting it too high can make the voice sound flat or monotonous.
Start near the default value and make small adjustments.
Lower stability when:
- The narration sounds too flat.
- Emotional sentences lack energy.
- A character needs a more expressive performance.
Raise stability when:
- The voice changes tone unexpectedly.
- The speed varies between sentences.
- The delivery sounds chaotic.
- Several generations need to sound more consistent.
Even with identical settings, repeated generations may produce slightly different performances. The difference is usually more noticeable at lower stability values.
Use Eleven v3 Stability Modes
Eleven v3 may present stability through three simplified modes:
Creative allows the greatest emotional range and responds more strongly to expressive directions, but it also has the highest risk of unusual or unintended output.
Natural provides a more balanced performance that remains relatively close to the voice’s original character.
Robust prioritizes consistency and stability but responds less strongly to emotional instructions and audio tags.
Use Natural as the starting point for most v3 projects. Move toward Creative for dramatic narration and toward Robust when consistency matters more than expression.
Adjust Similarity
Similarity controls how closely ElevenLabs attempts to match the selected voice’s original characteristics.
Increasing similarity can help a cloned or community voice remain closer to its source. However, when the source recording contains background noise, echo, compression, or other imperfections, a very high similarity setting may reproduce some of those problems.
Raise similarity when:
- A cloned voice does not sound recognizable enough.
- The voice loses its expected accent or vocal character.
- Different passages sound too far removed from the original speaker.
Lower similarity when:
- The audio contains distortion or unwanted background texture.
- The clone exaggerates imperfections from the source recording.
- The result sounds less clear at higher values.
Similarity cannot repair a poor voice-cloning sample. When the source audio contains significant noise or inconsistent performance, creating a better recording may improve the result more than increasing the slider.
Use Speaker Boost Carefully
Speaker Boost attempts to increase resemblance to the original speaker. It requires additional processing and may slightly increase generation latency. ElevenLabs describes the effect as generally subtle.
Turn it on when a cloned voice needs a small improvement in similarity. Compare the result with and without Speaker Boost before applying it throughout a long project.
Do not assume it automatically improves every voice. If the difference is minimal or the generation becomes less reliable, leave it disabled.
Adjust Style Exaggeration
Style exaggeration amplifies characteristics from the original speaker’s performance. It may strengthen qualities such as enthusiasm, dramatic delivery, vocal rhythm, or character personality.
ElevenLabs recommends leaving this setting at 0 for most projects. Increasing it uses more processing, can add latency, and may make the output less stable. Potential problems include inconsistent speed, mispronunciations, unusual sounds, and exaggerated delivery.
Increase style only when:
- The selected voice has a clear performance style you want to emphasize.
- The default narration sounds too restrained.
- The project is creative or character-focused.
- You are prepared to generate and compare several attempts.
For tutorials, educational narration, and professional business content, leaving style at zero is usually the safer choice.
Adjust the Speaking Speed
The Speed control changes how quickly the voice reads the script.
The default is 1.0. Values below 1.0 slow the narration, while values above 1.0 make it faster. The supported range is currently 0.7 to 1.2. Extreme values may reduce speech quality.
A practical process is:
- Generate the passage at 1.0.
- Decide whether the problem is truly the speed.
- Make a small adjustment, such as 0.95 or 1.05.
- Generate the same passage again.
- Compare clarity and naturalness.
- Avoid moving immediately to the minimum or maximum.
Use a slower speed for detailed instructions, unfamiliar terminology, accessibility-focused narration, language-learning content, and important warnings.
A slightly faster speed may suit short advertisements, social-media narration, simple transitions, or energetic introductions.
Do Not Use Speed to Fix a Poor Script
When narration feels too slow, the script may contain unnecessary words. When it feels rushed, the sentences may be too long or complicated.
Before changing the speed significantly, try:
- Shortening repetitive wording
- Dividing long sentences
- Adding punctuation
- Creating another paragraph
- Removing unnecessary introductions
- Splitting the passage into separate generations
Rewriting the script often produces more natural pacing than forcing the voice to speak much faster or slower.
Adjust One Setting at a Time
Changing stability, similarity, style, and speed simultaneously makes it difficult to understand which adjustment improved or damaged the result.
Use the same short test script and follow this order:
- Confirm the voice and model.
- Adjust stability.
- Review similarity when available.
- Test Speaker Boost when needed.
- Leave style at zero unless the project requires more expression.
- Adjust speed last.
Download or label strong versions so they are not confused with weaker tests.
Use Different Settings for Special Sections
Most projects should use consistent settings, but an individual section may need a different delivery.
For example, a tutorial could use stable narration for the instructions and slightly lower stability for an energetic introduction. Inside ElevenCreative Studio, users can override voice settings for selected passages or apply changes across every paragraph using the same voice. Changing the voice or settings may require the affected audio to be generated again.
Use overrides sparingly. Large differences between sections can make the narrator sound like a different person.
Troubleshoot Common Problems
When the voice sounds monotonous, lower stability slightly, improve the script’s punctuation, or try a more expressive voice.
When it sounds too chaotic, increase stability and return style exaggeration to zero.
When a clone does not sound similar enough, increase similarity gradually, test Speaker Boost, and review the quality of the original recording.
When the voice speaks too quickly or slowly, adjust speed in small increments and simplify the script.
When the result contains strange sounds or mispronunciations, shorten the passage, reduce style exaggeration, adjust stability or similarity, and generate it again. ElevenLabs notes that voice type, text length, model choice, and voice settings can all affect these problems.
Save the Approved Settings
After finding a combination that works, record:
- Voice name or ID
- Speech model
- Stability
- Similarity
- Speaker Boost status
- Style exaggeration
- Speaking speed
Keep these details with the project files. AI generations will not be perfectly identical, but reusing the same voice, model, and settings improves consistency across future sections.
Complete a Voice-Settings Check
Before generating the complete script, confirm that:
- The voice is expressive without becoming unpredictable.
- The similarity setting does not reproduce unwanted noise.
- Style exaggeration is being used only when necessary.
- The speed remains natural and understandable.
- Speaker Boost provides a meaningful improvement when enabled.
- The same settings work across several different sentences.
- A short test has been downloaded for comparison.
Voice settings should refine a suitable voice rather than rescue the wrong one. When extensive adjustments still produce weak results, selecting another voice or model may be more effective.
Correct Pronunciation and Pacing Problems
Even a strong ElevenLabs voice may mispronounce names, acronyms, technical terms, numbers, or words from another language. Narration can also contain pauses that feel too long, sentences that move too quickly, or changes in rhythm between separate generations.
Most problems can be improved by correcting the script first, then testing pronunciation tools, punctuation, voice selection, and speed settings.
Proofread the Script First
Check that every word is spelled correctly before changing the voice settings. ElevenLabs generally tries to pronounce the text exactly as written rather than automatically correcting misspellings.
Pay particular attention to:
- Names of people and places
- Company and product names
- Technical terminology
- Acronyms
- Foreign words
- Dates, prices, and measurements
- Website and email addresses
A small spelling error may sound like a problem with the voice when the actual issue is the script.
Test the Difficult Word in a Sentence
Do not test only the isolated word. Place it inside a natural sentence so the model has enough context to interpret its language, meaning, and emphasis.
Instead of generating:
“Vaxali.”
Test:
“Visit Vaxali to discover beginner-friendly guides to the latest AI tools.”
The surrounding sentence can help ElevenLabs choose a more natural rhythm and pronunciation.
Rewrite the Word Phonetically
The easiest beginner solution is often to spell the word according to how it should sound.
For example:
AI→A.I.FAQ→F.A.Q.SQL→sequelorS.Q.L., depending on the intended pronunciationVaxali→ a phonetic version that matches the brand’s official pronunciation
Try using spaces, periods, dashes, apostrophes, or capitalization to emphasize individual sounds. ElevenLabs also recommends phonetic substitutions or alias rules when a model does not support direct phoneme instructions.
Keep the visible spelling separate from the spoken version when the official word will also appear in captions or on-screen text.
Expand Acronyms
When an acronym continues to sound incorrect, replace it with the full phrase.
For example:
UN→United NationsCEO→chief executive officerCRM→customer relationship managementAPI→application programming interface
A pronunciation dictionary can automate this replacement across a longer project by connecting the acronym with an alias.
Write Numbers as Words
Numbers can be interpreted differently depending on the model, language, and context. ElevenLabs enables text normalization by default in its website Text to Speech workflow, but important figures should still be reviewed carefully.
When accuracy matters, write the exact spoken form:
2026→twenty twenty-six$1,399→one thousand three hundred ninety-nine dollars4.7%→four point seven percent7/28/2026→July twenty-eighth, twenty twenty-six610-555-0184→six one zero, five five five, zero one eight four
This is especially important for prices, financial information, phone numbers, addresses, statistics, and measurements.
Rewrite Websites and Email Addresses
A narrator may not know whether a period, slash, hyphen, or domain ending should be spoken literally.
Instead of:
Visit vaxali.com/guides.
Write:
Visit vaxali dot com slash guides.
ElevenLabs specifically recommends converting URLs into the wording that should be spoken.
For an email address, write each part clearly:
support at vaxali dot com
Generate a short test because some voices may still read letters or domain endings differently.
Use the Correct Voice for the Accent
The written script determines the language, while the selected voice strongly influences the accent. A voice that is not native to the target language may keep its original accent or drift between accents.
When pronunciation remains weak:
- Open the Voice Library.
- Filter by the target language.
- Select the intended regional accent.
- Generate the difficult sentence again.
- Compare it with the original voice.
For example, use a voice trained in Spanish from Spain when the script is intended for a Spanish audience rather than relying on an English voice to read translated Spanish text.
Use a Pronunciation Dictionary
Pronunciation dictionaries create reusable rules for words that appear repeatedly. They are useful for brand names, employee names, locations, technical terminology, and acronyms.
A dictionary can use:
- Alias rules, which replace the written term with another spelling or phrase
- Phoneme rules, which describe the exact pronunciation using a phonetic alphabet
For example, an alias can instruct ElevenLabs to pronounce UN as United Nations every time it appears. Pronunciation dictionaries can be connected to supported Studio, Dubbing Studio, agent, and API workflows.
Add a Rule in Studio
Inside a supported ElevenCreative Studio project:
- Open the project settings.
- Find the pronunciation controls or Pronunciations Editor.
- Add the word that is being mispronounced.
- Choose a phonetic pronunciation or word substitution.
- Save the rule to the project’s pronunciation dictionary.
- Regenerate the affected narration.
- Listen to every sentence containing the word.
Adding or changing a dictionary can mark affected passages for regeneration because the existing audio no longer follows the updated rule.
Use Alias Rules for Simple Corrections
Alias rules are easier for beginners because they replace a difficult word with ordinary written language.
For example:
Claughton→ClofftonUN→United NationsAI→A.I.- A brand name → a spelling that produces the correct sound
Aliases are also useful with models such as Multilingual v2 that do not support the same direct phoneme controls as certain other models.
Use IPA With Eleven v3
Eleven v3 can interpret International Phonetic Alphabet notation written directly between forward slashes. This provides more precise pronunciation control across supported languages.
The format is:
/IPA pronunciation/
IPA is powerful, but it requires accurate symbols and stress marks. A small transcription mistake may create a worse result than the original pronunciation.
When using IPA:
- Verify the transcription with a reliable pronunciation reference.
- Include stress markers for multisyllable words.
- Test the pronunciation with the selected voice.
- Generate several versions when consistency matters.
ElevenLabs notes that the same IPA transcription may occasionally produce slightly different results, so important lines should still be reviewed manually.
Understand Model Compatibility
Pronunciation controls do not work identically across every speech model.
ElevenLabs’ current documentation states that pronunciation-dictionary phoneme tags work with Eleven Flash v2 and Eleven v3. Other models may ignore those phoneme rules and require aliases or phonetic rewriting instead. IPA or CMU-based pronunciation outside English requires Eleven v3 in supported dictionary workflows.
When a rule appears to have no effect, confirm that the selected model supports that type of instruction before repeatedly regenerating the audio.
Create Separate Rules for Capitalization
Pronunciation-dictionary matching can be case-sensitive. A rule for Vaxali may not automatically apply to vaxali or VAXALI.
Add separate versions when the word appears with different capitalization:
VaxalivaxaliVAXALI
Keep the dictionary organized and document why each rule exists.
Correct a Voice That Speaks Too Quickly
First determine whether the entire voice is too fast or only one section feels rushed.
When the complete narration is too quick:
- Open the voice settings.
- Lower the speed slightly.
- Test a value such as 0.95.
- Generate the same passage again.
- Continue in small increments when necessary.
ElevenLabs currently supports speed values from 0.7 to 1.2, with 1.0 representing the original pace. Extreme settings can reduce speech quality.
When only one sentence sounds rushed, rewrite or divide it instead of slowing the complete project.
Correct a Voice That Speaks Too Slowly
Before increasing the speed, remove unnecessary wording and long pauses.
For example:
Slow and repetitive:
“At this particular point in the process, what you are going to want to do next is select the settings button.”
Improved:
“Next, select the Settings button.”
When the script is already concise, increase the speed slightly, such as from 1.0 to 1.05. Large increases may make consonants less clear and reduce the natural quality of the voice.
Add a Natural Pause With Punctuation
Use a period or paragraph break for a normal stop. Use a comma for a shorter pause.
For example:
“Your voiceover is ready. Next, download the audio.”
A dash can create a short interruption:
“You can regenerate the line — but review the credit cost first.”
An ellipsis can create hesitation:
“I thought the file was gone… but it was still in History.”
Ellipses may also make the narrator sound uncertain or nervous, so they should not be used for every ordinary pause.
Add a Precise Pause With a Break Tag
Supported Text to Speech workflows can use an SSML break tag:
<break time="1.5s" />
For example:
Your first section is complete. <break time="1.0s" /> Now, let’s create the next one.
ElevenLabs supports break durations of up to three seconds. Excessive break tags can cause faster speech, additional noise, or other audio artifacts, so use them only when normal punctuation does not provide enough control.
Control Pauses in Eleven v3
Eleven v3 does not use SSML break tags in the same way as earlier workflows. Use punctuation, paragraph structure, or expressive pause tags instead.
Examples include:
[pause][short pause][long pause]
For example:
This setting changes the voice across the entire project. [short pause] Review it before continuing.
Test the result because different voices may interpret the same pause instruction differently.
Correct Unnatural Pauses
When the narrator pauses in the wrong location:
- Remove unnecessary commas.
- Check for copied formatting or invisible paragraph breaks.
- Divide the sentence differently.
- Replace parentheses with normal spoken wording.
- Generate the sentence separately.
- Test another voice when the problem continues.
A sentence containing several clauses may create pauses that are technically correct but still sound unnatural. Rewriting it into two sentences usually provides more control.
Match Pacing Between Separate Clips
Separately generated sections may sound slightly different, even when they use the same voice and settings.
To improve consistency:
- Use the same voice and model.
- Keep stability, similarity, style, and speed unchanged.
- Generate sections of similar length.
- Include enough surrounding context.
- Avoid switching between different emotional styles.
- Compare the ending of one clip with the beginning of the next.
When replacing one line, include a sentence before or after it if possible. Additional context can help the model match the surrounding tone more naturally.
Correct Whispering, Accent Changes, or a Breaking Voice
Long or unstable generations may occasionally begin whispering, shift accents, change tone, or lose clarity.
When this happens:
- Divide the script into shorter sections.
- Regenerate only the affected paragraph.
- Increase stability slightly.
- Return style exaggeration to zero.
- Confirm that the voice matches the script’s language.
- Test another speech model.
- Check whether copied symbols or formatting appear near the problem.
ElevenLabs recommends paragraph-based workflows because they reduce the need to regenerate an entire long passage when one section develops a problem.
Compare Several Generations
Some pronunciation and pacing differences are caused by the natural variability of AI generation rather than a clear script error.
When the settings appear correct:
- Generate the line again.
- Compare both versions.
- Save the stronger result.
- Avoid changing several controls unless the problem repeats consistently.
This is particularly useful with Eleven v3, where emotional and conversational delivery can vary more noticeably.
Review the Audio in Context
A sentence may sound acceptable by itself but feel too fast, too emotional, or too quiet when placed between other clips.
Listen to:
- The sentence before it
- The corrected sentence
- The sentence after it
Check whether the transition sounds natural and whether the speaker’s energy remains consistent.
Complete a Pronunciation and Pacing Check
Before approving the narration, confirm that:
- Names and technical terms are pronounced correctly.
- Acronyms are read in the intended form.
- Prices, dates, and numbers are unambiguous.
- The accent matches the target language.
- Pauses appear in natural locations.
- No break tags are being overused.
- The speaking speed remains comfortable.
- Separate clips sound consistent.
- Pronunciation rules have been tested with the selected model.
- The approved version has been downloaded or saved.
Pronunciation tools can improve difficult words, but script clarity and voice selection remain the most important starting points. Once the narration sounds accurate and well paced, the next step is browsing and saving additional voices from the Voice Library.
Browse and Save Voices From the Voice Library
The ElevenLabs Voice Library is a marketplace containing community-shared Professional Voice Clones. It gives users access to more than 10,000 voices for narration, education, advertisements, entertainment, character work, social media, and conversational projects.
Open the Voice Library
- Sign in to ElevenLabs.
- Select Voices from the sidebar.
- Open Explore.
- Browse the featured collections or search for a specific voice.
- Play voice samples before selecting one.
- Click Use Voice to open Text to Speech with that voice selected.
Users can generate with a Voice Library voice without saving it first. Selecting Use Voice sends it directly to the Text to Speech workspace.
Browse the Handpicked Collections
ElevenLabs organizes selected voices into collections based on language, genre, use case, and vocal style. These collections can help beginners find strong options without searching the complete library individually.
A collection may focus on:
- Narrators
- Educational voices
- Conversational speakers
- Character voices
- Social-media narration
- Advertisements
- Entertainment
The collections are updated as ElevenLabs adds and reviews new community voices.
Search by Name or Keyword
Use the search bar when you already know the voice name or need a particular style.
Useful searches might include:
- Warm female narrator
- Spanish educational voice
- Calm audiobook narrator
- Energetic social-media voice
- Professional American male
- British documentary narrator
- Conversational podcast voice
Searches can also use a voice ID when one has been provided by another creator or developer.
Search With an Audio Sample
ElevenLabs allows users to upload or drag an audio recording into the Voice Library search. The platform attempts to identify the original public voice when available and may also suggest similar voices.
This can be useful when you hear a voice in an existing ElevenLabs-generated clip but do not know its name.
The search examines public Voice Library options. It cannot identify private custom voices belonging to other users.
Filter by Language and Accent
Choose the target language first. Once a language has been selected, the Accent filter becomes available for supported regional variations.
Voices tagged with a particular language generally perform best in that language, even though ElevenLabs voices can often generate speech in other supported languages.
For example, a Spanish project may require:
- Spanish from Spain
- Mexican Spanish
- American Spanish
- Another supported regional accent
Test the real script before approving the voice because a language tag does not guarantee perfect pronunciation for every name, phrase, or regional expression.
Filter by Category
The Category filter helps match a voice to its intended use.
Current categories include:
- Conversational
- Narration
- Characters
- Social Media
- Educational
- Advertisement
- Entertainment
A voice tagged for narration may be more comfortable during a long tutorial, while a social-media voice may sound stronger in a short and energetic video.
Filter by Gender and Age
Users can filter voices by available gender and age labels.
Current gender filters include:
- Male
- Female
- Neutral
Age filters include:
- Young
- Middle Aged
- Old
These labels are useful for narrowing the library, but the audio preview is more important than the category. Two voices with the same labels can sound very different in tone, energy, clarity, and emotional range.
Filter for Studio Quality
Select Studio Quality when professional recording quality is important.
ElevenLabs uses this label for voices that have been recorded with suitable equipment, mixed properly, and reviewed for common problems such as echo, reverb, distortion, and other audio artifacts.
Studio Quality voices are a useful starting point for:
- Audiobooks
- Paid courses
- Client projects
- Professional YouTube narration
- Business presentations
- Long-form educational content
The label does not guarantee that the voice will suit every project, so the actual script should still be tested.
Sort the Results
The Voice Library can be sorted by:
- Trending
- Latest
- Most users
- Character usage
Trending and most-used voices may provide reliable starting points, but popular voices may also appear across many other creators’ videos. A less common voice may help a brand or series sound more distinctive.
Preview the Voice
Click a voice to hear its sample. When several language previews are available, choose the language you intend to use from the preview player.
Do not approve a narrator based only on the default demonstration. Generate a short section of the real script containing:
- The project’s main terminology
- Names and locations
- Numbers
- Questions
- Long sentences
- Emotional wording
A voice may sound excellent in its prepared sample but perform differently with the actual content.
Save a Voice to My Voices
Saving is optional, but it makes a voice easier to locate later.
- Find the voice in the Voice Library.
- Click the + button.
- Open My Voices.
- Confirm that the saved voice appears.
- Select the T button when you want to open Text to Speech with it selected.
Saved Voice Library voices become available in the voice-selection menus throughout ElevenLabs. They do not use the custom voice slots reserved for voices created through Voice Design or Voice Cloning.
Understand My Voices
The My Voices area contains:
- Voices saved from the Voice Library
- Voices created through Voice Design
- Instant Voice Clones
- Professional Voice Clones
Each voice may display its training language, suggested category, voice type, and notice period when one applies.
Voice-type icons may include:
- A yellow tick for a Professional Voice Clone
- A black tick for a Studio Quality Professional Voice Clone
- A lightning icon for an Instant Voice Clone
- No icon for a Voice Design voice
Rename a Saved Voice
Open the voice’s three-dot menu and select Edit Voice when you want to change its name or description inside your own account.
For example, rename a voice from its public name to:
- Main Vaxali Narrator
- Spanish Guide Voice
- Calm Course Narrator
- Short-Form Social Voice
These personal changes do not alter the public Voice Library listing.
View Previous Generations
Open the three-dot menu beside a voice and select View History to find earlier Text to Speech generations made with that narrator.
This is useful when you need to:
- Download an approved version again
- Compare earlier settings
- Confirm which voice was used
- Avoid spending credits regenerating the same script
Save important audio outside ElevenLabs as well, since a community voice may not remain available forever.
Check the Notice Period
A notice period tells users how long they can continue using a voice after its owner decides to remove it from the Voice Library.
Current notice periods range from three months to two years. When a voice with a notice period is withdrawn, users who previously saved or used it receive email and in-app notifications and retain access until that period ends.
When a voice has no notice period, access can be lost immediately after the owner stops sharing it.
Use the Notice Period filter when selecting a voice for:
- A recurring video series
- An audiobook collection
- A long-term course
- A game receiving future updates
- Ongoing brand narration
A longer notice period provides more time to replace the voice when necessary, but it does not guarantee permanent access.
Understand When the Notice Period Is Saved
You do not need to add the voice to My Voices for its notice period to be recorded. Using the voice at least once saves its current notice period to the account.
Saving the voice is still helpful for organization and faster access.
Check for a Credit Multiplier
Some older community voices have custom credit multipliers. A voice with a 2x multiplier uses twice as many credits as a standard-rate voice for the same script.
The multiplier appears as a label in the Voice Library and as a warning in the speech-generation interface. These custom rates are a legacy feature and cannot be added to newly shared voices.
Voices with a credit multiplier are not available on the Free plan. Test the credit cost before using one for a long project.
Check Live Moderation
Some voices have Live Moderation enabled and display a shield label. When these voices are used, ElevenLabs checks the generated text against additional prohibited-content categories.
This extra processing may increase generation time. The Voice Library provides a filter for excluding voices with Live Moderation when that distinction matters to the workflow.
All ElevenLabs content remains subject to the platform’s broader safety and usage policies regardless of whether a particular voice has this additional label.
Remove a Saved Voice Carefully
To remove a voice:
- Open My Voices.
- Find the voice.
- Open its three-dot menu.
- Select Delete Voice.
- Confirm the deletion.
Deleting a voice from My Voices is permanent. When the voice has already been removed from the public Voice Library, it cannot be saved again after deletion.
Do not remove a voice that is still needed for an active project, especially when its notice period is already running.
Avoid Depending on One Community Voice
Community voices provide variety, but their owners control whether they remain publicly shared. For long-term production, keep a backup option.
A practical approach is to:
- Choose the preferred community voice.
- Save a second similar voice.
- Record the voice name and ID.
- Keep the voice and model settings in the project notes.
- Save every approved audio file locally.
For a permanent personal or brand narrator, consider creating an authorized Voice Design voice or voice clone rather than relying entirely on a community option.
Complete a Voice-Library Check
Before choosing a community voice for the complete project, confirm that:
- The language and accent match the audience.
- The category fits the content.
- The voice sounds natural with the real script.
- Its recording quality is acceptable.
- The credit multiplier fits the budget.
- The notice period is suitable for the project’s lifespan.
- A backup voice has been identified.
- The voice has been saved to My Voices.
- The voice name, ID, model, and settings have been recorded.
The Voice Library makes it easier to find a narrator without creating one from the beginning. Once a suitable voice has been saved, the next step is learning how to create an original synthetic voice with Voice Design.
Create an Original Voice With Voice Design
Voice Design creates a completely new synthetic voice from a written description. It is useful when the Voice Library does not contain the exact accent, age, tone, personality, or speaking style required for a project. Unlike voice cloning, Voice Design does not require a recording of a real speaker.
Open Voice Design
- Select Voices from the ElevenLabs sidebar.
- Open My Voices.
- Click Add a New Voice.
- Select Voice Design.
- Enter a description of the voice.
- Add or generate preview text.
- Adjust the available controls.
- Click Generate Voice.
- Listen to the three generated options.
- Save the strongest version.
Voice Design is available to all users, including Free-plan accounts. A saved Voice Design voice occupies one of the account’s custom voice slots.
Describe the Voice Clearly
The prompt tells ElevenLabs what kind of speaker to create. Include the characteristics that matter most to the project, such as:
- Language and regional accent
- Gender or vocal presentation
- Approximate age
- Pitch and vocal texture
- Personality
- Emotional tone
- Speaking pace
- Delivery style
- Recording quality
More detailed prompts generally provide greater control, although a simple description can work for a neutral narrator.
For example:
“Native American English. Female, early 30s. Studio-quality audio. Warm and trustworthy educational narrator with a smooth, clear tone. Relaxed conversational pacing with gentle emphasis and confident delivery.”
This gives ElevenLabs more useful direction than:
“Professional female voice.”
Use a Consistent Prompt Structure
ElevenLabs recommends organizing the description in a predictable order:
- Language and regional variety
- Gender and age range
- Desired audio quality
- Brief persona
- Emotional qualities
- Timbre, pacing, and delivery
A practical format is:
“Native [language and region]. [Gender], [age range]. [Audio quality]. Persona: [brief role]. Emotion: [key qualities]. [Description of tone, rhythm, and delivery].”
For example:
“Native English, neutral American accent. Male, 40–50. Excellent audio quality. Persona: experienced technology instructor. Emotion: patient, confident, approachable. Deep but smooth timbre with steady pacing and clear emphasis on important instructions.”
Specify the Language First
State the language and regional variety at the beginning of the prompt. This can reduce accent drift and help ElevenLabs understand whether the voice should sound American, British, European Spanish, Mexican Spanish, or another supported variation.
Avoid vague descriptions such as:
“Foreign accent.”
Use something specific:
“Native Hungarian speaker using clear, neutral Hungarian pronunciation.”
or:
“Native Spanish speaker from Spain, without Latin American pronunciation.”
Also avoid using the word accent when you actually mean emphasis or intonation. ElevenLabs warns that this can cause an unintended regional dialect change.
Describe the Vocal Texture
Tone or timbre describes the physical quality of the voice rather than its emotion.
Useful descriptions include:
- Warm
- Smooth
- Deep
- Rich
- Gravelly
- Raspy
- Breathy
- Airy
- Resonant
- Light
- Mellow
- Low-pitched
- High-pitched
For example:
“A low-pitched female voice with a warm, slightly husky texture.”
This provides more useful direction than describing only the speaker’s gender.
Describe the Pacing
Explain how quickly and rhythmically the speaker should deliver the script.
Possible instructions include:
- Relaxed and conversational
- Slow and deliberate
- Fast and energetic
- Even and consistent
- Measured with clear pauses
- Rhythmic and expressive
- Quick with occasional dramatic pauses
For an instructional narrator, try:
“Steady, measured pacing with short pauses after important information.”
For a social-media voice, try:
“Fast, energetic pacing with clear articulation and strong emphasis.”
Pacing instructions influence the personality of the generated voice, so they should match the intended use.
Include the Desired Audio Quality
Add a quality description when the voice should sound clean and professionally recorded.
Examples include:
- Good audio quality
- Excellent audio quality
- Studio-quality recording
- Broadcast-quality audio
- Clean, noise-free signal
ElevenLabs notes that quality instructions can improve clarity, although an extremely detailed or unusual prompt may sometimes reduce how accurately every other characteristic is followed.
Avoid words such as echo, reverb, telephone, or tape unless low-quality or stylized audio is intentional. These terms can cause unwanted noise or processing effects.
Add Preview Text
The preview text is the passage the generated voices will read. It should reflect the voice’s intended purpose and emotional style.
For an educational narrator, use something like:
“Welcome to the course. In this lesson, we’ll review the three steps required to complete your account setup safely and correctly.”
For a dramatic character:
“I warned them not to enter the forest after dark. But no one listened.”
The preview should not contradict the description. A calm, reflective voice may sound inconsistent when tested with an angry or highly energetic passage.
Use Enough Preview Context
A full sentence or short paragraph generally provides a more useful preview than one or two isolated words. Longer preview text gives the model more context for demonstrating pacing, emotion, pronunciation, and vocal texture.
Include content similar to what the finished voice will normally read. When designing a technical narrator, test it with technical language. When designing a fictional character, use dialogue that reflects the character’s personality.
The voice description can contain between 20 and 1,000 characters, while optional custom preview text can contain between 100 and 1,000 characters.
Generate the Voice Options
Click Generate Voice after completing the prompt and preview text. ElevenLabs creates three different voice options from the same description.
Listen to all three before choosing one. Compare:
- Accent
- Vocal age
- Tone and texture
- Clarity
- Speaking speed
- Emotional delivery
- Pronunciation
- Background artifacts
The three options may interpret the same description differently, so the first result is not automatically the best one.
Understand the Credit Cost
Voice Design consumes credits when the voice options are generated. ElevenLabs charges according to the number of characters in the preview text, but the account is charged only once even though three previews are produced.
Shorter preview text costs fewer credits, but it may not demonstrate the voice as accurately. Use enough text to evaluate the voice without generating an unnecessarily long passage.
Adjust Loudness
The Loudness control changes the volume of the generated preview and the saved voice.
Keep it near a natural level unless the voice is clearly too quiet or too loud. An excessively loud setting may make the recording sound harsh, while a low setting may require additional volume adjustment during editing.
Adjust the Guidance Scale
Guidance Scale controls how strictly ElevenLabs follows the written voice description.
A higher value keeps the result closer to the prompt but may reduce audio quality when the requested voice is highly specific or unusual.
A lower value gives the model more freedom and may produce smoother audio, but some requested characteristics may become less accurate.
Use a higher setting when the accent or vocal character is essential. Use a lower setting when natural performance and clean audio matter more than matching every prompt detail exactly.
Make small adjustments rather than moving immediately to the highest value.
Refine the Prompt
When none of the three voices match the project, revise the description instead of repeatedly generating the same prompt.
If the voice sounds too young, specify an older age range.
If the accent is incorrect, place the exact language and region at the beginning.
If the delivery is too energetic, request a calm, measured pace.
If the voice sounds artificial, simplify the prompt and request studio-quality, natural speech.
Change one or two characteristics at a time so it is easier to identify what improves the result.
Save the Preferred Voice
After selecting the best preview:
- Choose the preferred generation.
- Enter a recognizable voice name.
- Add an optional description.
- Save the voice.
- Open My Voices to confirm that it appears.
- Test it with a longer passage in Text to Speech.
Saving a Voice Design result uses one custom voice slot. Voices saved from the community Voice Library do not use these slots, but voices created through Voice Design or Voice Cloning do.
Understand Voice-Slot Limits
The number of custom voices that can remain saved depends on the plan. ElevenLabs currently lists:
- Free: 3 custom voice slots
- Starter: 10
- Creator: 30
- Pro: 160
- Scale: 660
- Business: 660
Deleting an unused custom voice frees a slot, but deletion cannot be undone.
Do not delete a voice that is still used in an active project. Save important audio and record the voice settings before removing anything.
Test the Saved Voice
After saving the voice, open Text to Speech and generate several different passages.
Test:
- A normal informational sentence
- A question
- A sentence containing names
- Numbers and dates
- A longer paragraph
- Emotional wording
- The target language and accent
A preview may sound impressive while the saved voice performs less consistently across other types of content. Testing several passages helps confirm whether it is suitable for a complete project.
Understand the Quality Limits
ElevenLabs describes Voice Design as an experimental tool intended for quick exploration and iteration. Results can vary according to the prompt and use case.
Professional Voice Clones generally provide greater consistency and production-ready quality when a suitable verified voice is available. However, Voice Design remains useful when the project needs an original synthetic narrator or fictional character rather than a reproduction of a real speaker.
Do Not Use Voice Design to Imitate Someone
Voice Design should be used to create an original voice, not to reproduce a celebrity, public figure, employee, family member, or another identifiable person.
When the goal is to use a real person’s voice, use ElevenLabs’ authorized voice-cloning workflow and obtain the required permission. Voice Design is better for creating a general narrator, fictional character, virtual assistant, or new brand voice.
Understand Sharing Restrictions
Voice Design voices remain inside the creator’s account and currently cannot be shared with other users. Professional Voice Clones have different sharing options, but synthetic Voice Design voices are private account assets.
This may be limiting for teams that want several members to use the same designed voice. Confirm the account and workspace arrangement before building a major production around it.
Save the Voice Details
Record the following information in the project notes:
- Voice name
- Original prompt
- Preview text
- Loudness setting
- Guidance Scale
- Intended language
- Intended use
- Preferred speech model
- Approved Text to Speech settings
Keeping the original description makes it easier to understand the voice’s intended identity and create a similar replacement if the saved voice is later deleted.
Complete a Voice Design Check
Before using the voice throughout a project, confirm that:
- The language and accent are correct.
- The apparent age fits the intended speaker.
- The tone matches the project.
- The pacing remains comfortable.
- The audio sounds clean.
- Important words are pronounced correctly.
- The voice performs consistently across several scripts.
- The credit cost and voice-slot use are acceptable.
- The prompt and settings have been saved.
Voice Design is most useful when you need an original synthetic speaker that does not already exist in the Voice Library. Once the designed voice has been tested and approved, the next step is learning how to create an authorized clone of your own voice.
Clone Your Own Voice With Permission
Voice cloning allows ElevenLabs to reproduce the vocal characteristics of a real speaker, including their tone, accent, rhythm, and speaking style. ElevenLabs offers two cloning methods: Instant Voice Cloning for fast results from a short recording and Professional Voice Cloning for a more accurate model trained with a much larger collection of audio.
Only clone a voice that you own or have clear permission to use. Professional Voice Cloning has stricter rules: you can create a Professional Voice Clone only of your own voice, even when another person has given permission. That person must create and verify the Professional Voice Clone through their own ElevenLabs account before sharing it with you.
Choose Between Instant and Professional Voice Cloning
Use Instant Voice Cloning when:
- You want a clone that is ready almost immediately.
- You have approximately one to two minutes of clean audio.
- You are testing whether voice cloning works for your project.
- You need a narrator for general videos, courses, or personal content.
- The speaker does not have a particularly unusual voice or rare accent.
Instant Voice Cloning does not train a dedicated model on the speaker. It uses the uploaded audio to estimate the voice based on patterns ElevenLabs has already learned. This makes the process fast, but it may be less accurate for unusual voices or accents.
Use Professional Voice Cloning when:
- You need a more accurate reproduction of your own voice.
- The voice will be used regularly or professionally.
- Your accent or vocal characteristics are difficult for Instant Voice Cloning to reproduce.
- You can provide at least 30 minutes of clean recordings.
- You are willing to complete voice verification and wait for training.
Professional Voice Cloning trains a dedicated model using a larger collection of recordings. It is designed to reproduce the source voice more accurately and consistently than an Instant Voice Clone.
Check Your Plan
Instant Voice Cloning is available on the Starter plan and above. It is not included with the Free plan.
Professional Voice Cloning requires the Creator plan or above. Current base Professional Voice Clone limits include:
- Free and Starter: no Professional Voice Clone slots
- Creator and Pro: one slot
- Scale: three slots
- Business: ten slots
- Enterprise: a custom number
Plan limits can change, so review the current subscription page before upgrading specifically for voice cloning.
Record a Clean Voice Sample
The source recording has the greatest effect on the quality of the clone. A clean one-minute recording can create a better Instant Voice Clone than several minutes of inconsistent or noisy audio.
Record in a quiet room with:
- One speaker only
- Minimal background noise
- Little or no room echo
- A consistent microphone distance
- Clear speech
- Consistent volume
- A natural but steady speaking style
- No music or sound effects
- No extremely long pauses
ElevenLabs recommends one to two minutes of good audio for Instant Voice Cloning and approximately 30 to 180 minutes for Professional Voice Cloning. Audio quality is more important than the number of separate files.
Use a Consistent Microphone
Record the complete sample with the same microphone and settings whenever possible.
Avoid combining:
- Phone recordings with studio recordings
- Quiet narration with shouted dialogue
- Different rooms with noticeably different echo
- Recordings made from different distances
- Audio with changing background noise
- Clips that have received very different processing
A clone can reproduce imperfections contained in the source material. Noise, distortion, echo, compression, and inconsistent volume may therefore appear in the generated speech.
Keep the Speaking Style Consistent
The clone learns more than the physical sound of the voice. It also reflects the speaking style present in the recordings.
For example, use:
- Calm educational narration for a tutorial voice
- Conversational speech for a podcast voice
- Clear storytelling for an audiobook voice
- Energetic delivery for advertising narration
Do not combine several completely different performances in one dataset unless that variation is intentional. ElevenLabs recommends keeping the style consistent because the delivery contained in the source recordings influences the generated output.
Prepare an Instant Voice Clone Recording
For a beginner’s first Instant Voice Clone, record between one and two minutes of continuous speech.
The recording should contain:
- Complete sentences
- Different vowel and consonant sounds
- Questions and statements
- Short and longer sentences
- Natural pauses
- The speaker’s normal accent
- The vocal tone you want the clone to reproduce
Do not deliberately change your voice while recording. Speak naturally and use the delivery you expect to use most often.
ElevenLabs notes that adding more than two or three minutes to an Instant Voice Clone usually provides little improvement and can sometimes reduce stability.
Choose a File Format
ElevenLabs accepts several audio-file types for voice cloning. Its current recommendation is an MP3 file with a bitrate of at least 192 kilobits per second.
An uncompressed WAV file does not necessarily create a noticeably better clone. A clean recording matters more than using the largest possible file format.
Keep the original recording after uploading it. ElevenLabs does not allow users to export a voice clone itself, so the source samples are needed when you want to recreate the voice later.
Create an Instant Voice Clone
To create the clone:
- Open Voices from the ElevenLabs dashboard.
- Select My Voices or the personal voice area.
- Click the plus icon or Create Voice.
- Choose Instant Voice Clone.
- Upload the clean recording or record directly through the browser.
- Enter a recognizable name.
- Add an optional description.
- Confirm that you have the required permission to clone the voice.
- Complete the creation process.
- Open the clone in Text to Speech.
Use a name that explains the purpose, such as:
- Mildred Tutorial Voice
- English Course Narrator
- Spanish Product Voice
- Calm YouTube Narrator
A clear name becomes especially useful when the account contains several custom voices.
Test the Instant Voice Clone
Do not begin with the complete project. Generate a short test containing different types of speech.
For example:
“Welcome to today’s tutorial. First, we’ll open the dashboard and review the main settings. Does everything look correct? Great. Let’s continue.”
Review:
- Voice similarity
- Accent
- Clarity
- Speaking speed
- Emotional delivery
- Pronunciation
- Volume
- Consistency between sentences
Create a second test containing names, numbers, and terminology from the real project.
Adjust the Text to Speech Settings
After creating the clone, open its voice settings and begin near the standard defaults.
Similarity controls how closely the output follows the uploaded voice. Increasing it may improve resemblance, but it can also reproduce unwanted noise or imperfections from the original recording.
Stability controls how consistent or expressive the performance sounds. Higher stability generally creates more controlled narration, while lower stability allows more variation.
Adjust only one setting at a time. When the clone still sounds inaccurate after several small adjustments, improve the recording rather than forcing the settings to compensate for weak source audio.
Correct an Inaccurate Instant Clone
When the clone does not sound similar enough:
- Confirm that only one speaker appears in the recording.
- Remove clips containing background noise or echo.
- Use a recording with consistent volume.
- Make sure the speaker uses their normal accent.
- Remove sections containing an unusual or exaggerated performance.
- Create a new clone with the improved sample.
- Test another speech model.
A clone’s accent and tone are largely determined by its source samples. ElevenLabs states that the main way to change an inaccurate accent or tone is to create the clone again using different recordings.
Prepare Audio for a Professional Voice Clone
Professional Voice Cloning requires significantly more material.
ElevenLabs recommends:
- At least 30 minutes of high-quality speech
- Preferably two or more hours for stronger results
- Up to approximately three hours of useful training audio
- One speaker throughout
- Consistent recording quality
- Consistent delivery and tone
- Separate files of roughly 30 minutes when uploading several hours
Adding low-quality audio only to increase the total length can weaken the clone. Use the clearest and most representative recordings available.
Professional Voice Clones currently support spoken audio rather than singing samples.
Create a Professional Voice Clone
To begin:
- Open Voices.
- Select Create Voice.
- Choose Professional Voice Clone.
- Name the voice.
- Select the primary language when requested.
- Upload the prepared recordings.
- Review the platform’s feedback about the recording length.
- Process the uploaded audio.
- Complete the voice-verification step.
- Submit the clone for fine-tuning.
Do not begin the process using another person’s recordings. Professional Voice Cloning requires the person creating the clone to verify that the uploaded voice is their own.
Complete Voice Verification
ElevenLabs asks the voice owner to record a verification sample through the microphone. This confirms that the person creating the Professional Voice Clone matches the speaker in the uploaded recordings.
For the best chance of successful verification:
- Allow the browser to access the microphone.
- Use the same or similar microphone used for the original recordings.
- Find a quiet location.
- Speak in a similar tone and style.
- Avoid background audio.
- Confirm that the microphone is not muted.
When all verification attempts fail, ElevenLabs currently requires users to wait 24 hours before trying again or contact its support team for assistance.
Wait for Fine-Tuning to Finish
Professional Voice Clones are not ready immediately. After verification, ElevenLabs places the voice in a training queue and fine-tunes the model using the uploaded recordings.
The process commonly takes several hours but may take longer depending on the dataset, processing status, and number of voices waiting in the queue. ElevenLabs sends an email and in-app notification when the clone becomes available.
Check the progress under My Voices. The displayed status may show whether the clone is incomplete, waiting for verification, queued, generating its dataset, fine-tuning, or ready for use.
Test the Professional Voice Clone
Once training is complete:
- Open Voices.
- Select the personal voice area.
- Find the finished Professional Voice Clone.
- Click Use.
- Open it in Text to Speech.
- Generate a short passage.
- Compare the result with the original speaker.
- Test several speaking styles before beginning the full project.
The clone reproduces characteristics found in the training recordings. When the dataset contains mostly calm narration, the clone may perform most naturally with calm narration. Strong emotional performances may require source recordings containing a similar delivery.
Understand Sharing Restrictions
Instant Voice Clones cannot be shared with other ElevenLabs users. Professional Voice Clones can be shared privately by their verified owners and may have additional sharing options through the Voice Library.
When another person wants you to use their Professional Voice Clone:
- They must create the clone in their own account.
- They must complete voice verification.
- They must wait for training to finish.
- They can then provide an authorized sharing link.
Do not ask them to send you recordings so that you can create the Professional Voice Clone through your account. ElevenLabs does not permit that workflow, even when they have verbally agreed.
Review the Data-Use Setting
ElevenLabs currently allows users to review whether newly submitted data may be used to improve its general models.
To find the option:
- Open the profile menu.
- Select Terms and privacy.
- Open Data use.
- Review Improve the models for everyone.
- Change the setting when appropriate.
- Save the selection.
ElevenLabs states that disabling the option prevents new submitted data from being used for general model improvement. It also states that it does not use customer data to clone someone’s voice without consent.
Save the Original Recordings
Keep secure copies of:
- The unedited source recording
- The cleaned recording
- Every uploaded file
- The recording script
- The microphone and recording settings
- The clone’s voice name
- The selected speech model
- The approved Text to Speech settings
Voice clones cannot be exported as standalone models. Recreating one later requires the original source recordings, and a newly created clone may not sound exactly identical even when the same audio is used.
Use the Clone Responsibly
Do not use a cloned voice to:
- Impersonate someone deceptively
- Make a person appear to say something they did not authorize
- Mislead customers, family members, employers, or the public
- Bypass identity or security checks
- Commit fraud
- Conceal the real source of harmful content
ElevenLabs’ policies prohibit illegal, harmful, fraudulent, and abusive uses of its voice technology, and generated audio can be traced back to the account responsible for creating it.
When the audience could reasonably mistake generated speech for a real recording, clearly disclosing that the audio was AI-generated is the safest and most transparent approach.
Complete a Voice-Cloning Check
Before using the cloned voice in a complete project, confirm that:
- You own the voice or have the required permission.
- A Professional Voice Clone was created only by its actual owner.
- The source audio contains one clear speaker.
- Background noise and echo are minimal.
- The recordings use a consistent style.
- The accent and vocal tone sound accurate.
- Important words have been tested.
- The clone remains consistent across several paragraphs.
- The original recordings are stored securely.
- The planned use follows ElevenLabs’ policies.
Voice cloning can create a consistent personal narrator without recording every new script manually. Once the clone has been tested and approved, the next step is learning how to transform an existing performance with Voice Changer.
Clone Your Own Voice With Permission
Voice cloning allows ElevenLabs to reproduce the vocal characteristics of a real speaker, including their tone, accent, rhythm, and speaking style. ElevenLabs offers two cloning methods: Instant Voice Cloning for fast results from a short recording and Professional Voice Cloning for a more accurate model trained with a much larger collection of audio.
Only clone a voice that you own or have clear permission to use. Professional Voice Cloning has stricter rules: you can create a Professional Voice Clone only of your own voice, even when another person has given permission. That person must create and verify the Professional Voice Clone through their own ElevenLabs account before sharing it with you.
Choose Between Instant and Professional Voice Cloning
Use Instant Voice Cloning when:
- You want a clone that is ready almost immediately.
- You have approximately one to two minutes of clean audio.
- You are testing whether voice cloning works for your project.
- You need a narrator for general videos, courses, or personal content.
- The speaker does not have a particularly unusual voice or rare accent.
Instant Voice Cloning does not train a dedicated model on the speaker. It uses the uploaded audio to estimate the voice based on patterns ElevenLabs has already learned. This makes the process fast, but it may be less accurate for unusual voices or accents.
Use Professional Voice Cloning when:
- You need a more accurate reproduction of your own voice.
- The voice will be used regularly or professionally.
- Your accent or vocal characteristics are difficult for Instant Voice Cloning to reproduce.
- You can provide at least 30 minutes of clean recordings.
- You are willing to complete voice verification and wait for training.
Professional Voice Cloning trains a dedicated model using a larger collection of recordings. It is designed to reproduce the source voice more accurately and consistently than an Instant Voice Clone.
Check Your Plan
Instant Voice Cloning is available on the Starter plan and above. It is not included with the Free plan.
Professional Voice Cloning requires the Creator plan or above. Current base Professional Voice Clone limits include:
- Free and Starter: no Professional Voice Clone slots
- Creator and Pro: one slot
- Scale: three slots
- Business: ten slots
- Enterprise: a custom number
Plan limits can change, so review the current subscription page before upgrading specifically for voice cloning.
Record a Clean Voice Sample
The source recording has the greatest effect on the quality of the clone. A clean one-minute recording can create a better Instant Voice Clone than several minutes of inconsistent or noisy audio.
Record in a quiet room with:
- One speaker only
- Minimal background noise
- Little or no room echo
- A consistent microphone distance
- Clear speech
- Consistent volume
- A natural but steady speaking style
- No music or sound effects
- No extremely long pauses
ElevenLabs recommends one to two minutes of good audio for Instant Voice Cloning and approximately 30 to 180 minutes for Professional Voice Cloning. Audio quality is more important than the number of separate files.
Use a Consistent Microphone
Record the complete sample with the same microphone and settings whenever possible.
Avoid combining:
- Phone recordings with studio recordings
- Quiet narration with shouted dialogue
- Different rooms with noticeably different echo
- Recordings made from different distances
- Audio with changing background noise
- Clips that have received very different processing
A clone can reproduce imperfections contained in the source material. Noise, distortion, echo, compression, and inconsistent volume may therefore appear in the generated speech.
Keep the Speaking Style Consistent
The clone learns more than the physical sound of the voice. It also reflects the speaking style present in the recordings.
For example, use:
- Calm educational narration for a tutorial voice
- Conversational speech for a podcast voice
- Clear storytelling for an audiobook voice
- Energetic delivery for advertising narration
Do not combine several completely different performances in one dataset unless that variation is intentional. ElevenLabs recommends keeping the style consistent because the delivery contained in the source recordings influences the generated output.
Prepare an Instant Voice Clone Recording
For a beginner’s first Instant Voice Clone, record between one and two minutes of continuous speech.
The recording should contain:
- Complete sentences
- Different vowel and consonant sounds
- Questions and statements
- Short and longer sentences
- Natural pauses
- The speaker’s normal accent
- The vocal tone you want the clone to reproduce
Do not deliberately change your voice while recording. Speak naturally and use the delivery you expect to use most often.
ElevenLabs notes that adding more than two or three minutes to an Instant Voice Clone usually provides little improvement and can sometimes reduce stability.
Choose a File Format
ElevenLabs accepts several audio-file types for voice cloning. Its current recommendation is an MP3 file with a bitrate of at least 192 kilobits per second.
An uncompressed WAV file does not necessarily create a noticeably better clone. A clean recording matters more than using the largest possible file format.
Keep the original recording after uploading it. ElevenLabs does not allow users to export a voice clone itself, so the source samples are needed when you want to recreate the voice later.
Create an Instant Voice Clone
To create the clone:
- Open Voices from the ElevenLabs dashboard.
- Select My Voices or the personal voice area.
- Click the plus icon or Create Voice.
- Choose Instant Voice Clone.
- Upload the clean recording or record directly through the browser.
- Enter a recognizable name.
- Add an optional description.
- Confirm that you have the required permission to clone the voice.
- Complete the creation process.
- Open the clone in Text to Speech.
Use a name that explains the purpose, such as:
- Mildred Tutorial Voice
- English Course Narrator
- Spanish Product Voice
- Calm YouTube Narrator
A clear name becomes especially useful when the account contains several custom voices.
Test the Instant Voice Clone
Do not begin with the complete project. Generate a short test containing different types of speech.
For example:
“Welcome to today’s tutorial. First, we’ll open the dashboard and review the main settings. Does everything look correct? Great. Let’s continue.”
Review:
- Voice similarity
- Accent
- Clarity
- Speaking speed
- Emotional delivery
- Pronunciation
- Volume
- Consistency between sentences
Create a second test containing names, numbers, and terminology from the real project.
Adjust the Text to Speech Settings
After creating the clone, open its voice settings and begin near the standard defaults.
Similarity controls how closely the output follows the uploaded voice. Increasing it may improve resemblance, but it can also reproduce unwanted noise or imperfections from the original recording.
Stability controls how consistent or expressive the performance sounds. Higher stability generally creates more controlled narration, while lower stability allows more variation.
Adjust only one setting at a time. When the clone still sounds inaccurate after several small adjustments, improve the recording rather than forcing the settings to compensate for weak source audio.
Correct an Inaccurate Instant Clone
When the clone does not sound similar enough:
- Confirm that only one speaker appears in the recording.
- Remove clips containing background noise or echo.
- Use a recording with consistent volume.
- Make sure the speaker uses their normal accent.
- Remove sections containing an unusual or exaggerated performance.
- Create a new clone with the improved sample.
- Test another speech model.
A clone’s accent and tone are largely determined by its source samples. ElevenLabs states that the main way to change an inaccurate accent or tone is to create the clone again using different recordings.
Prepare Audio for a Professional Voice Clone
Professional Voice Cloning requires significantly more material.
ElevenLabs recommends:
- At least 30 minutes of high-quality speech
- Preferably two or more hours for stronger results
- Up to approximately three hours of useful training audio
- One speaker throughout
- Consistent recording quality
- Consistent delivery and tone
- Separate files of roughly 30 minutes when uploading several hours
Adding low-quality audio only to increase the total length can weaken the clone. Use the clearest and most representative recordings available.
Professional Voice Clones currently support spoken audio rather than singing samples.
Create a Professional Voice Clone
To begin:
- Open Voices.
- Select Create Voice.
- Choose Professional Voice Clone.
- Name the voice.
- Select the primary language when requested.
- Upload the prepared recordings.
- Review the platform’s feedback about the recording length.
- Process the uploaded audio.
- Complete the voice-verification step.
- Submit the clone for fine-tuning.
Do not begin the process using another person’s recordings. Professional Voice Cloning requires the person creating the clone to verify that the uploaded voice is their own.
Complete Voice Verification
ElevenLabs asks the voice owner to record a verification sample through the microphone. This confirms that the person creating the Professional Voice Clone matches the speaker in the uploaded recordings.
For the best chance of successful verification:
- Allow the browser to access the microphone.
- Use the same or similar microphone used for the original recordings.
- Find a quiet location.
- Speak in a similar tone and style.
- Avoid background audio.
- Confirm that the microphone is not muted.
When all verification attempts fail, ElevenLabs currently requires users to wait 24 hours before trying again or contact its support team for assistance.
Wait for Fine-Tuning to Finish
Professional Voice Clones are not ready immediately. After verification, ElevenLabs places the voice in a training queue and fine-tunes the model using the uploaded recordings.
The process commonly takes several hours but may take longer depending on the dataset, processing status, and number of voices waiting in the queue. ElevenLabs sends an email and in-app notification when the clone becomes available.
Check the progress under My Voices. The displayed status may show whether the clone is incomplete, waiting for verification, queued, generating its dataset, fine-tuning, or ready for use.
Test the Professional Voice Clone
Once training is complete:
- Open Voices.
- Select the personal voice area.
- Find the finished Professional Voice Clone.
- Click Use.
- Open it in Text to Speech.
- Generate a short passage.
- Compare the result with the original speaker.
- Test several speaking styles before beginning the full project.
The clone reproduces characteristics found in the training recordings. When the dataset contains mostly calm narration, the clone may perform most naturally with calm narration. Strong emotional performances may require source recordings containing a similar delivery.
Understand Sharing Restrictions
Instant Voice Clones cannot be shared with other ElevenLabs users. Professional Voice Clones can be shared privately by their verified owners and may have additional sharing options through the Voice Library.
When another person wants you to use their Professional Voice Clone:
- They must create the clone in their own account.
- They must complete voice verification.
- They must wait for training to finish.
- They can then provide an authorized sharing link.
Do not ask them to send you recordings so that you can create the Professional Voice Clone through your account. ElevenLabs does not permit that workflow, even when they have verbally agreed.
Review the Data-Use Setting
ElevenLabs currently allows users to review whether newly submitted data may be used to improve its general models.
To find the option:
- Open the profile menu.
- Select Terms and privacy.
- Open Data use.
- Review Improve the models for everyone.
- Change the setting when appropriate.
- Save the selection.
ElevenLabs states that disabling the option prevents new submitted data from being used for general model improvement. It also states that it does not use customer data to clone someone’s voice without consent.
Save the Original Recordings
Keep secure copies of:
- The unedited source recording
- The cleaned recording
- Every uploaded file
- The recording script
- The microphone and recording settings
- The clone’s voice name
- The selected speech model
- The approved Text to Speech settings
Voice clones cannot be exported as standalone models. Recreating one later requires the original source recordings, and a newly created clone may not sound exactly identical even when the same audio is used.
Use the Clone Responsibly
Do not use a cloned voice to:
- Impersonate someone deceptively
- Make a person appear to say something they did not authorize
- Mislead customers, family members, employers, or the public
- Bypass identity or security checks
- Commit fraud
- Conceal the real source of harmful content
ElevenLabs’ policies prohibit illegal, harmful, fraudulent, and abusive uses of its voice technology, and generated audio can be traced back to the account responsible for creating it.
When the audience could reasonably mistake generated speech for a real recording, clearly disclosing that the audio was AI-generated is the safest and most transparent approach.
Complete a Voice-Cloning Check
Before using the cloned voice in a complete project, confirm that:
- You own the voice or have the required permission.
- A Professional Voice Clone was created only by its actual owner.
- The source audio contains one clear speaker.
- Background noise and echo are minimal.
- The recordings use a consistent style.
- The accent and vocal tone sound accurate.
- Important words have been tested.
- The clone remains consistent across several paragraphs.
- The original recordings are stored securely.
- The planned use follows ElevenLabs’ policies.
Voice cloning can create a consistent personal narrator without recording every new script manually. Once the clone has been tested and approved, the next step is learning how to transform an existing performance with Voice Changer.
keep going
Transform a Recording With Voice Changer
Voice Changer converts an existing performance into a different voice while preserving much of the original speaker’s timing, emotion, accent, pacing, and delivery. It was previously called Speech to Speech. Unlike Text to Speech, you perform the line yourself first and then select the voice that should replace your original voice.
This tool is useful when Text to Speech does not produce the exact performance you need. You can record the desired pause, whisper, laugh, sigh, emphasis, or emotional delivery yourself and allow ElevenLabs to transfer that performance to another voice.
Understand How Voice Changer Works
Voice Changer uses two main elements:
- The source recording, which provides the words, timing, emotion, and performance
- The output voice, which determines who the finished recording sounds like
For example, you could record a tutorial sentence using your own voice and convert it into a saved narrator voice. The finished audio should follow your original pacing while sounding like the selected output speaker.
Voice Changer can preserve performance details such as:
- Whispering
- Laughing
- Crying
- Sighing
- Emotional tone
- Speaking cadence
- Accent
- Language
- Pauses and emphasis
The quality of the result still depends on the source recording, selected voice, model, and settings.
Open Voice Changer
- Sign in to ElevenLabs.
- Open the Playground.
- Select Voice Changer.
- Choose whether to record directly or upload an existing file.
- Select the voice that should replace the original speaker.
- Review the model and voice settings.
- Generate the converted audio.
- Listen to the complete result before downloading it.
The interface displays the estimated credit cost before generation. Recording or uploading the source does not consume the generation allowance by itself; credits are charged when the user confirms the generation.
Record Directly in the Browser
Recording directly is useful when you need to perform a short line, correction, reaction, or emotional passage.
Before recording:
- Connect the intended microphone.
- Allow the browser to access it.
- Move to a quiet location.
- Reduce background noise.
- Keep a consistent distance from the microphone.
- Read the script aloud once.
- Record the final performance.
Use headphones when background audio is playing nearby so it does not enter the microphone.
Speak naturally and focus on the delivery rather than trying to imitate the selected output voice. The Voice Changer will replace your vocal identity while attempting to preserve the performance.
Upload an Existing Recording
Upload a file when the performance has already been recorded or edited.
ElevenLabs currently accepts several common audio and video formats for Voice Changer, including formats such as MP3, WAV, M4A, FLAC, OGG, OGA, MKV, and WebM.
Use a recording that contains:
- One clear speaker
- Minimal background noise
- Little room echo
- Consistent volume
- No background music when possible
- No overlapping dialogue
- No heavy distortion
A clean source helps the tool identify the speech and performance more accurately.
Keep Each Recording Under Five Minutes
The maximum input length for one Voice Changer conversion is currently five minutes. Longer recordings must be divided into smaller sections before processing.
Dividing longer content into sections also makes it easier to:
- Correct one mistake
- Replace one sentence
- Compare voices
- Maintain consistent pacing
- Avoid regenerating approved audio
- Organize the final project
Use logical divisions such as paragraphs, scenes, speakers, or chapters rather than cutting in the middle of a sentence.
Understand the Credit Cost
Voice Changer currently costs 1,000 credits for each minute of processed audio. Billing is based on the recording’s duration rather than the number of written characters.
A recording shorter than one minute may still use a proportional amount according to the duration shown in the interface.
Before processing a long recording:
- Remove unnecessary silence.
- Delete incorrect takes.
- Trim the beginning and ending.
- Confirm that the correct voice is selected.
- Review the displayed credit cost.
- Generate only after the source is ready.
Do not repeatedly convert the complete five-minute recording to fix one weak sentence. Isolate that sentence, process it separately, and replace it during editing.
Select the Output Voice
Open the voice menu and choose the speaker that should appear in the finished audio.
The output voice can come from:
- The Voice Library
- My Voices
- Voice Design
- An Instant Voice Clone
- A Professional Voice Clone
Custom cloned or designed voices saved in the account can be used as Voice Changer output voices.
Test the selected voice with a short passage first. Some voices preserve the original performance more naturally than others, especially when the source contains strong emotion, unusual pacing, or a different language.
Match the Voice to the Performance
Choose an output voice that can convincingly perform the source material.
For example:
- Use a calm narrator for educational speech.
- Use an expressive voice for dramatic dialogue.
- Use a conversational voice for podcasts.
- Use an energetic voice for advertisements.
- Use a voice trained in the target language and accent.
A soft narrator may not handle shouting convincingly, while a highly energetic character voice may sound unnatural during calm instructional content.
Choose the Voice Changer Model
ElevenLabs currently provides an English-only Speech-to-Speech model and a multilingual model supporting 29 languages. Its documentation recommends the multilingual model in many cases because it may perform better even for English recordings.
Use the multilingual model when:
- The recording is not in English.
- Accent preservation matters.
- The English-only model produces weaker results.
- The content contains multilingual speech.
- You need broader language support.
Use the English-only model only when the source is fully English and testing shows that it produces the stronger result.
Preserve the Original Language and Accent
Voice Changer generally attempts to retain the source recording’s language and accent. This means an English output voice may still speak Spanish when the source recording is in Spanish, although the voice’s original training can affect the final accent and pronunciation.
For a more natural result:
- Select an output voice compatible with the target language.
- Speak clearly in the source recording.
- Avoid switching languages rapidly.
- Test regional terminology.
- Review names and borrowed words carefully.
When the output develops the wrong accent, try another voice trained in the target region.
Perform the Delivery You Want
Voice Changer follows the performance contained in the source recording. This gives you more direct control than writing emotional instructions inside a Text to Speech script.
During the recording, control:
- Speaking speed
- Pauses
- Emphasis
- Emotion
- Volume
- Questions
- Laughter
- Whispering
- Hesitation
- Character reactions
For example, when you need the output to whisper, whisper during the source recording. When you need a dramatic pause, perform the pause naturally instead of relying on punctuation.
Do Not Overact Unnecessarily
Voice Changer can preserve strong emotions, but exaggerated performances may create unstable or unnatural output.
Record several versions when a line is important:
- A subtle version
- A moderate version
- A more dramatic version
Convert the strongest source and compare the results. A performance that feels slightly exaggerated in the original recording may sometimes transfer well, but this varies by voice.
Use Voice Changer to Fix Pronunciation
Voice Changer can replace a mispronounced word or phrase from an existing Text to Speech recording.
A practical workflow is:
- Listen to the generated narration.
- Identify the incorrect word or sentence.
- Record yourself saying it correctly.
- Match the surrounding speed and emotion.
- Convert the recording into the same narrator voice.
- Place the corrected audio into the project.
- Compare the transition with the surrounding narration.
ElevenLabs specifically identifies pronunciation correction as one use for Voice Changer.
Record the complete phrase or sentence instead of only the individual word. This usually creates a smoother transition because the converted clip includes natural context.
Use Voice Changer to Direct a Performance
Voice Changer can help when Text to Speech delivers the correct words but not the intended rhythm or emotion.
For example, instead of repeatedly adjusting punctuation and stability:
- Record the passage with the exact desired delivery.
- Include the intended pauses and emphasis.
- Convert it into the project’s narrator voice.
- Replace the original Text to Speech version.
Inside ElevenCreative Studio, this type of workflow may appear as Direct speech with your voice or Actor Mode. Studio can also apply Voice Changer to existing audio while preserving the narration’s delivery.
Remove Background Noise
Voice Changer includes an option for automatically reducing background noise from the source recording.
Enable it when the source contains:
- Light fan noise
- Air-conditioning noise
- Low environmental hum
- Mild room noise
- Small microphone artifacts
However, noise removal cannot fully repair every recording. Heavy music, overlapping voices, loud traffic, echo, or severe distortion may remain noticeable.
When possible, create a cleaner recording rather than depending entirely on automatic cleanup.
Use Voice Isolator for Difficult Recordings
When a source contains significant background audio, Voice Isolator can separate the main voice from surrounding noise before the file is processed with Voice Changer. ElevenLabs supports uploading an existing file or recording directly in Voice Isolator.
A practical workflow is:
- Upload the noisy recording to Voice Isolator.
- Generate the isolated vocal track.
- Download and review it.
- Upload the cleaned voice to Voice Changer.
- Select the output voice.
- Generate the transformed recording.
Voice isolation may still create artifacts when the speech overlaps heavily with music or other people.
Adjust the Voice Settings
Voice Changer provides familiar controls such as stability, similarity, and other settings supported by the selected voice. It also adds the background-noise removal control.
Start near the default settings.
Raise stability when:
- The output changes tone unexpectedly.
- The voice becomes too emotional.
- The pacing feels inconsistent.
- Several sections need a similar delivery.
Lower stability slightly when:
- The output sounds flat.
- Emotional details from the original performance are being lost.
- The voice does not follow the source naturally.
Raise similarity when the selected voice does not sound recognizable enough. Lower it when the output reproduces unwanted artifacts or loses clarity.
Change One Setting at a Time
Use the same short recording while testing.
A useful order is:
- Confirm the output voice.
- Confirm the model.
- Generate with default settings.
- Adjust stability.
- Review similarity.
- Test background-noise removal.
- Compare the results.
Changing every control simultaneously makes it difficult to identify which adjustment improved or weakened the audio.
Compare the Source and Output
Listen to the source recording first and then the transformed version.
Compare:
- Timing
- Emotion
- Pronunciation
- Accent
- Pauses
- Volume
- Voice similarity
- Background artifacts
- Overall naturalness
The finished voice should sound like the selected speaker while retaining the important performance details from the source.
Troubleshoot a Weak Result
When the output sounds robotic:
- Use a more natural source performance.
- Select a more suitable voice.
- Reduce excessive noise processing.
- Try the multilingual model.
- Divide the recording into shorter sections.
When the output does not preserve emotion:
- Perform the emotion more clearly.
- Lower stability slightly.
- Choose a more expressive output voice.
- Avoid speaking too quietly.
- Process a complete sentence rather than one isolated word.
When the output changes the accent:
- Choose a voice trained in the target language.
- Use the multilingual model.
- Keep the source recording in one language.
- Test another output voice.
When the output contains background artifacts:
- Enable background-noise removal.
- Clean the source with Voice Isolator.
- Re-record in a quieter room.
- Remove music and overlapping speakers.
- Lower similarity when it reproduces source imperfections.
Maintain Consistency Across Several Clips
When converting a long project in separate sections, keep the following unchanged:
- Microphone
- Recording environment
- Distance from the microphone
- Output voice
- Voice Changer model
- Stability
- Similarity
- Background-noise setting
- Recording volume
Use similar performance energy throughout the project. A calm first clip and an overly energetic second clip may sound inconsistent even when the same output voice is selected.
Match Replacement Audio to Existing Narration
When replacing one sentence inside an existing recording:
- Listen to the previous sentence.
- Match its speaking speed.
- Match its emotional tone.
- Record the replacement sentence.
- Leave a small amount of space before and after it.
- Convert it with the same voice and settings.
- Align the replacement during editing.
- Add a short crossfade when needed.
The converted sentence may not match the original recording’s loudness perfectly. Adjust the level during editing rather than regenerating repeatedly for a small volume difference.
Use Voice Changer in Studio
Inside ElevenCreative Studio, users can select audio and apply Voice Changer while preserving the existing delivery. Studio also provides tools for removing background audio and directing speech through a user-recorded performance.
This workflow is useful when the project already contains:
- Narration
- Dialogue
- Music
- Video
- Captions
- Sound effects
Changing the voice inside the same Studio project reduces the need to export, edit, and re-upload every individual clip.
Download the Converted Audio
After approving the result:
- Locate the generated Voice Changer output.
- Play the full recording.
- Click the download control.
- Select an available format.
- Save the file to the correct project folder.
- Rename it clearly.
Voice Changer generations can be downloaded through the ElevenLabs interface after processing.
Use filenames such as:
Tutorial-Introduction-Voice-Changer-Final.wav
or:
Character-Two-Scene-Four-Take-Three.mp3
Save the Source Recording
Keep both the original and transformed files.
A useful folder structure is:
- Source recordings
- Cleaned recordings
- Voice Changer tests
- Approved output
- Rejected versions
- Final edited project
The source recording may be needed later when you want to use another voice, change the settings, or create a higher-quality conversion.
Use Authorized Voices Only
Voice Changer should not be used to deceptively impersonate another person or create speech that misleads listeners about who actually recorded it.
Use:
- Your own recordings
- Performances you are authorized to modify
- Voices you are permitted to use
- Licensed Voice Library voices
- Authorized custom clones
- Original Voice Design voices
When the transformed audio could reasonably be mistaken for a real recording, clearly identifying it as AI-generated or voice-modified is the most transparent approach.
Complete a Voice Changer Check
Before approving the transformed recording, confirm that:
- The source audio is clear.
- Only the intended speaker is audible.
- The recording is under five minutes.
- The correct output voice is selected.
- The correct language model is active.
- The emotional delivery is preserved.
- The accent and pronunciation sound natural.
- Background noise is not distracting.
- The credit cost has been reviewed.
- The approved version has been downloaded.
- The source recording has been saved.
- The voice and performance are authorized for use.
Voice Changer is most useful when you already know exactly how the line should be performed but need it to sound like another authorized voice. Once this workflow feels familiar, the next step is creating realistic conversations with multiple speakers.
Create Multi-Speaker Dialogue
ElevenLabs can generate conversations containing multiple voices through Dialogue mode. This feature uses Eleven v3 to create natural exchanges with changes in tone, emotional context, pauses, interruptions, and reactions between speakers. It is useful for podcasts, audiobooks, fictional scenes, training conversations, advertisements, games, and scripted interviews.
Open Dialogue Mode
- Sign in to ElevenLabs.
- Open Text to Speech.
- Select Eleven v3 as the speech model.
- Add more than one speaker to the generation.
- Assign a different voice to each speaker.
- Enter each person’s dialogue in the correct speaker block.
- Generate the conversation.
- Listen to the complete exchange before downloading it.
Dialogue mode becomes available when multiple speakers are used through the ElevenLabs website. Text to Dialogue is currently powered only by Eleven v3.
Choose the Speakers
Select a separate voice for each person in the conversation. Voices can come from:
- The Voice Library
- My Voices
- Voice Design
- An authorized Instant Voice Clone
- An authorized Professional Voice Clone
The voices should be different enough for listeners to recognize who is speaking. A conversation may become confusing when both speakers have nearly identical tone, pitch, pacing, and accents.
For example, pair:
- A calm, lower-pitched narrator with an energetic, higher-pitched speaker
- A professional instructor with a casual student
- A serious interviewer with a conversational guest
- Two character voices with clearly different personalities
ElevenLabs recommends choosing voices that naturally match the required delivery. A serious voice may not respond convincingly to playful or exaggerated instructions, while an energetic voice may struggle with restrained narration.
Test Each Voice Separately
Before generating the complete dialogue, test each voice with one or two of its real lines.
Listen for:
- Accent
- Pronunciation
- Speaking speed
- Emotional range
- Clarity
- Character suitability
- Compatibility with Eleven v3
A voice that works well with another model may behave differently with Eleven v3. ElevenLabs notes that Voice Library voices can produce more variable results with v3, so testing them before a long generation is important.
Give Each Speaker a Clear Personality
Write a brief description of each speaker before drafting the dialogue.
For example:
Speaker 1: Calm technology instructor who explains each step clearly.
Speaker 2: Curious beginner who asks short, practical questions.
This helps keep the conversation consistent. The speakers should not suddenly exchange personalities unless the story intentionally requires it.
A useful character outline can include:
- Role
- Age range
- Attitude
- Energy level
- Relationship to the other speaker
- Speaking style
- Important emotional traits
The descriptions are for planning and do not need to appear in the generated script unless they should be spoken aloud.
Write Natural Conversation
Dialogue should sound like people speaking to each other rather than an article divided between two voices.
Avoid:
Speaker 1:
“ElevenLabs is a platform that provides numerous artificial-intelligence-based audio-production capabilities.”
Speaker 2:
“That is correct. It offers several tools that can assist users with their audio-production requirements.”
Use:
Speaker 1:
“Have you tried ElevenLabs yet?”
Speaker 2:
“Not really. I opened it, but I wasn’t sure where to start.”
Speaker 1:
“Start with Text to Speech. It’s the easiest tool to learn.”
Natural conversation usually includes questions, short replies, reactions, interruptions, incomplete thoughts, and differences in sentence length.
Keep Speaker Turns Short
Long speeches can make a conversation feel like two separate narrators reading essays. Divide complicated information into smaller exchanges.
For example:
Speaker 1:
“First, open Text to Speech.”
Speaker 2:
“Do I need a paid account?”
Speaker 1:
“No. You can test it using the Free plan.”
Speaker 2:
“Okay. What do I do next?”
Short turns create a more believable rhythm and give Eleven v3 clearer emotional context between speakers.
Use Audio Tags for Delivery
Eleven v3 supports natural-language audio tags placed inside square brackets. These instructions can guide emotion, delivery, and audible reactions.
Examples include:
[excited][thoughtful][sad][angry][whispering][laughs][sighs][clears throat][sarcastic][nervous]
Place the tag near the dialogue it should affect:
Speaker 1:[excited] I finally finished the voiceover!
Speaker 2:[surprised] Already? That was fast.
Audio tags should describe an audible emotion, reaction, or delivery style. They should not be used for silent visual actions such as [smiling], [walking], or [looking away].
Match Audio Tags to the Voice
The selected voice must be capable of producing the requested performance.
For example, a calm professional narrator may handle:
[thoughtful][serious][reassuring][quietly]
A playful character voice may respond better to:
[giggling][excited][mischievously][dramatically]
ElevenLabs warns that tags may perform poorly when they conflict with the voice’s original character or training style.
Do Not Tag Every Line
A natural conversation does not require an emotional instruction before every sentence.
Too many tags may make the scene sound exaggerated, inconsistent, or distracting. Use them when:
- A speaker’s emotion changes
- Someone whispers or shouts
- A reaction is important
- A line could be interpreted several ways
- The scene includes laughter, sighing, or hesitation
- A speaker interrupts another person
Allow ordinary lines to rely on the dialogue’s wording and punctuation.
Use Punctuation to Control the Flow
Eleven v3 interprets punctuation and text structure as part of the performance. Use periods, commas, dashes, ellipses, question marks, and exclamation points to guide the exchange. Eleven v3 does not support SSML break tags, so pauses should be controlled through punctuation, text structure, and audio tags.
For example:
A normal pause:
“I understand. Let’s try it again.”
Hesitation:
“I don’t know… maybe we should wait.”
Interruption:
“But I thought you said—”
A strong reaction:
“You deleted the whole project?”
Excitement:
“It worked!”
Write Interruptions Carefully
Dialogue mode can produce conversations with overlapping or interrupted speech. Use a dash at the end of the interrupted speaker’s line and begin the next line as a direct response. ElevenLabs provides interruption-style prompting examples in its official dialogue guidance.
For example:
Speaker 1:I was going to download the file, but—
Speaker 2:[interrupting] Wait. Did you check the pronunciation first?
Speaker 1:No. That’s a good point.
Do not create an interruption in every exchange. Too many overlapping lines can make the scene difficult to understand.
Use Ellipses for Hesitation
Ellipses can suggest that a speaker is thinking, hesitating, or trailing off.
For example:
Speaker 1:I thought I saved the recording…
Speaker 2:You did save it, right?
Speaker 1:I’m checking.
ElevenLabs uses ellipses in its official dialogue examples to represent unfinished or uncertain speech.
Add Reactions Naturally
Reactions make dialogue feel less scripted.
Examples include:
Speaker 1:[sighs] I generated the wrong version again.
Speaker 2:That’s okay. It should still be in your history.
or:
Speaker 1:I finally fixed the pronunciation.
Speaker 2:[relieved laugh] Good. I thought we’d have to record everything again.
A reaction should support the conversation rather than replace meaningful dialogue.
Keep the Emotional Context Clear
The same words can sound very different depending on the situation.
For example:
That’s great.
could sound sincere, sarcastic, relieved, disappointed, or surprised.
Provide context through the surrounding lines or add a suitable tag:
[sarcastic] That’s great. Now we have to start over.
or:
[relieved] That’s great. I was worried we had lost everything.
Eleven v3 interprets emotional context from the text, punctuation, and audio tags rather than relying only on a separate emotion control.
Use the Enhance Feature Carefully
The ElevenLabs interface may provide an Enhance option that adds relevant audio tags to dialogue automatically. ElevenLabs states that this feature preserves the original wording while inserting tags intended to make the performance more expressive.
After using Enhance:
- Read every added tag.
- Remove unnecessary emotions.
- Check that no tag changes the intended meaning.
- Confirm that the tag matches the selected voice.
- Generate a short test.
- Compare it with the unenhanced version.
Automatic enhancement can provide ideas, but it should not replace manual review.
Choose the Stability Mode
Eleven v3 provides stability modes that affect how expressive or consistent the generation becomes:
- Creative allows stronger emotion but carries a higher risk of unexpected output.
- Natural offers a balanced performance close to the original voice.
- Robust prioritizes consistency but responds less strongly to expressive directions.
Start with Natural for most dialogue.
Use Creative when:
- The scene is dramatic
- Emotional variation matters
- The characters need strong reactions
- You are prepared to compare several generations
Use Robust when:
- The dialogue must remain controlled
- The speakers change tone unexpectedly
- The scene sounds too chaotic
- Consistency matters more than emotional range
Understand That Speed Works Differently With v3
Eleven v3 does not provide the same standard speed-slider workflow described for other Text to Speech models. Its pacing is primarily influenced through voice choice, script structure, punctuation, and audio tags.
To slow a speaker down:
- Use shorter sentences
- Add meaningful punctuation
- Use a measured voice
- Add instructions such as
[slowly]or[deliberately] - Divide long lines
To create faster delivery:
- Use shorter pauses
- Select a naturally energetic voice
- Use instructions such as
[quickly]or[excitedly] - Keep the dialogue concise
Always test the result because different voices may interpret the same instruction differently.
Keep the First Dialogue Short
For the first test, create a conversation containing four to eight speaker turns.
For example:
Speaker 1:[friendly] Have you created your first ElevenLabs voiceover yet?
Speaker 2:Not yet. I’m still choosing a voice.
Speaker 1:Try generating the same sentence with two different voices.
Speaker 2:That makes sense. Then I can compare them directly.
Speaker 1:Exactly. Keep the first script short so you don’t waste credits.
A short test makes it easier to review the speakers, timing, emotion, and pronunciation before producing a complete podcast or scene.
Generate the Dialogue
Once the voices and script are ready:
- Review every speaker assignment.
- Confirm that Eleven v3 is selected.
- Check the script for incorrect speaker names or duplicated lines.
- Review the displayed credit cost.
- Click Generate.
- Listen to the complete conversation.
- Compare the transitions between speakers.
- Use an available regeneration when another performance is needed.
- Download the strongest result.
Dialogue output is nondeterministic, meaning repeated generations may produce different pacing, emotional delivery, and interactions even when the text and voices remain unchanged. ElevenLabs recommends comparing several generations when the first result does not match the intended performance.
Review Each Speaker
Listen to each voice separately and then review the complete conversation.
Check:
- Whether the correct voice reads each line
- Whether either speaker changes accent
- Whether one voice is much louder
- Whether interruptions sound understandable
- Whether reactions occur in the correct place
- Whether the emotional delivery matches the scene
- Whether names and technical terms are pronounced correctly
- Whether the conversation moves too quickly or slowly
A strong individual voice does not automatically guarantee that the two voices sound natural together.
Regenerate Weak Sections
When one part of the dialogue sounds incorrect, avoid rewriting the complete script immediately.
First try:
- Correcting the punctuation
- Changing one audio tag
- Shortening the affected line
- Selecting a more suitable voice
- Changing the stability mode
- Generating another version
The model may produce a stronger performance on another attempt because its output varies between generations.
Avoid Extremely Short Prompts
Very short Eleven v3 prompts may produce less consistent results. ElevenLabs recommends providing enough context for the model to understand the intended voice and emotion.
Instead of:
Speaker 1:Fine.
Use:
Speaker 1:[frustrated] Fine. We’ll do it your way.
The longer version explains the intended delivery more clearly.
Split Long Conversations Into Scenes
Do not generate an entire long podcast, audiobook chapter, or dramatic scene in one request.
Divide it into:
- Introduction
- Topic sections
- Individual scenes
- Interview questions
- Story chapters
- Conflict and resolution
- Closing conversation
Shorter sections are easier to revise and organize. They also reduce the risk of one weak line requiring the entire conversation to be generated again.
ElevenLabs recommends keeping Text to Dialogue API requests at or below approximately 2,000 total text characters for reliable results, which also provides a practical reason to divide longer conversations into smaller sections.
Maintain Continuity Between Sections
When creating separate dialogue clips, record:
- Speaker names
- Voice names or IDs
- Stability mode
- Emotional state
- Scene context
- Pronunciation decisions
- Approved audio tags
- The ending tone of the previous clip
Begin the next section with the same emotional context unless the scene intentionally changes.
For example, when one clip ends with a speaker feeling worried, the next clip should not begin with that person sounding cheerful without an explanation.
Use Studio for Longer Dialogue Projects
For longer productions, move the approved dialogue into ElevenCreative Studio. Studio provides a timeline where narration, dialogue, video, music, captions, and sound effects can be organized together.
Studio is better suited for:
- Podcast episodes
- Audiobooks
- Video scenes
- Training conversations
- Fictional stories
- Advertisements with several speakers
- Projects requiring music and sound effects
Each voice can be placed in the correct scene, and individual clips can be repositioned or regenerated without rebuilding the complete project.
Balance the Speaker Volume
Two voices may generate at different loudness levels.
When one speaker sounds noticeably louder:
- Check whether the voice itself is naturally louder.
- Review the generation settings.
- Regenerate the weaker or stronger line.
- Adjust clip volume inside Studio or an audio editor.
- Listen through headphones and speakers.
- Avoid increasing volume until distortion appears.
Volume differences are often easier to correct during editing than through repeated speech generation.
Add Music and Sound Effects Later
Create the dialogue first before adding background music or sound effects.
A practical order is:
- Approve the script.
- Choose the voices.
- Generate the dialogue.
- Correct pronunciation and pacing.
- Arrange the clips.
- Add sound effects.
- Add music.
- Mix the final audio.
Adding music too early can make it harder to hear voice problems or compare generations accurately.
Download and Name the Dialogue
After approving the result:
- Download the generated audio.
- Select an appropriate available format.
- Save it inside the project folder.
- Give it a descriptive filename.
- Save the script beside the audio.
For example:
Podcast-Episode-One-Introduction-Dialogue-V2.wav
or:
Training-Conversation-Account-Setup-Final.mp3
When separate scenes exist, number them in order:
Scene-01-Introduction.wavScene-02-Problem.wavScene-03-Solution.wav
Use the Voices Responsibly
Every speaker voice should be original, licensed, publicly available for the permitted use, or cloned with proper authorization.
Do not create dialogue that deceptively makes a real person appear to participate in a conversation they never recorded or approved. When synthetic speech could be mistaken for an authentic recording, clearly disclosing that the conversation uses AI-generated voices is the most transparent approach.
Complete a Multi-Speaker Dialogue Check
Before approving the conversation, confirm that:
- Eleven v3 is selected.
- Every speaker has the correct voice.
- The voices are easy to distinguish.
- The dialogue sounds conversational.
- Speaker turns are not unnecessarily long.
- Audio tags match the selected voices.
- Interruptions remain understandable.
- Emotional changes make sense.
- Pronunciation and accents are correct.
- Volume levels are reasonably balanced.
- The approved settings have been recorded.
- The final audio and script have been saved.
- Every voice is authorized for the intended use.
Multi-speaker dialogue is most effective when the script, voices, emotions, and timing work together. Once the conversation sounds natural, the next step is organizing longer voiceovers, dialogue, music, and video inside ElevenCreative Studio.
Build Longer Projects in ElevenCreative Studio
ElevenCreative Studio is ElevenLabs’ timeline-based workspace for creating longer audio and video projects. It combines narration, dialogue, video, captions, music, and sound effects in one place, making it better suited to audiobooks, courses, podcasts, narrated articles, and complete video voiceovers than the standard Text to Speech Playground.
Open ElevenCreative Studio
- Sign in to ElevenLabs.
- Select Studio from the sidebar.
- Choose one of the available project options.
- Follow the setup instructions.
- Enter a recognizable project name.
- Select the initial voice and other available settings.
- Click Create.
Studio is currently available on all ElevenLabs plans, including the Free plan. Certain specialized workflows, such as the GenFM podcast-creation feature, may require a paid subscription.
Choose a Starting Option
Studio provides several prepared starting points depending on what you want to create. The available options may include projects for audiobooks, article narration, voiceovers, captions, dubbing, soundtracks, faceless videos, and blank audio or video productions.
Use a prepared option when you want Studio to configure part of the project automatically.
Choose a blank project when you want complete control over:
- Narration
- Chapters
- Scenes
- Speakers
- Video
- Captions
- Music
- Sound effects
- Timing
Beginners should start with a simple audio project containing only narration. Add video, music, and sound effects after learning the basic timeline.
Import Existing Content
Studio can create a project from an existing document, written script, webpage, or book. Supported import sources currently include:
- EPUB
- DOCX
- TXT
- HTML
- A webpage URL
EPUB is particularly useful for audiobooks because a properly structured EPUB can automatically divide the content into chapters. ElevenLabs recommends formatting each chapter heading as Heading 1 so Studio can recognize the chapter structure.
To import content:
- Select the appropriate project type.
- Upload the document or enter the webpage URL.
- Wait for Studio to process the content.
- Review the imported headings and paragraphs.
- Remove anything that should not be spoken.
- Correct formatting problems before generating narration.
Imported documents may contain page numbers, citations, menus, image descriptions, headers, footers, or other material that should not become part of the voiceover. Review every chapter before converting it to speech.
Create a Blank Project
A blank project is useful when the script is still being written or the content needs to be added manually.
After opening the project:
- Add a scene, chapter, or narration block.
- Paste the first section of the script.
- Divide the text into short paragraphs.
- Assign the correct voice.
- Select the speech model.
- Generate one test paragraph.
- Review it before generating the remaining content.
Do not paste the entire project into one paragraph. Separate paragraphs are easier to regenerate, rearrange, lock, and assign to different speakers.
Understand the Timeline
The Studio timeline appears along the bottom of the workspace and provides a visual representation of the project’s audio and media. Different types of content appear on separate horizontal tracks.
Tracks can include:
- Narration
- Dialogue
- Video
- Captions
- Music
- Sound effects
Each clip’s position determines when it plays. Moving a clip to the right creates a delay, while overlapping music and narration allows both to play at the same time.
The overall project may become longer than the narration when a video, music track, or sound effect extends beyond the final spoken paragraph. Review the end of every track before exporting.
Navigate the Timeline
Use the timeline to:
- Move clips
- Trim their beginning or ending
- Split longer sections
- Arrange scenes
- Align audio with video
- Position sound effects
- Adjust background music
- Review the complete production
Zoom in when making precise timing changes. Zoom out when reviewing the overall chapter or project structure.
For a beginner’s first project, keep the timeline simple:
- Add the narration.
- Correct the narration.
- Arrange the paragraphs.
- Add visuals when needed.
- Add sound effects.
- Add music last.
This prevents background media from distracting you while reviewing pronunciation and pacing.
Organize the Project Into Chapters
Chapters help divide long content into manageable sections. They are useful for audiobooks, courses, podcast episodes, training modules, and lengthy narrated articles.
When Studio recognizes chapters during an audiobook import, it can create them automatically. Existing projects can also be managed by opening Project options and selecting Manage chapters. From there, users can add, rename, remove, and rearrange chapters.
Use chapter names that clearly describe the content:
- Introduction
- Account Setup
- Text to Speech Basics
- Voice Selection
- Advanced Tools
- Final Recommendations
Avoid vague names such as “Part 1” unless the meaning is obvious from the rest of the project.
Select the Default Voice
Choose a default narrator for the project before generating several paragraphs. Studio currently uses Eleven Multilingual v2 as the default model for many newly created projects, although other models, including Eleven v3, can be selected through the project settings.
The default voice should match:
- The project’s language
- The audience
- The subject
- The intended emotional tone
- The required accent
- The expected project length
Test the narrator with several types of sentences before applying it to the full project.
Assign Different Voices to Sections
Studio allows users to assign different voices and settings to individual paragraphs, characters, or sections. This is useful for dialogue, audiobooks, interviews, training conversations, and projects containing quotations from several speakers.
To assign another voice:
- Select the paragraph or dialogue section.
- Open its voice menu.
- Choose a voice from the available collection.
- Review the speech model and settings.
- Generate the section.
- Listen to the transition between speakers.
Keep the number of voices manageable. A project containing too many similar speakers can become confusing.
Apply Voice Settings Consistently
Use the same voice, model, stability, similarity, style, and speed settings across related paragraphs unless a deliberate change is needed.
For example, an audiobook narrator should not suddenly become faster or more emotional in the middle of a normal descriptive passage.
Create a production note containing:
- Voice name or ID
- Speech model
- Stability
- Similarity
- Style
- Speed
- Pronunciation rules
- Intended emotional tone
This helps maintain consistency when revising the project later.
Generate One Paragraph First
Before generating an entire chapter:
- Choose a representative paragraph.
- Confirm the narrator.
- Check the model.
- Review the settings.
- Generate the audio.
- Listen for pronunciation and pacing.
- Correct the script when necessary.
- Approve the paragraph before continuing.
This test should contain the type of language used throughout the project, including important names, numbers, or technical terminology.
Generating one paragraph first can prevent spending credits on an entire chapter that uses the wrong voice or settings.
Generate Narration in Sections
Studio lets users generate and regenerate individual paragraphs or words rather than rebuilding the entire production. This makes it easier to correct a small problem without replacing approved audio elsewhere in the project.
A practical workflow is:
- Generate several paragraphs.
- Listen to them in order.
- Correct weak passages.
- Approve the strongest versions.
- Continue to the next group.
- Review the complete chapter afterward.
Do not generate every chapter before listening to the beginning. A repeated pronunciation or voice-setting problem could affect the whole project.
Regenerate a Weak Paragraph
When a paragraph sounds incorrect:
- Select the affected paragraph.
- Check the written text.
- Correct punctuation or pronunciation.
- Review its voice and settings.
- Generate another version.
- Compare it with the previous result.
- Keep the stronger version.
Studio’s Generation History can help users review or restore earlier versions of generated sections.
Do not delete a strong generation until the replacement has been reviewed in context.
Lock Approved Sections
Studio allows users to lock sections after they are approved. Locking helps protect completed narration from accidental changes while other parts of the project are still being edited.
Lock a paragraph when:
- The wording is final.
- The pronunciation is correct.
- The voice and pacing are approved.
- Its position on the timeline is correct.
- No further regeneration should be necessary.
Unlock it only when a correction is required.
Use Generation History
Generation History stores previous audio versions, allowing users to revisit earlier results and recover a version that sounded better than the latest attempt.
Use it when:
- A regeneration sounds worse.
- A voice setting was changed accidentally.
- You need to compare several performances.
- An approved version was replaced.
- You want to download an earlier result.
AI speech can vary between generations, so recreating an earlier performance exactly may not be possible. Keep important exports outside ElevenLabs as well.
Add Video and Images
Studio can import video and image assets for narrated video projects. The timeline provides a video track and caption layer so narration can be synchronized with visual content.
After adding a video:
- Place it on the video track.
- Add or generate the narration.
- Move the narration to the correct starting point.
- Trim the video when necessary.
- Align spoken instructions with the matching visuals.
- Review the complete scene.
Do not rely only on the waveform. Watch the full video to confirm that each sentence appears at the correct moment.
Add Captions
Captions can be included on a separate layer in Studio video projects. They help viewers follow the narration when watching without sound and improve accessibility.
Review captions for:
- Misspelled names
- Incorrect punctuation
- Poor line breaks
- Words appearing too early or late
- Captions covering important visuals
- Differences between the caption and narration
When the spoken script changes, confirm that the captions are updated before exporting again.
Add Music
Music can be imported or created and placed on its own timeline track.
Add music only after the narration has been approved. This makes it easier to judge speech clarity without background audio.
When adding a track:
- Place it underneath the narration.
- Trim it to the required duration.
- Lower its volume.
- Fade it in and out when appropriate.
- Listen through headphones and regular speakers.
- Confirm that every spoken word remains understandable.
Background music should support the project rather than compete with the narrator.
Add Sound Effects
Sound effects can be imported or generated and placed on a separate track. Studio allows users to combine them with narration, video, music, and captions in the same production.
Sound effects are useful for:
- Scene transitions
- Interface demonstrations
- Environmental ambience
- Character actions
- Notifications
- Dramatic moments
- Podcast production
Use them selectively. Too many effects can distract from the information or make a professional tutorial feel less polished.
Review the Project Efficiently
Studio currently allows playback-speed adjustment between approximately 0.8× and 2.0×. This can help users review long projects more efficiently without changing the actual exported narration speed.
Use faster playback to check:
- Missing sections
- Repeated paragraphs
- Long empty gaps
- Incorrect clip order
- Overall project structure
Return to normal speed when checking pronunciation, emotion, timing, music levels, and transitions.
Share the Project for Feedback
Studio includes collaboration tools that allow a project to be shared with others for feedback and comments.
Before sharing:
- Generate the sections that should be reviewed.
- Lock approved material.
- Name the chapters and scenes clearly.
- Explain which areas still need feedback.
- Avoid presenting unfinished audio as final.
Ask reviewers to comment on specific issues such as pronunciation, pacing, clarity, speaker selection, or synchronization rather than simply asking whether the project sounds good.
Review the Credit Cost
Studio may charge credits when narration or other generated content is created. When exporting, any paragraphs that have not yet been converted—or that were edited after conversion—may be generated automatically and charged at that time. ElevenLabs displays the required cost before the export is confirmed.
Before exporting:
- Confirm that the correct voice is applied.
- Check every ungenerated paragraph.
- Remove unused text.
- Review edited sections.
- Make sure you have enough credits.
- Avoid exporting unfinished chapters accidentally.
When every paragraph has already been generated, exporting the project does not require additional speech-generation credits.
Export the Project
When the project is ready:
- Click Export.
- Choose the current chapter or complete project.
- Review the displayed generation cost.
- Select the available output format.
- Confirm the export.
- Wait for processing to finish.
- Download the completed file.
- Listen to the exported version from beginning to end.
Audio-only projects can currently be exported as MP3 or WAV. Projects containing video tracks or captions may offer a video export. Free and Starter-plan video exports include a watermark, while watermark-free video exports require a Creator plan or higher.
Export Chapters Separately
A project containing several chapters can be exported:
- One chapter at a time
- As one complete file
- As a ZIP file containing separate chapter files
Single-chapter projects generally provide MP3 or WAV audio options, or a video option when the timeline contains applicable video or caption content.
Exporting chapters separately is useful for courses, podcast sections, modular training, or audiobooks that will be uploaded one chapter at a time.
Export Again After Making Changes
Editing the Studio project does not automatically update a file that was already downloaded. After changing narration, timing, captions, video, music, or another element, create a new export to include the revisions.
To download the updated version:
- Click Export.
- Review the new credit cost.
- Create another export.
- Wait for rendering to finish.
- Download the newly generated file.
- Confirm that the correction appears.
Earlier exports remain available under the History tab inside the Export area.
Troubleshoot Export Problems
When the download does not begin:
- Confirm that you used Export rather than looking for the normal Text to Speech download button.
- Wait for the rendering process to finish.
- Check the Export History.
- Allow pop-ups and downloads from ElevenLabs.
- Temporarily disable an ad blocker.
- Try another supported browser.
ElevenLabs notes that browser settings, pop-up blockers, ad blockers, and certain browsers can interfere with Studio downloads.
Save a Local Project Archive
After exporting, save:
- The final audio or video
- Individual chapter exports
- The complete script
- Voice names and IDs
- Speech models
- Voice settings
- Pronunciation rules
- Music and sound-effect files
- Source images and videos
- Notes about revisions
Use clear filenames such as:
ElevenLabs-Course-Chapter-01-Final.wav
ElevenLabs-Course-Full-Project-V2.mp3
ElevenLabs-Tutorial-Final-Video.mp4
A local archive protects the project if a voice, asset, or earlier generation later becomes unavailable.
Complete a Studio Project Check
Before approving the project, confirm that:
- Every chapter is in the correct order.
- The same narrator and settings are used consistently.
- Names, numbers, and technical terms are pronounced correctly.
- No paragraphs are duplicated or missing.
- Approved sections are locked.
- Dialogue speakers are assigned correctly.
- Music does not overpower the narration.
- Sound effects appear at the correct time.
- Captions match the spoken audio.
- Video and narration remain synchronized.
- Empty gaps have been removed.
- The export cost has been reviewed.
- The latest changes appear in the exported file.
- The final project has been saved outside ElevenLabs.
ElevenCreative Studio provides more control than the basic Text to Speech Playground, but beginners should build the project in stages. Approve the narration first, then add video, captions, sound effects, and music after the spoken content is accurate and consistent.
Generate Sound Effects and Music
ElevenLabs can generate custom sound effects and complete music tracks from written descriptions. These tools are useful for videos, podcasts, audiobooks, games, advertisements, social-media content, and longer projects created inside ElevenCreative Studio.
Sound Effects is best for individual noises, ambience, transitions, and short musical elements. Eleven Music is designed for complete instrumental tracks or songs with vocals.
Open Sound Effects
- Sign in to ElevenLabs.
- Open Playground from the sidebar.
- Select Sound Effects.
- Enter a description of the sound you need.
- Review the duration, looping, and prompt-influence settings.
- Click the generation arrow.
- Listen to the generated variations.
- Download or save the strongest result.
The ElevenLabs website currently creates four variations each time a sound effect is generated. Previous results can be reopened through the History tab.
Describe One Clear Sound
For a basic effect, describe the source of the sound and what it is doing.
For example:
Glass shattering on a concrete floorHeavy wooden door slowly creaking openSoft footsteps moving across wet grassDistant thunder during a quiet rainstormShort digital confirmation sound for a mobile app
Clear prompts normally work better than vague descriptions such as:
Scary sound
A stronger version would be:
Low cinematic rumble followed by a sharp metallic impact, dark and suspenseful
ElevenLabs understands ordinary language as well as audio-production terminology.
Add Useful Details
Include details that affect how the sound should be produced.
You can describe:
- The object creating the sound
- The surface or environment
- The distance from the listener
- The intensity
- The speed
- The emotional mood
- Whether the effect should be realistic or cinematic
- Whether it should be a single impact or continuous ambience
- The approximate duration
- The intended use
For example:
Close-up recording of high heels walking slowly across a marble hallway, realistic indoor reverb
This gives the model more direction than:
Footsteps
Use Audio Terminology
Audio terms can make prompts more specific.
Useful terms include:
- Impact: A collision, hit, crash, or contact sound
- Whoosh: A fast movement through the air
- Ambience: Continuous environmental background sound
- One-shot: A single sound that does not repeat
- Loop: Audio intended to repeat continuously
- Stem: An isolated musical or audio component
- Braam: A large cinematic brass-like impact
- Glitch: A digital malfunction or distorted transition
- Drone: A continuous atmospheric tone
For example:
Dark cinematic drone with a low sub-bass texture and distant metallic ambience
or:
Fast futuristic whoosh, one-shot transition effect
ElevenLabs officially recommends these kinds of production terms when greater control is needed.
Keep the Prompt Focused
The Sound Effects prompt currently supports up to 450 characters. More detail can help, but a long prompt containing several unrelated events may produce a confused result.
When a scene needs several sounds, generate them separately.
Instead of:
A person walks through a hallway, unlocks a door, opens it, enters a kitchen, drops a glass, and a dog starts barking
Create separate effects for:
- Hallway footsteps
- Keys and door unlocking
- Door opening
- Glass dropping
- Dog barking
The individual effects can then be arranged on separate tracks inside Studio or another editor. ElevenLabs notes that complex sequences can be described in one prompt, but separate generations generally provide more control.
Set the Duration
Sound effects can currently be generated at a specified length of up to 30 seconds. Selecting Auto allows ElevenLabs to determine an appropriate length from the prompt.
Use shorter durations for:
- Button clicks
- Impacts
- Door sounds
- Notification tones
- Transitions
- Individual footsteps
- Object movements
Use longer durations for:
- Rain
- Wind
- Crowd ambience
- Room tone
- Forest environments
- Machinery
- Suspenseful drones
Do not automatically select 30 seconds for every effect. Longer generations consume more credits when duration is manually specified and may contain unnecessary material.
Understand the Sound-Effect Cost
On the ElevenLabs website, allowing the platform to choose the duration currently costs 200 credits per generation. Setting a specific duration costs 40 credits for each second, with four effects produced in that website generation.
For example, manually selecting ten seconds would cost more than a short automatic generation.
Before generating:
- Confirm that the prompt is correct.
- Select the shortest practical duration.
- Decide whether Auto is sufficient.
- Review the displayed cost.
- Generate only after the settings are ready.
The cost is based on the duration settings rather than the length of the written prompt.
Turn On Looping
Enable Looping when the sound should repeat continuously without an obvious beginning or ending.
Looping works well for:
- Rain
- Ocean waves
- Wind
- Forest ambience
- Traffic
- Office background noise
- Mechanical hum
- Game environments
- Atmospheric drones
ElevenLabs blends the end of the generated effect into its beginning so it can repeat more smoothly. This allows a sound of up to 30 seconds to be extended throughout a longer scene.
Do not use looping for a one-time event such as a glass breaking, door slamming, explosion, or notification sound.
Adjust Prompt Influence
Prompt Influence determines how strictly the generated effect follows the description.
A higher value creates a more literal interpretation of the prompt. A lower value gives the model more creative freedom and may introduce additional variation. The website’s current default is approximately 30%.
Increase Prompt Influence when:
- A specific object must be audible.
- The generated effects contain unrelated sounds.
- The timing or environment needs to follow the prompt more closely.
- You are creating a practical Foley effect.
Lower it when:
- The result sounds too plain.
- You want more creative cinematic variations.
- The exact interpretation is flexible.
- You are exploring possible effects rather than recreating one specific noise.
Change the setting gradually and compare the results.
Generate and Compare the Variations
After selecting Generate, listen to all four versions.
Compare:
- Realism
- Background noise
- Timing
- Intensity
- Whether the correct object is audible
- Whether the effect begins and ends cleanly
- Whether a looping version repeats smoothly
- Whether the sound fits the video or scene
Do not choose the first result automatically. One variation may have a cleaner ending, better timing, or more natural texture.
Save Favorites
Select the star icon to save a useful sound effect to the Favorites tab. This is helpful when exploring several possibilities before deciding which one belongs in the final project.
Save effects that may be useful later, such as:
- Brand notification sounds
- Standard transitions
- Background ambience
- Interface clicks
- Repeated character effects
- Podcast intro elements
Use descriptive filenames after downloading them so they remain recognizable outside ElevenLabs.
Download the Sound Effect
Generated sound effects can currently be downloaded as:
- MP3 at 44.1 kilohertz
- WAV at 48 kilohertz
WAV is available for non-looping effects, while MP3 is available more generally.
Use MP3 for smaller files and simple online content. Use WAV when the effect will receive substantial editing, mixing, or professional processing.
Example filenames include:
Wooden-Door-Creak-Variation-Three.wav
Soft-Rain-Loop-Final.mp3
Mobile-App-Confirmation-Sound.wav
Add the Effect to Studio
After approving the sound:
- Open the ElevenCreative Studio project.
- Add or select the Sound Effects track.
- Upload the downloaded effect or generate one from within Studio.
- Move it to the correct position.
- Trim unnecessary material.
- Adjust its volume.
- Add a fade when needed.
- Play it together with the narration and music.
Studio includes separate timeline tracks for narration, music, video, captions, and sound effects.
An effect that sounds good by itself may be too loud when placed underneath speech. Always review it inside the complete scene.
Generate Musical Elements With Sound Effects
Sound Effects can also create short musical components such as:
- Drum loops
- Bass lines
- Synth pads
- Brass hits
- Percussion
- Trailer impacts
- Short melodic samples
For example:
Nineties hip-hop drum loop, eighty-eight beats per minute, dusty vinyl texture
or:
Atmospheric synth pad in D minor with slow modulation
Use Sound Effects for individual musical components. Use Eleven Music when you need a complete song or structured instrumental track.
Create Music With Eleven Music
Eleven Music generates complete music from written prompts. It can create instrumental tracks, songs with vocals, multilingual lyrics, and structured compositions containing sections such as introductions, verses, choruses, bridges, solos, and outros.
Open Eleven Music
- Select Music from the ElevenLabs sidebar.
- Choose to begin a new song.
- Enter a description of the track.
- Add an optional Audio Reference or Music Finetune.
- Choose the number of variations.
- Select the duration.
- Generate the music.
- Open the strongest result in the editor.
- Refine its sections, styles, or lyrics.
- Download or share the finished version.
Eleven Music is currently available to all users through the ElevenLabs website. Music API access requires a paid subscription.
Decide Between Instrumental Music and Vocals
State clearly whether the track should contain singing.
For instrumental music, include phrases such as:
- Instrumental only
- No vocals
- Background score
- Underscore
- Ambient instrumental
- No spoken words
For a vocal track, describe:
- Vocal gender or presentation
- Vocal style
- Language
- Emotional delivery
- Whether the vocals should be solo or layered
- The topic of the lyrics
- The song structure
For example:
Upbeat modern pop song with confident female vocals, English lyrics about starting a new business, bright synthesizers, rhythmic bass, and an energetic chorus
Eleven Music supports both instrumental generation and multilingual vocals.
Write a Detailed Music Prompt
A useful music prompt can include:
- Genre
- Subgenre
- Mood
- Tempo
- Instruments
- Vocal style
- Language
- Song structure
- Production style
- Intended use
- Duration
- Type of ending
For example:
Warm, uplifting corporate instrumental for a beginner technology tutorial. Light piano, soft electronic drums, gentle acoustic guitar, and subtle synth textures. Medium tempo, polished modern production, no vocals, calm introduction, and a clean ending.
This is stronger than:
Background music for a video
ElevenLabs recommends combining genre, mood, instrumentation, tempo, and use case when describing the desired track.
Do Not Overcomplicate the Prompt
A detailed prompt can help, but longer does not always mean better. ElevenLabs notes that the relationship between prompt length and music quality is not absolute. A clear high-level description can sometimes produce a strong result without an extremely technical prompt.
Start with the elements that matter most:
- Genre
- Mood
- Main instruments
- Vocals or instrumental
- Intended use
Add more production details only when the first generation does not match the idea.
Specify the Intended Use
Telling Eleven Music how the track will be used can influence its structure and intensity.
Examples include:
- Background music for a tutorial
- Podcast introduction
- Dramatic film scene
- Mobile-game battle music
- Calm meditation
- Product advertisement
- Fashion video
- YouTube outro
- Audiobook ambience
- Social-media promotion
For example:
Thirty-second energetic electronic track for a software product advertisement, immediate attention-grabbing opening, strong rhythmic build, and clean final hit
A tutorial needs different music from a movie trailer, even when both use electronic instruments.
Select the Duration
Eleven Music can generate tracks ranging from approximately three seconds to five minutes. Users can select a fixed length or allow the platform to determine the duration automatically.
For a first test, generate approximately 30 seconds. ElevenLabs recommends beginning with a shorter section and extending the composition gradually when more control is needed.
Use:
- 5–15 seconds for a logo sound or transition
- 15–30 seconds for a social-media clip
- 30–60 seconds for an advertisement or introduction
- Longer tracks for podcasts, videos, films, games, or standalone songs
Do not generate five minutes before confirming that the genre, instrumentation, and mood are correct.
Choose the Number of Variations
The Music interface allows users to select how many versions should be created. More variations provide additional choices but also use more credits.
For a beginner:
- Start with one or two versions.
- Compare the structure and musical direction.
- Select the stronger foundation.
- Edit that version instead of generating many complete tracks.
A weaker song may need a clearer prompt rather than ten nearly identical generations.
Use an Audio Reference
Music v2 allows users to upload a short reference track of approximately 30 seconds. The reference guides characteristics such as instrumentation, tempo, mood, production style, and overall sound. It does not directly copy or remix the uploaded recording. Every Audio Reference is screened for copyright compliance.
Use an Audio Reference when:
- You created an original musical example.
- You own or control the uploaded recording.
- You need the new track to follow a specific production direction.
- Written descriptions are not capturing the intended sound.
Do not upload copyrighted commercial music that you do not have the right to use. Use your own original recording or an authorized reference.
Audio Reference is currently available on Music v2 plans, including the Free plan.
Choose a Music Finetune
A Music Finetune guides the generation toward a particular stylistic identity.
The interface may provide:
- Curated Finetunes created by ElevenLabs for selected musical styles
- Custom Finetunes trained using the user’s own original music
Selecting a Finetune is optional. When none is selected, Eleven Music uses its standard model.
Beginners should first test the standard model or a curated option. Custom Finetunes are more useful after developing a collection of original music and a consistent production style.
Review the Generated Track
After generation, listen from beginning to end.
Check:
- Whether the genre matches the prompt
- Whether the mood fits the project
- Whether the instruments are appropriate
- Whether the vocals are understandable
- Whether the lyrics follow the intended topic
- Whether the track changes direction unexpectedly
- Whether the ending is clean
- Whether the duration is useful
- Whether the music leaves space for narration
Background music should not contain distracting vocals or dramatic changes unless those elements are intentional.
Open the Music Editor
Eleven Music allows users to refine the track without rebuilding the entire song from the beginning.
The editor can be used to:
- Add sections
- Remove sections
- Extend sections
- Shorten sections
- Change lyrics
- Edit instrumental prompts
- Include specific musical styles
- Exclude unwanted styles
- Request changes through conversational instructions
- Generate another version of the same prompt
These controls make it possible to preserve a strong song while changing only the weak section.
Add a New Section
To extend the song:
- Find the section after which the new material should appear.
- Click the available + control.
- Add the new section.
- Drag it to change its duration.
- Enter lyrics or an instrumental instruction.
- Generate the revised track.
You might add:
- Another verse
- A second chorus
- A bridge
- An instrumental solo
- A breakdown
- An outro
- A transition
For an instrumental section, use a bracketed description such as:
[Energetic electric guitar solo with rising drums]
Eleven Music supports building songs section by section, which provides more control than creating the complete composition in one request.
Edit Lyrics
Select the text box for the affected section and rewrite the lyrics.
Review for:
- Grammar
- Repetition
- Rhyme
- Pronunciation
- Brand names
- Numbers
- Unwanted claims
- Language consistency
- Whether the lyrics fit the available time
Do not assume generated lyrics are ready to publish without editing. Review them as carefully as any other written content.
Edit an Instrumental Section
Instrumental sections use descriptions instead of sung lyrics.
For example:
[Soft piano introduction with distant ambient strings]
or:
[Drums become more energetic while the bass and synthesizers gradually build]
Change only the affected section when the rest of the composition already works.
Include or Exclude Styles
Individual sections can contain instructions for sounds that should be included or avoided.
Include examples:
- Gradual filter opening
- Fading hi-hats
- Long vocal delay
- Soft piano
- Distorted bass
- Layered harmonies
Exclude examples:
- Abrupt ending
- New instruments
- Heavy drums
- Vocals
- Sudden tempo changes
- Distorted guitar
Section-level style control is useful when one part of the song does not match the rest.
Make Changes With Direct Prompts
The conversation area in the Music editor accepts natural-language revision requests.
For example:
Make the chorus more energetic.Add a guitar solo after the second verse.Reduce the drums during the introduction.Make the ending softer and more gradual.Remove the vocals from the bridge.
Be specific about the section and the desired change. A request such as “make it better” does not explain what should be improved.
Generate the Edited Version
Editing the visible structure does not immediately change the audio. After making revisions, click Generate to create a new version that includes them.
Before generating:
- Review every edited section.
- Confirm that no section was deleted accidentally.
- Check the lyrics.
- Review included and excluded styles.
- Confirm the duration.
- Review the credit cost.
- Generate the updated version.
Keep the earlier version until the new track has been reviewed.
Add the Music to a Voiceover
After downloading the music:
- Add it to the Music track inside Studio.
- Position it beneath the narration.
- Lower the volume.
- Trim it to the required duration.
- Add a fade-in and fade-out.
- Review it with the voiceover.
- Reduce busy instruments during important speech.
- Confirm that the narration remains understandable.
Music that sounds quiet by itself may still compete with the human voice because both occupy similar frequency ranges.
For instructional content, simple instrumental music usually works better than music containing prominent vocals.
Download or Share the Track
Select Download to save the finished audio. The Music interface also provides a Share option that creates a link with a customizable visualizer.
Use a clear filename such as:
Vaxali-Tutorial-Background-Music-Final.wav
or:
Podcast-Introduction-Music-Version-Two.mp3
Save the prompt, lyrics, generation settings, and project notes beside the downloaded file.
Review Commercial Usage Rights
ElevenLabs states that Eleven Music is cleared for many commercial uses, including film, television, podcasts, social-media videos, advertising, and games. However, the exact rights available depend on the account’s plan and the applicable Music Terms.
Before publishing or monetizing a track:
- Review the current plan.
- Read the Music Terms.
- Confirm that the intended commercial use is covered.
- Keep records of the account and subscription used.
- Avoid unauthorized copyrighted references.
- Review any platform-specific requirements.
Do not assume that free-plan and paid-plan content always carry identical commercial permissions.
Complete a Sound Effects and Music Check
Before approving the audio, confirm that:
- The sound-effect prompt describes one clear result.
- The selected duration is appropriate.
- Looping is enabled only when needed.
- Prompt Influence produces the intended balance.
- The strongest variation has been saved.
- The music prompt includes the important genre, mood, and instrumentation.
- Vocals or instrumental-only output are specified clearly.
- The track leaves enough space for narration.
- Lyrics have been reviewed.
- Unwanted song sections have been edited.
- The latest edits were regenerated.
- Music and effects do not overpower speech.
- The files have descriptive names.
- The source prompts and settings have been saved.
- The planned usage complies with the current license and subscription terms.
Sound effects and music can make a voiceover feel more polished, but they should support the narration rather than distract from it. Add them only after the spoken content has been reviewed and approved.
- Sign in to ElevenLabs.
- Select Speech to Text from the dashboard.
- Click Transcribe Files.
- Upload an audio or video file.
- Review the available transcription options.
- Click Upload Files to begin processing.
- Open the completed file to review the transcript.
Speech to Text is different from Text to Speech. Text to Speech creates audio from writing, while Speech to Text creates writing from existing audio.
Choose the Right Recording
The quality of the original recording strongly affects the accuracy of the transcript.
Use audio with:
- Clear voices
- Minimal background noise
- Limited echo
- Consistent volume
- Little overlapping dialogue
- A stable internet connection during upload
- Speakers positioned close enough to the microphone
A noisy phone recording can still be transcribed, but it may require more corrections than a clean recording.
When several people speak over one another, Scribe may have difficulty identifying the exact words and assigning them to the correct speaker. Clean separation between speakers usually produces a more useful transcript.
Upload Audio or Video
ElevenLabs accepts both audio and video uploads through the website. The current website guide lists a maximum file size of 3 GB.
Supported formats include common audio and video types such as:
- MP3
- WAV
- M4A
- FLAC
- OGG
- MP4
- MOV
- AVI
- MKV
- WebM
The broader Speech to Text capability supports recordings as long as approximately ten hours in standard processing mode, although upload and account limits can still affect individual jobs.
For a first test, upload a short recording of one to five minutes. A shorter file makes it easier to understand the editor before processing a complete interview, podcast, or course.
Select the Language
During upload, choose the primary language when you already know it.
For example:
- English
- Spanish
- Hungarian
- German
- French
Leave the language set to Detect when the correct language is unknown or when the recording contains several languages. Scribe v2 can automatically identify and transcribe multilingual audio.
Selecting the known language may help reduce uncertainty, especially when the recording is short or contains names and borrowed words from other languages.
Handle Multilingual Recordings
Scribe v2 can recognize language changes within the same recording. This is useful for bilingual interviews, translated conversations, international meetings, and videos containing several languages.
However, review every language section carefully. Accuracy can become weaker when:
- Speakers change languages in the middle of a sentence
- Several accents appear
- The audio is noisy
- Foreign terms are used without context
- Speakers talk over one another
- The recording contains uncommon names
Add important words as keyterms before transcription when possible.
Add Keyterms
Keyterm prompting tells Scribe to pay closer attention to specific names, brands, technical terms, products, locations, or phrases.
Useful keyterms might include:
- Vaxali
- ElevenLabs
- ChatGPT
- Hungarian surnames
- Employee names
- Medical terminology
- Software names
- Industry abbreviations
Add the term using its correct spelling. This increases the chance that the transcript will recognize it instead of replacing it with a more common word.
ElevenLabs currently supports extensive keyterm prompting with Scribe v2. The option may increase the transcription cost, so review the displayed usage before submitting a long recording.
Enable Audio-Event Tags
Turn on Tag Audio Events when non-speech sounds provide useful context.
Scribe can identify events such as:
- Laughter
- Applause
- Coughing
- Music
- Door sounds
- Background reactions
- Other audible events
These labels can be helpful for interviews, podcasts, films, accessibility transcripts, and subtitle preparation.
Leave the option off when you need only clean spoken text and do not want environmental sounds included.
Begin the Transcription
After reviewing the settings:
- Confirm that the correct file is attached.
- Select or detect the language.
- Add important keyterms.
- Choose whether to tag audio events.
- Review the expected usage or cost.
- Click Upload Files.
- Wait for the file to finish processing.
Long recordings take longer to process than short clips. Do not upload the same file repeatedly because each successful transcription can consume additional credits.
Speech to Text billing is based on the duration of the submitted recording rather than the number of words produced. Rates can vary by model, plan, and optional features.
Open the Completed Transcript
When processing finishes, select the uploaded filename to open the result.
The transcript displays the spoken words alongside their timing information. Clicking a word can begin playback near the point where that word was spoken, making corrections easier.
Listen while reading rather than reviewing the text alone. A sentence may look reasonable while still differing from what the speaker actually said.
Review Speaker Separation
Scribe uses speaker diarization to identify when different people are speaking. Scribe v2 currently supports automatic separation for as many as 32 speakers.
The initial transcript may label speakers as:
- Speaker 1
- Speaker 2
- Speaker 3
Listen to several sections before renaming them. Do not assume that the first voice always belongs to the interviewer or host.
Speaker separation may require corrections when:
- Voices sound similar
- Speakers interrupt each other
- Someone moves away from the microphone
- The recording includes phone or video-call compression
- Background voices are audible
- One person changes their speaking style
Rename the Speakers
To rename a detected speaker:
- Open the transcript.
- Find the Speakers area.
- Click the edit control.
- Replace the automatic label with the correct name or role.
- Review the other segments assigned to that speaker.
Use labels such as:
- Interviewer
- Guest
- Instructor
- Student
- Customer
- Support Agent
Use actual names only when they are appropriate for the intended transcript and privacy requirements. ElevenLabs allows speaker names to be edited after transcription.
Edit the Transcript Text
The Transcript Editor allows direct text editing.
- Click inside the incorrect passage.
- Delete or replace the wrong wording.
- Add missing punctuation.
- Correct capitalization.
- Fix names and technical terms.
- Continue playing the recording to verify the change.
The editor uses a direct on-screen editing system, so text can be changed by clicking and typing.
Review carefully for common errors such as:
- Similar-sounding words
- Incorrect names
- Missing negative words such as “not”
- Incorrect prices or dates
- Wrong speaker labels
- Missing punctuation
- Confused acronyms
- Repeated phrases
Small transcription mistakes can significantly change the meaning of legal, financial, medical, employment, or instructional content.
Adjust Segment Timing
Each transcript section has a beginning and ending timestamp.
To change the timing:
- Select the affected segment.
- Drag its handles on the timeline.
- Move the beginning to the correct point.
- Move the ending after the speaker finishes.
- Play the section again.
Exact timestamps can also be entered through the transcript’s side panel.
Accurate timing becomes particularly important when the transcript will be exported as subtitles.
Split a Segment
Split a segment when it contains two separate ideas, speakers, or subtitle lines.
- Click inside the text where the split should occur.
- Press Enter.
- Review the new segments.
- Confirm that each section begins and ends at the correct time.
- Assign the correct speaker.
Splitting long segments can improve readability and create shorter subtitle lines.
Merge Segments
Merge segments when one sentence has been divided unnecessarily.
Two segments can be merged when:
- They are next to each other.
- They belong to the same speaker.
Select the merge control and review the combined timing afterward.
Do not merge long passages simply to reduce the number of segments. Shorter sections are usually easier to edit and subtitle.
Add or Delete Segments
The Transcript Editor also allows users to add a missing segment or remove an unnecessary one.
To add a segment:
- Select Add Segment.
- Choose its location on the timeline.
- Enter the missing text.
- Assign the correct speaker.
- adjust its timestamps.
To remove one, select the segment and use the Delete option.
Be careful when deleting a speaker. ElevenLabs warns that deleting a speaker also deletes all transcript segments assigned to that speaker.
Reassign Incorrect Speakers
When a passage has been assigned to the wrong person:
- Select the speaker indicator beside the segment.
- Choose the correct speaker.
- Repeat for other affected segments.
The editor also allows all segments belonging to one speaker label to be moved to another speaker in bulk.
Use the bulk option only after confirming that every segment under the original label belongs to the same person.
Realign Words After Editing
Changing a transcript can cause the written words and original timestamps to stop matching.
After making corrections, use Align Words to calculate updated word-level timing for the segment.
Realignment is especially important before exporting:
- SRT subtitles
- VTT captions
- Timestamped transcripts
- Searchable video text
- Content that will be synchronized with playback
Listen through the corrected segment after realignment to confirm that the timing remains accurate.
Adjust the Playback Speed
The Transcript Editor allows the source recording to be played at a different review speed.
Use slower playback when:
- A speaker talks quickly.
- Several people overlap.
- The recording contains an unfamiliar accent.
- A name is difficult to understand.
- A sentence is unclear.
Use faster playback when reviewing a long, clearly spoken recording for general mistakes.
Changing the review speed does not change the original file or exported transcript.
Export the Transcript
After editing, select Export in the upper-right corner.
The Transcript Editor currently supports:
- Plain text
- JSON
- HTML
- SRT
- VTT
Use plain text for notes, articles, summaries, and general editing.
Use JSON when structured transcript data and timestamps are needed for software or development.
Use HTML when the transcript will be displayed as a webpage or formatted document.
Use SRT for subtitles supported by many video editors and platforms.
Use VTT for web-based video captions and platforms that prefer WebVTT.
Create Subtitles
The Transcript Editor can also create and edit subtitles. Select the plus control beside Subtitles to add them to the transcript project.
Review subtitles separately from the full transcript. Good subtitles should:
- Remain on screen long enough to read
- Avoid covering important visuals
- Use manageable line lengths
- Match the spoken words
- Identify speakers when necessary
- Include useful sound labels
- Appear at the correct time
A transcript paragraph that works well for reading may still be too long for an individual subtitle.
Save a Local Copy
After finishing the transcript, save:
- The original audio or video
- The corrected plain-text transcript
- The subtitle file
- The speaker list
- Important terminology
- The final exported version
- Notes about uncertain passages
Use clear filenames such as:
Vaxali-Interview-Corrected-Transcript.txt
Vaxali-Interview-English-Subtitles.srt
Podcast-Episode-Four-Transcript-Final.html
Do not rely on only one online copy for important business, legal, research, or client material.
Review Sensitive Recordings Carefully
Audio may contain names, phone numbers, financial details, health information, private conversations, or other sensitive material.
Before uploading, confirm that:
- You have permission to process the recording.
- The account and workspace are correct.
- The file does not include unnecessary confidential sections.
- The finished transcript will be stored securely.
- The transcript will be shared only with authorized people.
Organizations requiring HIPAA-compliant processing must contact ElevenLabs and complete a Business Associate Agreement before using the service for applicable healthcare workflows.
Complete a Speech-to-Text Check
Before approving the transcript, confirm that:
- The correct recording was uploaded.
- The language was selected or detected properly.
- Important names and terms were added or corrected.
- Every speaker has the right label.
- Numbers, dates, and prices match the audio.
- Overlapping dialogue has been reviewed manually.
- Segment timestamps are accurate.
- Edited words have been realigned.
- Subtitle lines remain readable.
- The correct export format has been downloaded.
- The original recording and corrected transcript are stored securely.
Speech to Text provides a fast starting point, but important transcripts still require human review. Once the transcript is accurate, the next step is learning how to remove background noise and isolate the main speaker with Voice Isolator.
Remove Background Noise With Voice Isolator
Voice Isolator removes background noise and separates the main spoken voice from the rest of an audio recording. It is useful for cleaning interviews, podcasts, voiceovers, videos, phone recordings, and audio recorded in noisy environments.
The tool can reduce sounds such as:
- Wind
- Traffic
- Office noise
- Microphone feedback
- Background music
- Environmental ambience
- Other nearby conversations
The result is a new audio file containing a clearer version of the main speech. Voice Isolator works best when the intended speaker remains understandable in the original recording.
Open Voice Isolator
- Sign in to ElevenLabs.
- Open Audio Tools from the sidebar.
- Select Voice Isolator.
- Upload an existing recording or record a new one.
- Review the uploaded audio.
- Click Isolate Voice.
- Wait for the file to process.
- Play the cleaned result.
- Download it when the speech sounds clear.
ElevenLabs allows users to upload a file, drag and drop one into the workspace, or record directly through their device’s microphone.
Choose a Suitable Recording
Voice Isolator can improve noisy audio, but it cannot perfectly recover speech that was never captured clearly.
The best source recording contains:
- One main speaker
- Understandable speech
- Limited clipping or distortion
- Minimal overlapping dialogue
- Consistent microphone volume
- No extremely loud sound covering the voice
A recording containing moderate traffic, wind, fan noise, or background music may clean up well. A recording in which another person speaks directly over the main speaker may be more difficult because both voices occupy the same part of the audio.
Upload the File
Drag the recording into Voice Isolator or select the upload control and locate it on the device.
ElevenLabs currently accepts Voice Isolator uploads up to 500 MB or one hour in length. Longer files should be divided into smaller sections before processing.
Splitting a long recording can also make it easier to:
- Review individual sections
- Correct one noisy passage
- Avoid processing unnecessary silence
- Organize interviews or podcast chapters
- Replace only the sections that need cleaning
Cut the file between sentences or natural topic changes rather than in the middle of a word.
Record Directly in ElevenLabs
Voice Isolator also allows users to create a new recording with their device’s microphone.
Before recording:
- Select the correct microphone.
- Allow the browser to access it.
- Move close enough to the microphone.
- Reduce unnecessary noise when possible.
- Record a short test.
- Listen before creating the full recording.
Voice Isolator can reduce background noise afterward, but beginning with a cleaner recording usually produces a more natural result.
Trim Unnecessary Audio First
Before uploading, remove:
- Long silence
- Failed recording attempts
- Unneeded introductions
- Empty sections
- Audio after the conversation ends
- Sections containing no useful speech
Voice Isolator charges according to the recording’s duration, so removing unnecessary material can reduce credit usage.
The current cost is 1,000 credits per minute of processed audio.
For example, processing a five-minute recording uses approximately 5,000 credits.
Isolate the Voice
Once the recording is uploaded:
- Confirm that it is the correct file.
- Review its length.
- Check the expected credit cost.
- Click Isolate Voice.
- Allow the processing to finish.
- Play the isolated result.
The tool creates a new file rather than permanently changing the original recording. Keep the original until the cleaned version has been reviewed.
Compare the Original and Cleaned Versions
Listen to the same passage in both files.
Check whether Voice Isolator successfully reduced:
- Background conversations
- Wind
- Road noise
- Music
- Room ambience
- Electrical hum
- Microphone noise
- Other distractions
Also listen for problems introduced during isolation, such as:
- Metallic speech
- Missing syllables
- Unnatural silence
- Changing volume
- Muffled words
- Robotic artifacts
- Cut-off breaths
- Parts of another speaker remaining
Noise removal is useful only when the cleaned recording remains easy to understand.
Review the Entire File
Do not approve the result after listening to only the first few seconds. Background noise may change throughout the recording.
Review:
- The beginning
- Several middle sections
- The noisiest passage
- Quietly spoken sentences
- Areas containing music
- The ending
A file may sound clean during one section but contain artifacts where the noise becomes louder or overlaps with speech.
Handle Background Conversations Carefully
Voice Isolator can reduce background chatter, but separating two people who speak at similar volume can be difficult.
When another voice remains audible:
- Process a shorter section.
- Check whether the intended speaker is louder.
- Trim sections where the other speaker is unnecessary.
- Use an audio editor to lower the remaining background.
- Re-record the affected line when possible.
Do not expect the tool to reconstruct words that were completely covered by another speaker.
Clean Wind and Outdoor Noise
Wind can create low-frequency rumbling and sudden bursts that partially hide speech.
For better outdoor recordings:
- Use a windscreen on the microphone.
- Turn away from strong wind.
- Keep the microphone close to the speaker.
- Avoid touching the microphone.
- Record a short test first.
Voice Isolator can reduce street and wind noise, but preventing those sounds during recording will usually provide a clearer final voice. ElevenLabs presents wind and city-street recordings as supported use cases for the tool.
Remove Background Music
Voice Isolator can reduce music beneath spoken audio. This may be useful when:
- Editing an interview
- Recovering dialogue from a video
- Preparing audio for transcription
- Replacing an existing soundtrack
- Cleaning a podcast clip
Music that is significantly louder than the voice or contains prominent vocals may be more difficult to remove cleanly.
After isolation, listen for faint instrumental sounds and changes in the speaker’s tone. Additional editing may still be necessary before publishing the recording.
Understand Overlapping Vocals
When a song contains singing, Voice Isolator may interpret the vocals as speech or another voice that should remain. This can make it harder to separate a narrator from music containing lyrics.
Instrumental background music is generally easier to distinguish from spoken dialogue than another human voice.
When possible, use:
- The original narration track
- A version without background music
- Separate audio stems
- A clean microphone recording
Voice Isolator should be used as a cleanup tool rather than a replacement for properly separated source tracks.
Use Voice Isolator Before Speech to Text
Cleaning a noisy recording before transcription can improve the quality of the audio submitted to Speech to Text.
A practical workflow is:
- Upload the noisy recording to Voice Isolator.
- Generate the cleaned speech.
- Download the isolated file.
- Upload it to Speech to Text.
- Add important keyterms.
- Generate the transcript.
- Review names, numbers, and speaker labels manually.
This is especially useful for interviews, meetings, podcasts, and outdoor recordings.
Use Voice Isolator Before Voice Changer
Voice Changer performs better when the source contains clear speech without music or competing voices.
Use this workflow:
- Clean the recording with Voice Isolator.
- Download the isolated speech.
- Open Voice Changer.
- Upload the cleaned file.
- Select the authorized output voice.
- Generate the transformed version.
- Compare it with the source.
This helps Voice Changer focus on the original speaker’s timing, emotion, and delivery rather than the surrounding background audio.
Use Voice Isolator for Voice-Library Search
ElevenLabs allows users to search the Voice Library using an audio sample. When the only available sample contains background noise, ElevenLabs recommends cleaning it with Voice Isolator before using it for voice identification or similarity search.
A cleaner sample gives the search tool more useful vocal information and reduces interference from music or environmental sounds.
Prepare Cleaner Voice-Cloning Samples
Voice Isolator may also help clean a recording before testing an authorized voice clone. However, heavily processed audio is not always ideal for cloning.
Before using isolated audio:
- Listen for metallic artifacts.
- Confirm that no words were damaged.
- Compare it with the original.
- Use a naturally clean recording when one is available.
- Avoid combining isolated audio with recordings of very different quality.
For Professional Voice Cloning, recording new clean samples is generally preferable to relying on severely noisy audio that required heavy restoration.
Add the Cleaned Audio to Studio
After downloading the isolated file:
- Open the ElevenCreative Studio project.
- Upload the cleaned recording.
- Place it on the correct timeline track.
- Align it with the video or other audio.
- Trim unnecessary silence.
- Adjust the volume.
- Add music or effects afterward.
- Review the complete scene.
Do not immediately add loud background music back underneath the cleaned speech. Set the narration level first, then introduce other tracks gradually.
Adjust the Volume After Isolation
The cleaned file may sound quieter or louder than other project audio.
Inside Studio or an audio editor:
- Compare it with the surrounding clips.
- Adjust the gain gradually.
- Avoid pushing the volume into distortion.
- Add a short fade at edited boundaries.
- Listen through headphones and normal speakers.
Noise removal and volume balancing are separate steps. A clean recording may still need level adjustment before it matches the rest of a project.
Download the Isolated Voice
After reviewing the result:
- Click the download control.
- Save the file in the correct project folder.
- Rename it clearly.
- Keep the original recording separately.
ElevenLabs allows the processed result to be played in the app or downloaded for use in another project.
Useful filenames include:
Podcast-Interview-Isolated-Voice.wav
Outdoor-Recording-Cleaned-V2.mp3
Tutorial-Narration-Noise-Removed.wav
Avoid replacing the original file with the cleaned version. Keeping both allows you to return to the source when the isolation removed something important.
Save the Files in Separate Folders
A simple folder structure is:
- Original recordings
- Isolated voices
- Edited audio
- Transcripts
- Voice Changer output
- Final exports
This prevents an automatically cleaned version from being mistaken for an untouched source file.
Troubleshoot Weak Results
When too much noise remains:
- Process a shorter section.
- Trim sections with no useful speech.
- Start with a higher-quality source.
- Remove particularly loud sounds manually.
- Re-record the audio when possible.
When the speech sounds metallic:
- Compare it with the original.
- Avoid processing the same file repeatedly.
- Use less damaged source audio.
- Replace only the noisiest sections.
- Keep some natural room tone during editing.
When words disappear:
- Check whether the original speech was covered by noise.
- Restore the original passage.
- Re-record the missing sentence.
- Use subtitles or a transcript when rerecording is impossible.
When another speaker remains:
- Isolate a section where the main person speaks alone.
- Edit overlapping passages manually.
- Avoid using the result for voice cloning without careful review.
Do Not Process the File Repeatedly
Running an already isolated file through Voice Isolator several times may remove additional vocal detail and make the result sound less natural.
A better workflow is:
- Keep the original recording.
- Process it once.
- Compare the result.
- Edit remaining problems manually.
- Re-record severely damaged passages when possible.
Repeated processing also uses additional credits.
Protect Private Recordings
Before uploading interviews, calls, meetings, or personal recordings, confirm that you have permission to process the audio.
Recordings may contain:
- Personal names
- Contact information
- Financial details
- Workplace conversations
- Health information
- Private family discussions
- Confidential business material
Use the correct ElevenLabs account and workspace, limit access to the finished files, and store downloads securely.
Complete a Voice-Isolation Check
Before using the cleaned recording, confirm that:
- The correct file was processed.
- The recording is within the upload limits.
- Unnecessary silence was removed.
- The displayed credit cost was reviewed.
- The main speaker remains clear.
- Important words were not removed.
- Background noise is less distracting.
- No severe metallic artifacts were introduced.
- The entire file was reviewed.
- The cleaned and original versions were saved separately.
- The recording is authorized for processing.
- The isolated audio has been tested in its final project.
Voice Isolator can make noisy recordings easier to understand, transcribe, transform, and edit. It works best when the main voice is already audible and the tool is used to reduce distractions rather than recover completely hidden speech.
Translate and Dub Audio or Video
ElevenLabs Dubbing translates spoken content into another language while attempting to preserve each speaker’s original voice, tone, emotion, pacing, and delivery. It can also separate dialogue from background music and sound effects so the original soundtrack remains part of the finished dub. The current automatic dubbing system supports more than 90 languages.
This tool is useful for:
- YouTube videos
- Podcasts
- Interviews
- Online courses
- Product demonstrations
- Training materials
- Social-media videos
- Audiobooks and spoken stories
- International marketing content
Open the Dubbing Tool
- Sign in to ElevenLabs.
- Select Dubbing from the navigation menu.
- Choose to upload a file or paste a supported online video URL.
- Select one or more target languages.
- Review the advanced settings.
- Check the displayed cost.
- Click Generate.
- Wait for the dub to finish processing.
- Open the completed project and download each language version.
Automatic Dubbing is available on every ElevenLabs plan, including the Free plan. Dubs generated on the Free plan contain a watermark, while paid-plan dubs do not.
Choose Between a File and a URL
You can upload an audio or video file stored on your device. ElevenLabs currently accepts common formats including:
- MP3
- WAV
- M4A
- FLAC
- AAC
- MP4
- MOV
- AVI
- MKV
- WebM
You can also paste a supported online video URL, including content hosted on services such as YouTube, TikTok, Vimeo, and X. Only dub content that you own or have permission to reproduce, translate, and publish.
Uploading the original file is generally preferable when it provides better audio or video quality than the online version.
Check the Upload Limits
Automatic Dubbing with the current v2 model supports uploads of up to 2 GB and 180 minutes through the ElevenLabs website. The file must remain under both limits.
For a first test, use a short clip of approximately one or two minutes. This allows you to evaluate the translation, speaker detection, voice similarity, and timing without processing a complete video.
Prepare the Original Recording
Dubbing performs best when the source contains:
- Clear speech
- Limited background noise
- Minimal echo
- Distinct speakers
- Little overlapping dialogue
- Consistent recording volume
- Accurate original-language speech
- Background music that does not overpower the speakers
ElevenLabs can separate multiple speakers, including some overlapping speech, but its documentation recommends no more than approximately nine unique speakers per file for the strongest results.
Clean the recording with Voice Isolator first when loud noise makes the dialogue difficult to understand.
Remove Unnecessary Sections
Before uploading, remove:
- Long silence
- Failed takes
- Unneeded introductions
- Copyrighted material you cannot translate
- Private conversations
- Sections that will not appear in the finished video
- Repeated versions of the same scene
Dubbing cost depends partly on the duration and the number of target languages. Removing unused material can therefore reduce the total cost. ElevenLabs displays the final amount before the request is confirmed.
Select the Target Language
Open the language selector and choose the language into which the original recording should be translated.
For example:
- English to Spanish
- Spanish to English
- English to Hungarian
- English to German
- English to French
- English to Japanese
You can select several target languages in one project. Each selected language is charged separately, so begin with one language until the workflow and output quality have been tested.
Review the Source Language
ElevenLabs can identify the original language, but confirm it manually when the interface provides that option.
Automatic detection may be less reliable when:
- The recording is extremely short.
- Several languages appear.
- The opening contains only music.
- The speaker has a strong regional accent.
- The audio contains heavy background noise.
- The first words are names or technical terms.
Choosing the correct source language helps the system interpret the original speech before translating it.
Handle Multilingual Source Content Carefully
A recording may already contain more than one language. ElevenLabs supports multilingual dubbing, but review every language change carefully.
Problems are more likely when:
- A speaker switches languages in the middle of a sentence.
- Several regional accents are present.
- Foreign words appear without context.
- Speakers pronounce brand names differently.
- Dialogue overlaps.
For important multilingual content, divide the video into sections or use an editable workflow so each translation can be reviewed more carefully.
Understand Automatic Dubbing
ElevenLabs currently uses Dubbing v2 by default for new automatic dubbing projects. The process detects speakers, transcribes the original dialogue, translates it, recreates each speaker’s voice in the target language, and places the new speech over the original background audio.
Dubbing v2 is automatic. After it finishes, the translation and individual dialogue clips cannot currently be edited inside that automatic project.
Review the original file carefully before generating because correcting a small translation problem may require another workflow or a new dub.
Adjust Cloning Strength
Automatic Dubbing provides a Cloning Strength or speaker-similarity setting. The current default value of approximately 7 works well for many recordings.
A higher value attempts to keep the dubbed voice closer to the original speaker. However, it may also preserve more of the original accent or make the new language sound less natural.
A lower value gives ElevenLabs more freedom to produce natural pronunciation in the target language, but the finished voice may sound less similar to the original person.
Start with the default setting. Adjust it only after comparing the first result.
Generate the Dub
Before confirming:
- Check the uploaded file.
- Confirm the source language.
- Review every target language.
- Check the Cloning Strength.
- Review the duration.
- Review the displayed cost.
- Confirm that you have permission to translate the content.
- Click Generate.
The cost is based on the recording’s duration and the number of selected target languages. The exact total appears before the job is submitted.
Wait for Processing to Finish
The project will appear in the list of dubs while it is being processed.
Longer videos, multiple languages, several speakers, and complicated background audio may require more processing than a short recording with one clear speaker.
ElevenLabs currently allows up to five simultaneous dubbing jobs on its self-service plans.
Avoid submitting the same project again while the original job is still processing.
Review Every Speaker
Once the dub is ready, listen to each speaker carefully.
Check:
- Whether the correct person speaks each line
- Whether voices remain distinguishable
- Whether the translation matches the original meaning
- Whether names are pronounced correctly
- Whether the emotion remains appropriate
- Whether the pacing matches the scene
- Whether any dialogue begins too early or late
- Whether the background audio remains clear
- Whether the target language sounds natural
Do not review only the first speaker. A project may handle one voice correctly while producing weaker results for another.
Check the Translation
AI dubbing can provide a fast translation, but important content still requires review by someone who understands the target language.
Pay particular attention to:
- Brand names
- Personal names
- Prices and numbers
- Legal or medical terminology
- Cultural references
- Humor
- Idioms
- Slang
- Gendered language
- Formal and informal forms of address
A literal translation may be grammatically correct while sounding unnatural or changing the original meaning.
Check the Timing
Translated sentences may be longer or shorter than the original speech. ElevenLabs attempts to preserve the original timing and pacing, but some languages naturally require more words to express the same idea.
Watch the video while listening and check whether:
- Speech begins when the person starts talking.
- The sentence finishes before the scene changes.
- Long pauses remain appropriate.
- Dialogue does not overlap incorrectly.
- Important visual actions match the narration.
- The speaker does not sound unnaturally rushed.
When timing is poor, a shorter translation may work better than forcing a long sentence into the original space.
Understand Lip-Sync Limitations
ElevenLabs’ main Dubbing tool does not currently include automatic lip synchronization. The translated audio may match the general timing of the original performance without matching every visible mouth movement.
For videos where the speaker’s face is clearly visible, review the result carefully. Lip-sync options may exist through separate image-and-video workflows or third-party models, but they are not part of the standard Dubbing output.
Use Dubbing Studio When Editing Is Essential
Dubbing Studio provides more detailed control over:
- Original transcripts
- Translations
- Speaker assignments
- Individual clips
- Voice settings
- Regeneration
- Timing
- Additional voiceover tracks
- Sound-effect tracks
However, Dubbing Studio currently works through the legacy v1 dubbing model, is in maintenance mode, and receives only critical bug fixes.
To access it:
- Begin creating a new dub.
- Open Advanced Settings.
- Select Use Legacy V1 Dubbing Model.
- Enable Create a Dubbing Studio Project.
- Submit the project.
- Open the completed project from the dubbing list.
An existing automatic v2 dub cannot be converted into a Dubbing Studio project afterward. Choose the editable workflow before creating the dub.
Edit the Transcript in Dubbing Studio
Speaker cards display the original transcript and its translation.
To correct a passage:
- Select the speaker card.
- Click inside the original transcript or translated text.
- Correct the wording.
- Confirm that the correct speaker is assigned.
- Review the clip-level voice settings.
- Regenerate the affected section.
- Listen to it in context.
Both the original transcript and translated version can be edited directly.
Correct transcription errors before correcting the translation. A mistranslated sentence may begin with an incorrect interpretation of the source dialogue.
Reassign a Speaker
If a line uses the wrong voice:
- Select the affected dialogue clip.
- Open its speaker assignment.
- Choose the correct speaker.
- Confirm the voice settings.
- Regenerate the clip.
- Review the transition between speakers.
Speaker reassignment is one of the main reasons to use Dubbing Studio rather than the automatic workflow.
Regenerate Only the Weak Clip
When one sentence sounds unnatural, regenerate only that clip instead of recreating the complete project.
Before regenerating:
- Shorten an overly long translation.
- Correct punctuation.
- Fix names and numbers.
- Confirm the speaker.
- Review the voice settings.
- Check the clip timing.
Listen to the sentence before and after the replacement to make sure the new version fits naturally.
Add a Manual Dub
Dubbing Studio also provides a Manual Dub option for users who already have an accurate transcript and translation. This helps preserve specific clip divisions and speaker assignments rather than relying entirely on automatic detection.
A manual workflow is useful when:
- A professional translator prepared the script.
- Speaker assignments are already known.
- Precise wording is required.
- Technical terminology must remain exact.
- The automatic transcript repeatedly fails.
- The production contains complicated dialogue.
Add Voiceover or Sound-Effect Tracks
Dubbing Studio can add new voiceover tracks and sound-effect tracks to the timeline.
A voiceover track can be used to insert a new translated introduction, explanation, or narrator. Sound-effect tracks can generate new effects from written prompts and place them at the required point in the timeline.
These additions should support the original production rather than unnecessarily replacing its existing background audio.
Download the Automatic Dub
When the automatic dub is ready:
- Open the completed dubbing project.
- Select the target language.
- Click the download option.
- Choose an available format.
- Save the file with a clear name.
- Watch or listen to the downloaded version completely.
Automatic dubbing outputs can include MP4 video or audio files, depending on the uploaded source and selected download.
Use filenames such as:
Vaxali-Tutorial-Spanish-Dub-Final.mp4
Podcast-Episode-One-Hungarian-Dub.aac
Product-Demo-German-Dub-V2.mp4
Export From Dubbing Studio
Dubbing Studio currently supports exports including:
- AAC audio
- MP3 audio
- WAV audio
- Separate audio tracks in a ZIP file
- Separate audio clips in a ZIP file
- AAF timeline data
- SRT subtitles
- CSV transcript and translation data
Select the correct target language before exporting.
Separate speaker tracks are useful when the dub will receive additional mixing inside professional video or audio-editing software.
Export Subtitles
Use SRT when the translated dialogue will also appear as subtitles or captions.
Before exporting, review:
- Spelling
- Speaker names
- Punctuation
- Translation accuracy
- Timing
- Line length
- Reading speed
- Language selection
The subtitle file should match the final dubbed version rather than an earlier translation draft.
Keep the Original Background Audio
One benefit of ElevenLabs Dubbing is that it attempts to preserve music, environmental sounds, and effects from the source while replacing the dialogue.
After processing, check whether:
- Music remains at the correct volume.
- Important sound effects are still present.
- Background voices were handled correctly.
- The translated speech remains understandable.
- No soundtrack section disappeared.
- Dialogue and music remain balanced.
A complex soundtrack may still require additional editing and mixing.
Troubleshoot a Weak Dub
When the voice sounds too different:
- Increase Cloning Strength slightly.
- Use a clearer source recording.
- Confirm that speakers were detected correctly.
- Use Dubbing Studio to reassign the voice when available.
When the target language sounds unnatural:
- Lower Cloning Strength slightly.
- Shorten the translation.
- Replace literal wording with natural phrasing.
- Have a fluent speaker review it.
- Regenerate the affected passage.
When the timing is poor:
- Shorten long translated sentences.
- Remove unnecessary filler words.
- Adjust the clip boundaries in Dubbing Studio.
- Divide one long sentence into two clips.
When background audio sounds damaged:
- Use a higher-quality source file.
- Avoid heavily compressed online video.
- Upload the original audio or video.
- Export separate tracks for additional mixing.
Protect the Speakers’ Rights
Before dubbing a recording, confirm that you have permission to use and translate every voice and piece of content included.
Do not use dubbing to:
- Misrepresent what someone originally said
- Create a false endorsement
- Change the meaning deceptively
- Impersonate a speaker
- Translate private recordings without permission
- Publish copyrighted content without authorization
When viewers could reasonably believe that the dubbed speaker recorded the translated language personally, disclose that AI dubbing was used.
Save the Complete Project Archive
Store:
- The original video or audio
- The original transcript
- The corrected translation
- The automatic dub
- Edited language versions
- Subtitle files
- Separate speaker tracks
- The selected Cloning Strength
- Notes about pronunciation
- Proof of the content and voice permissions
Keep a separate folder for each target language.
For example:
- Original English
- Spanish dub
- Hungarian dub
- German dub
- Subtitle files
- Final exports
Complete a Dubbing Check
Before publishing the translated content, confirm that:
- You have permission to dub the original recording.
- The correct source language was selected.
- The correct target language was generated.
- Every speaker has the appropriate voice.
- Names and terminology are pronounced correctly.
- The translation preserves the original meaning.
- Cultural references have been adapted appropriately.
- Dialogue remains synchronized with the scene.
- Background music and effects remain audible.
- The watermark and plan restrictions are understood.
- The correct language version was downloaded.
- Subtitles match the finished dub.
- A fluent speaker has reviewed important content.
- The original and translated files are stored securely.
Dubbing can help creators reach audiences in other languages without rerecording every speaker. However, translation accuracy, cultural context, and timing still require careful human review before the finished version is published.
Download, Organize, and Reuse Your Audio
After generating a voiceover, sound effect, music track, or cleaned recording, download and organize the approved file before continuing to the next part of the project. Clear filenames and folders prevent finished audio from being confused with test generations.
ElevenLabs also provides History and Assets tools for reopening previous generations, but important files should still be saved locally.
Download Audio Immediately
After generating speech:
- Listen to the complete recording.
- Confirm that the wording is correct.
- Check the pronunciation and pacing.
- Click the download button.
- Select the preferred format.
- Save the file in the correct project folder.
- Rename it clearly.
Text to Speech generations can be downloaded immediately after creation or reopened later through the History panel.
Do not wait until the entire project is finished to download approved sections. A voice, project setting, or earlier generation may become harder to locate later.
Choose the Right File Format
ElevenLabs currently provides standard downloads in MP3 and WAV, with additional options such as higher-bitrate MP3, M4A, and FLAC available through the advanced download menu.
Use MP3 when:
- A smaller file is preferred
- The audio will be uploaded to a website
- The recording is intended for social media
- Only basic editing is required
- Storage space is limited
Use WAV when:
- The audio will be edited extensively
- Maximum quality is important
- The project includes professional mixing
- Several audio clips must be combined
- The final file will be compressed later
Use M4A when the application or device works better with that format.
Use FLAC when lossless quality is needed but a smaller file than an uncompressed WAV is preferred.
For most beginners, MP3 is sufficient for a finished online voiceover. WAV is the safer choice when the recording will still be edited.
Select an Appropriate MP3 Quality
The normal Text to Speech download provides an MP3 at 128 kilobits per second. The advanced menu currently offers higher-quality MP3 options at 192 and 256 kilobits per second.
Use:
- 128 kbps for drafts and simple spoken content
- 192 kbps for polished voiceovers and podcasts
- 256 kbps when higher audio quality is worth the larger file
Spoken narration does not normally require the same bitrate as complex music, but avoid repeatedly compressing the same MP3 during editing. Repeated compression can gradually reduce clarity.
Use History to Find Earlier Generations
To reopen a previous Text to Speech result:
- Open Text to Speech.
- Select the History tab.
- Find the correct generation.
- Play it to confirm the content.
- Click the download icon.
- Choose the desired format.
Voice Changer generations can be reopened through their own History panel in a similar way.
History is useful when:
- An earlier generation sounded better
- A downloaded file was misplaced
- You need another format
- You want to confirm which voice was used
- You want to avoid spending credits on the same script again
Compare Versions Before Deleting Anything
AI-generated speech can vary between attempts, even when the same script, voice, model, and settings are used.
Keep the strongest versions until the project is complete.
For example:
Introduction-V1.mp3Introduction-V2.mp3Introduction-V3-Approved.wav
Listen to each version in context before deciding which one should be final. A sentence that sounds strong by itself may not transition naturally from the previous paragraph.
Use Studio Generation History
Inside ElevenCreative Studio, each paragraph can have its own Generation History. This allows users to listen to, restore, download, or remove earlier versions of that paragraph.
To restore an earlier version:
- Select the paragraph.
- Open its Generation History.
- Listen to the available generations.
- Find the preferred version.
- Select Restore previous generation.
- Play the paragraph with the surrounding audio.
- Lock it when it is approved.
Each individual generation can also be downloaded through its three-dot menu.
Removing a Studio generation is permanent, so do not delete earlier versions until the replacement has been reviewed fully.
Lock Approved Studio Paragraphs
After restoring or approving a generation, lock the paragraph to protect it from accidental changes.
Lock a section when:
- The script is final
- Pronunciation is correct
- The voice and model are correct
- The pacing matches nearby sections
- No additional regeneration is expected
Locking does not replace a local backup. Download important narration even when it remains saved inside Studio.
Open the Assets Area
The Assets section provides a central library for creative files used across ElevenCreative. It can store uploaded media and outputs from tools such as Text to Speech, Voice Changer, Voice Isolator, Sound Effects, and supported image or video workflows.
To open it:
- Select Assets from the main navigation.
- Review the available audio, video, voice, and media assets.
- Use search or filters to locate a specific item.
- Preview the asset before reusing it.
- Download, rename, move, or delete it as needed.
Assets is useful for material that may be reused across several projects.
Add Generated Content to Assets
Depending on the tool, content can be added by:
- Uploading a file directly
- Dragging a generated output into Assets
- Selecting Save to Assets
- Using the three-dot action menu
ElevenLabs currently supports common audio formats such as MP3, WAV, FLAC, and AAC in the Assets system.
Save reusable items such as:
- A standard channel introduction
- A brand narrator sample
- A podcast transition
- A notification sound
- Background ambience
- Approved music
- Frequently used voiceovers
- Cleaned recordings
Do not save every rejected test. An overcrowded asset library can become as difficult to manage as an unorganized download folder.
Rename Assets Clearly
Replace automatic filenames with names that describe the content.
Weak filename:
elevenlabs_2026_07_28_184739.mp3
Better filename:
Vaxali-ElevenLabs-Guide-Introduction-Warm-Narrator-V2.mp3
A useful filename can include:
- Project name
- Section or scene
- Voice or speaker
- Language
- Version number
- Approval status
- File format
For example:
Product-Demo-Scene-03-Spanish-Final.wav
Keep filenames readable. Adding every setting to the filename can make it unnecessarily long.
Use Consistent Version Numbers
Choose one versioning system and use it throughout the project.
For example:
V1V2V3FinalFinal-Approved
Avoid filenames such as:
FinalFinal-NewFinal-RealFinal-Updated-AgainFinal-Use-This-One
A clearer sequence would be:
Introduction-V1.wavIntroduction-V2.wavIntroduction-V3-Approved.wav
When an approved file is changed later, increase its version number instead of overwriting the earlier file.
Create Project Folders
A simple folder structure might be:
ElevenLabs Beginner Guide
- Scripts
- Voice tests
- Approved narration
- Dialogue
- Sound effects
- Music
- Transcripts
- Dubbing
- Studio exports
- Final project
For a multilingual project, create a folder for each language:
- English
- Spanish
- Hungarian
- German
Inside each language folder, keep the script, narration, subtitles, and final export together.
ElevenLabs Assets also supports folders, nested folder structures, search, and moving items between folders.
Organize Assets by Project or Purpose
There are two useful organization methods.
By project:
- YouTube Tutorial
- Podcast Episode
- Online Course
- Product Demonstration
By asset type:
- Narration
- Music
- Sound Effects
- Voice Samples
- Background Ambience
Project-based folders work well when files are used only once. Asset-type folders work better for reusable brand material.
A combined structure may be best:
- Brand Assets
- Reusable Audio
- Active Projects
- Archived Projects
Save the Script With the Audio
Keep the final written script beside the approved recording.
The script should include:
- Exact spoken wording
- Speaker names
- Audio tags
- Pronunciation substitutions
- Important pauses
- Voice name or ID
- Speech model
- Voice settings
- Date generated
This information helps when a line must be recreated or updated later.
For example:
Voice: Warm NarratorModel: Multilingual v2Stability: 50Similarity: 75Style: 0Speed: 1.0
The same settings may not reproduce an identical performance, but they improve consistency.
Save Voice IDs and Model Names
Voice Library names may be edited, duplicated, or removed. Record the voice ID when it is available, especially for a long-term project.
Keep:
- Voice name
- Voice ID
- Voice type
- Speech model
- Language
- Accent
- Credit multiplier
- Notice period
- Voice settings
This makes it easier to confirm which narrator was used and identify a replacement when necessary.
Keep Source and Final Files Separate
Do not place raw recordings and finished audio in the same folder without clear labels.
Use separate folders for:
- Source recordings
- Cleaned recordings
- Generated speech
- Edited versions
- Final exports
For Voice Changer, keep both:
- The original performance
- The transformed voice
For Voice Isolator, keep both:
- The noisy source
- The isolated voice
The original may be needed when the processed version removes an important sound or introduces an artifact.
Reuse an Approved Generation
Before generating a new line, check whether the required audio already exists.
A reusable recording might include:
- “Welcome back.”
- A standard legal disclaimer
- A podcast introduction
- A company name
- A call to action
- A recurring character reaction
- A transition between lessons
Reusing an approved file saves credits and maintains consistency. However, make sure its tone and pacing still fit the new context.
A cheerful introduction may not fit a serious announcement even when the words are identical.
Reuse Assets Inside Studio
Saved assets can be added to compatible ElevenCreative projects.
For example, you can reuse:
- An intro sound
- Background music
- A narrator clip
- A cleaned interview
- A transition effect
- A logo animation
- A standard outro
Preview the file before adding it. Confirm that it is the correct version, language, and length.
Share Assets Within a Workspace
Users working in a multi-seat ElevenLabs workspace can share asset folders with other workspace members. Folder roles currently include Viewer, Editor, and Admin. External folder sharing is not supported through the Assets system.
Use:
- Viewer when someone should use assets but not modify them
- Editor when someone should organize and update the contents
- Admin only when the person should also manage access and deletion
Be careful when granting deletion privileges because removed assets cannot be recovered.
Download Files Before Sharing Externally
To send an asset to someone outside the ElevenLabs workspace:
- Download the approved file.
- Confirm that it opens correctly.
- Rename it clearly.
- Share it through an authorized file-transfer method.
- Include the related script or instructions when necessary.
Do not send an unreviewed generation directly from History simply because it is the newest version.
Do Not Rely Only on ElevenLabs Storage
History and Assets are convenient, but they should not be the only location for important work.
Store local or cloud backups of:
- Final narration
- Original recordings
- Scripts
- Voice settings
- Music
- Sound effects
- Subtitles
- Transcripts
- Dubbing exports
- Complete Studio projects and exports
This protects the project when:
- A voice becomes unavailable
- A file is deleted accidentally
- A workspace changes
- An account is closed
- A collaborator removes an asset
- An earlier generation can no longer be located
Be Careful When Deleting Assets
Deleting an item from Assets is permanent and cannot currently be undone. ElevenLabs asks for confirmation, but there is no recovery option afterward.
Before deleting:
- Preview the file.
- Check whether it is used in an active project.
- Confirm that a local backup exists.
- Verify that it is not shared with another workspace member.
- Make sure it is not the only approved version.
Archive completed projects instead of immediately deleting their assets.
Create a Final Project Archive
After completing a project, save one archive containing:
- Final script
- Approved narration
- Original recordings
- Voice names and IDs
- Model and setting notes
- Music and sound effects
- Captions and subtitles
- Transcripts
- Dubbing versions
- Final audio or video
- Licensing and permission records
Use a folder name containing the project and completion date:
Vaxali-ElevenLabs-Guide-Completed-2026-07-28
A complete archive makes future corrections, translations, and updates easier.
Complete a File-Management Check
Before considering the audio complete, confirm that:
- The strongest generation was downloaded.
- The correct format was selected.
- The file has a descriptive name.
- Earlier versions remain available until approval.
- The script is stored beside the audio.
- Voice and model details have been recorded.
- Source and processed files are separated.
- Reusable content has been added to Assets.
- Important files have local or cloud backups.
- Shared folders have appropriate permissions.
- No active project depends on an asset being deleted.
- The final archive contains everything required to update the project later.
Good file organization may seem unnecessary during a short test, but it becomes essential once a project contains several voices, languages, revisions, and exports. The next step is learning how to manage credits, subscription limits, and usage more efficiently.
Manage Credits, Subscription Limits, and Usage
ElevenLabs uses credits to measure activity across its creative tools. The number of credits consumed depends on the product, speech model, selected voice, generation length, account plan, and whether the content is created through the website or API. Credits were previously described as characters, but ElevenLabs now uses one shared credit system across the platform.
Understanding credits before starting a long project helps prevent unfinished narration, unexpected purchases, or wasted generations.
Check the Remaining Credits
To view the current balance:
- Sign in to ElevenLabs.
- Click the profile icon.
- Select Subscription.
- Review the active plan.
- Check the included, used, and remaining credits.
- Confirm the next renewal date.
- Review any available Pay As You Go balance.
The Subscription page also displays plan features, custom voice limits, audio-quality options, billing information, and available upgrades.
Check the balance before starting:
- A long Studio project
- An audiobook
- Several language dubs
- Music generation
- Voice Changer processing
- Voice isolation
- A large transcription
- Several voice or model comparisons
Understand How Credits Are Used
Different ElevenLabs tools charge credits differently.
Text to Speech is generally based on the amount of text generated and the selected model. Other tools may charge according to audio duration, the number of target languages, or generation settings.
For example:
- Text to Speech usually depends on the written character count and model.
- Voice Changer depends on the duration of the source recording.
- Voice Isolator depends on the processed audio duration.
- Dubbing depends partly on duration and the number of target languages.
- Sound Effects depends on the selected duration settings.
- Music depends on the generated track and selected options.
- Speech to Text depends on the duration of the uploaded recording.
The precise cost should appear in the interface before the generation is confirmed. ElevenLabs notes that credit usage varies by product, model, plan, and whether the website or API is being used.
Review the Current Self-Service Plans
ElevenLabs currently lists the following monthly self-service plans:
| Plan | Monthly price | Included monthly credits |
|---|---|---|
| Free | $0 | 10,000 |
| Starter | $6 | 30,000 |
| Creator | $22 | 121,000 |
| Pro | $99 | 600,000 |
| Scale | $299 | 1,800,000 |
| Business | $990 | 6,000,000 |
Enterprise pricing is customized for larger organizations. ElevenLabs may also offer temporary introductory discounts, and annual billing can change the effective monthly cost. Review the Subscription page before upgrading because pricing and allowances can change.
Use the Free Plan for Learning
The Free plan is enough for testing short scripts, experimenting with voices, learning Text to Speech, and exploring several creative tools.
It currently includes 10,000 monthly credits and three custom voice slots that can be used with Voice Design. Voice cloning requires the Starter plan or higher. Content generated through the Free plan is intended for noncommercial use with attribution, while eligible paid plans provide commercial usage rights under ElevenLabs’ applicable terms.
The Free plan is best for:
- Learning the interface
- Testing several narrators
- Creating short personal samples
- Practicing script formatting
- Comparing speech models
- Trying Voice Design
- Deciding whether a paid plan is necessary
Do not begin a large audiobook, course, or multilingual project on the Free plan without estimating the expected usage first.
Choose Starter for Basic Paid Use
Starter currently includes 30,000 monthly credits and unlocks voice cloning. It can suit beginners who need commercial rights, occasional short voiceovers, or an authorized Instant Voice Clone but do not yet produce large volumes of audio.
Starter may be appropriate for:
- Short YouTube videos
- Occasional social-media narration
- Small business demonstrations
- Basic podcast introductions
- Testing an Instant Voice Clone
- Limited monthly commercial content
A creator producing several long videos each week may reach the Starter allowance quickly.
Choose Creator for Regular Production
Creator currently includes 121,000 monthly credits and one Professional Voice Clone slot. It is designed for more regular production and provides substantially more capacity than Starter.
Creator may suit:
- Weekly YouTube videos
- Online courses
- Recurring podcast narration
- Longer commercial voiceovers
- A personal Professional Voice Clone
- Frequent Studio projects
- Regular multilingual content
Estimate the expected script length and other tool usage before upgrading solely because one feature requires the Creator tier.
Consider Pro, Scale, or Business for High Usage
Pro, Scale, and Business provide larger monthly credit pools for creators, production teams, agencies, applications, and organizations processing significant amounts of content.
Current monthly allowances are:
- Pro: 600,000 credits
- Scale: 1,800,000 credits
- Business: 6,000,000 credits
Scale currently includes three seats, while Business includes ten seats.
These plans are more suitable when several people share a workspace, multiple projects are produced simultaneously, or the monthly volume consistently exceeds Creator.
Do not choose a larger plan only because one unusually large project requires additional credits. Compare the cost of temporarily upgrading, purchasing Pay As You Go credits, or dividing the project across billing cycles.
Estimate Text to Speech Usage
Before generating a long script:
- Check the script’s character count.
- Confirm the selected speech model.
- Review the displayed generation cost.
- Add an allowance for corrections.
- Add another allowance for voice tests.
- Reserve credits for other tools being used in the project.
Do not budget only for the final script. A realistic project often includes:
- Voice comparisons
- Pronunciation tests
- Regenerated paragraphs
- Alternate introductions
- Corrected numbers
- Replaced sentences
- Different emotional versions
- Final exports
A useful beginner approach is to reserve approximately 20% to 30% of the expected project allowance for corrections and additional generations. This is a planning recommendation rather than an ElevenLabs billing rule.
Test With Short Scripts First
Before generating a 10-minute narration, test one representative paragraph.
The test should include:
- The selected voice
- The intended model
- Important names
- Numbers
- Technical terminology
- The expected emotional tone
- Typical sentence length
A short test can reveal the wrong narrator, accent, model, or settings before a large amount of credits is used.
Review the Cost Before Clicking Generate
ElevenLabs displays the expected cost in supported workflows before the request is confirmed.
Before generating, check:
- Script length
- Audio duration
- Number of target languages
- Selected voice
- Voice credit multiplier
- Model
- Number of variations
- Music duration
- Sound-effect duration
- Optional transcription features
Do not click Generate repeatedly while a request is processing. A slow generation may eventually complete, and submitting it again can create another charge.
Avoid Regenerating Approved Content
When one sentence is incorrect, regenerate only that sentence or paragraph rather than the complete chapter.
Inside Studio:
- Select the affected paragraph.
- Correct the script.
- Regenerate that section.
- Compare it with the surrounding narration.
- Lock it after approval.
For ordinary Text to Speech, divide the script into logical sections so one mistake does not require several minutes of speech to be recreated.
Use History Before Generating Again
Check Text to Speech History, Voice Changer History, Studio Generation History, and Assets before recreating missing content.
An earlier approved version may already be available for download. Reusing it prevents unnecessary credit usage and may preserve a performance that cannot be reproduced exactly.
Check Community Voice Multipliers
Some Voice Library voices may use custom credit multipliers. A voice with a 2x multiplier consumes twice the normal credits for the same generation.
Before using a community narrator for a long project:
- Open the voice details.
- Check for a credit multiplier.
- Generate a short test.
- Review the displayed cost.
- Compare it with a standard-rate alternative.
A voice may sound excellent but be impractical for a long audiobook or recurring series when it consumes significantly more credits.
Understand Credit Renewal
Paid self-service plans receive a new credit allocation each month, including when the subscription is billed annually. Annual plans are paid yearly, but their self-service credits still reset monthly.
The renewal date is based on the billing cycle rather than necessarily occurring on the first day of each calendar month.
For example, a subscription beginning on July 15 may renew around the fifteenth rather than August 1.
Check the exact renewal date on the Subscription page instead of assuming the balance resets at the start of every month.
Understand Credit Rollover
On standard paid subscriptions, unused credits can roll into the next billing cycle as long as the account remains on the same eligible plan. ElevenLabs allows up to two months’ worth of unused credits to roll over, meaning the maximum available balance can reach approximately three times the plan’s normal monthly quota after the new allocation is added. Free-plan credits do not roll over.
For example, a plan with a normal allowance of 100,000 credits could potentially hold:
- 100,000 current-month credits
- Up to 200,000 rolled-over credits
- 300,000 total available credits
This example illustrates the rollover rule and is not a current ElevenLabs plan allowance.
Stay on the Same Plan to Preserve Rollover
Rollover depends on remaining on the same eligible subscription.
When a paid subscription continues normally, unused credits within the rollover limit carry forward. When the subscription is canceled or downgraded, unused subscription credits are lost when that change takes effect at the end of the current billing cycle.
Before canceling or downgrading:
- Check the remaining balance.
- Complete important generations.
- Download approved files.
- Export unfinished Studio projects when possible.
- Save voice and project settings.
- Confirm when the change will take effect.
Do not generate unnecessary content solely to use expiring credits. Create only material that is genuinely useful.
Understand What Happens During an Upgrade
Upgrading during an active billing cycle starts a new billing cycle. ElevenLabs charges for the new plan, adds its new allowance, and carries the remaining unused subscription credits from the previous plan into the new quota.
Before upgrading:
- Check the remaining current credits.
- Review the new monthly allowance.
- Confirm the new renewal date.
- Review the immediate charge.
- Check whether annual or monthly billing is selected.
- Confirm which features the upgrade unlocks.
An upgrade is useful when the monthly usage will continue at the higher level, not only when one poorly planned generation used more credits than expected.
Understand What Happens During a Downgrade
A downgrade normally takes effect at the end of the current billing cycle. The account remains on its existing plan until that date, but unused subscription credits are lost when the downgrade becomes effective.
A lower plan may also reduce:
- Custom voice slots
- Professional Voice Clone access
- Workspace seats
- Audio-quality options
- Monthly credits
- Other plan-specific features
ElevenLabs states that account content is not automatically deleted when a paid subscription ends, but access to paid features such as cloning may be restricted until the user subscribes again.
Understand What Happens After Cancellation
Canceling schedules the subscription to end after the current billing cycle. The account then returns to the Free plan. ElevenLabs does not currently offer a pause option for paid subscriptions.
To cancel or downgrade:
- Open the profile menu.
- Select Subscription.
- Click Manage subscription.
- Choose Cancel subscription or Downgrade.
- Complete the confirmation process.
- Check that the scheduled change appears.
The subscription can generally be resumed before the current billing period ends.
Understand Pay As You Go
For new self-service subscriptions, ElevenLabs now offers Pay As You Go, or PAYG, as the replacement for the older usage-based billing system.
PAYG allows users to purchase additional credits in advance. Those credits are used after the included monthly allowance has been exhausted and remain valid for 12 months after purchase. PAYG can be used with or without a paid subscription, subject to account eligibility.
PAYG may be useful when:
- One month has unusually high usage.
- A large project exceeds the normal allowance.
- The user does not want to move permanently to a higher subscription.
- Additional credits are needed before the next renewal.
- An account needs prepaid rather than postpaid spending.
Understand How PAYG Differs From Subscription Credits
Subscription credits:
- Are included with the plan.
- Reset according to the billing cycle.
- May roll over within the permitted limit.
- Can be lost after cancellation or downgrade.
PAYG credits:
- Are purchased separately.
- Remain valid for 12 months.
- Are used after the monthly subscription allowance.
- Are stored according to a monetary balance and converted using the current plan’s credit rate.
When the subscription plan changes, the monetary value of the remaining PAYG balance stays the same, but the number of credits represented by that balance may change because each plan can have a different credit rate.
Do Not Confuse PAYG With Legacy Usage-Based Billing
Older ElevenLabs accounts may still show usage-based billing. This was a postpaid feature that allowed eligible users to continue generating after exhausting their monthly quota and receive the additional charge later.
ElevenLabs now classifies usage-based billing as a legacy feature. It is not available for new self-service subscriptions. New self-service customers use prepaid PAYG credits instead, while Enterprise and eligible legacy accounts may still have the older overage system.
Follow the controls shown in the actual Subscription page because older tutorials may describe options that are no longer available to new accounts.
Set a Spending Budget
Before purchasing additional credits, decide the maximum amount that can be spent on the project.
A simple budget might include:
- Primary narration
- Correction allowance
- Music
- Sound effects
- Dubbing
- Transcription
- Emergency reserve
For example, when narration is the main requirement, avoid spending most of the available balance on several experimental music tracks before the speech is complete.
Track Usage by Project
When several projects share one account, record the credits used by each one.
A simple usage log can include:
| Date | Project | Tool | Credits before | Credits after | Credits used |
|---|---|---|---|---|---|
| July 28 | Tutorial intro | Text to Speech | 20,000 | 18,750 | 1,250 |
| July 28 | Rain ambience | Sound Effects | 18,750 | 18,550 | 200 |
| July 28 | Spanish version | Dubbing | 18,550 | 12,000 | 6,550 |
The example numbers are illustrative rather than fixed ElevenLabs rates.
Tracking usage helps identify which tools consume the most credits and whether the current subscription remains appropriate.
Separate Testing From Production
Create two stages for every project.
Testing stage:
- Short scripts
- One or two voices
- Default settings
- Brief music previews
- One target language
- Low-cost experiments
Production stage:
- Approved voice
- Approved model
- Final script
- Final settings
- Correct pronunciation
- Confirmed export format
Do not begin large production generations while still changing the narrator, script, or model.
Generate Difficult Sections Before Easy Ones
Test the most challenging paragraph first.
A difficult section may contain:
- Names
- Foreign words
- Technical terms
- Several numbers
- Emotional delivery
- Unusual pacing
- Dialogue
- Acronyms
When the selected voice cannot handle the difficult section, choosing another voice early prevents regenerating all the easier content later.
Remove Unnecessary Silence From Duration-Based Tools
Voice Changer, Voice Isolator, transcription, and dubbing use the duration of the uploaded recording as an important part of the cost.
Before uploading:
- Remove long silence.
- Delete failed takes.
- Trim unused introductions.
- Remove repeated sections.
- Export only the material that needs processing.
Do not upload a 30-minute file when only five minutes require isolation or transformation.
Begin Dubbing With One Language
Selecting several target languages multiplies the scope and cost of a dubbing project.
For a first test:
- Choose one target language.
- Use a short representative clip.
- Review the translation.
- Check speaker similarity.
- Confirm the timing.
- Correct the production workflow.
- Add other languages afterward.
Generating five complete language versions before reviewing the first one may repeat the same translation or speaker-detection problem across every output.
Keep Sound Effects Short
A notification, impact, click, or transition normally needs only a few seconds. Selecting a long fixed duration for a one-shot effect wastes credits and creates unnecessary audio.
Use Auto duration or the shortest practical length, then extend or loop ambience inside Studio when appropriate.
Generate Short Music Tests
Begin with approximately 15 to 30 seconds of music before creating a full track.
Confirm:
- Genre
- Mood
- Instruments
- Vocal style
- Production quality
- Whether it fits beneath narration
Extend the approved musical direction rather than generating several complete five-minute songs that do not match the project.
Monitor Workspace Usage
When several members share an ElevenLabs workspace, their activity may draw from the same account credit pool.
Workspace owners should establish:
- Project budgets
- Approved models
- Voice-selection rules
- Generation limits
- File-naming standards
- Responsibility for reviewing costs
- A process for purchasing additional credits
A shared workspace can exhaust its allowance quickly when several members test long scripts without coordinating.
Understand Immediate Credit-Use Limits
ElevenLabs places limits on how many credits can be used at once to reduce abuse. Current limits are set at twice the normal monthly quota for the applicable subscription and include:
- Starter: 60,000 credits
- Creator: 200,000 credits
- Pro: 1,000,000 credits
- Scale: 4,000,000 credits
- Business: 22,000,000 credits
These limits do not necessarily represent the balance available in the account. They restrict how much can be consumed in one period or operation under the platform’s safeguards.
Review Annual Billing Carefully
ElevenLabs offers monthly and annual subscription options but does not currently offer quarterly plans. Annual plans renew automatically unless canceled. Self-service annual customers still receive credits on a monthly reset schedule rather than receiving the entire year’s allowance immediately.
Before selecting annual billing:
- Confirm the full amount charged today.
- Review the effective monthly savings.
- Make sure the service will be used throughout the year.
- Confirm the included monthly credits.
- Understand the cancellation date.
- Check whether the plan can be upgraded later.
An annual plan may reduce the effective monthly cost, but it also creates a longer financial commitment.
Review Invoices and Billing Information
To access invoices:
- Open Subscription.
- Select Manage subscription.
- Open Manage billing information.
- Locate Invoice History.
- Select an invoice.
- Download the invoice or receipt.
ElevenLabs also emails invoices and receipts to the account’s registered email address.
Keep invoices for:
- Business bookkeeping
- Client billing
- Tax records
- Subscription reviews
- Confirming commercial-plan usage
- Tracking project expenses
Avoid Accidental Charges
Before confirming a purchase or upgrade:
- Check whether monthly or annual billing is selected.
- Review the full amount due.
- Confirm the correct plan.
- Check for introductory pricing that changes later.
- Review automatic renewal.
- Verify the payment method.
- Save the receipt.
Do not assume that a displayed monthly equivalent means the service will be billed monthly. Annual plans may display a monthly average while charging the full year upfront.
Review the Refund Conditions Before Purchasing
ElevenLabs’ current billing documentation states that a refund may be available when the request is submitted within 14 days of payment and no credits from the relevant purchased quota have been used. Refund requests must be submitted through ElevenLabs support from the email address associated with the account.
Because using credits can affect refund eligibility, review the plan and payment carefully before beginning production.
Know When to Upgrade
Consider upgrading when:
- The account repeatedly runs out of credits.
- PAYG purchases cost more than the next plan would.
- A required feature is unavailable.
- Professional Voice Cloning is needed.
- Several projects are produced every month.
- A team needs additional seats.
- Higher audio quality is necessary.
- Commercial work has grown beyond the current tier.
Do not upgrade only because one test used more credits than expected. First determine whether the problem was the plan size or inefficient generation.
Know When to Downgrade
Consider downgrading when:
- Most monthly credits remain unused.
- Production has slowed permanently.
- Advanced cloning features are no longer needed.
- The account no longer supports several active projects.
- The next plan down still provides sufficient capacity.
Before downgrading, finish important generations and understand which features or custom voices may become inaccessible.
Complete a Credit-Management Check
Before starting a major project, confirm that:
- The active subscription is correct.
- The remaining credits are visible.
- The next renewal date is known.
- The project’s estimated usage has been calculated.
- Additional credits are reserved for corrections.
- The selected voice does not have an unexpected multiplier.
- Short tests were completed before full generation.
- Duration-based uploads were trimmed.
- Existing History and Assets were checked.
- Rollover rules are understood.
- Cancellation or downgrade consequences are understood.
- PAYG purchases fit the project budget.
- Annual and monthly billing have not been confused.
- Workspace members understand the usage limits.
- Important files are downloaded before subscription changes.
Managing credits does not require calculating every generation perfectly. The main goal is to test on a small scale, approve the workflow, and use the remaining balance for the final production rather than repeated experiments.
Tips for Getting Better Results With ElevenLabs
ElevenLabs can create realistic audio quickly, but the first generation will not always be the strongest one. Voice selection, script quality, model choice, pronunciation, recording quality, and generation length can all affect the result.
The most reliable workflow is to test a short section, identify the specific problem, and change only one element at a time.
Start With the Right Voice
Choose a voice that already matches the project’s intended language, accent, emotion, and delivery style.
A calm narrator will usually perform instructional content more naturally than an exaggerated character voice. Similarly, a voice recorded with a Spanish accent may produce more natural Spanish narration than a voice originally created for American English.
ElevenLabs identifies voice selection as one of the most important factors affecting Text to Speech quality, particularly with Eleven v3. Emotional instructions also work better when they match the voice’s natural character.
Before creating the complete project, test the voice with:
- A normal sentence
- A question
- Important names
- Numbers and dates
- Technical terms
- A longer paragraph
- The strongest emotional section
A voice that sounds impressive in its prepared preview may not suit the real script.
Match the Model to the Project
Use the model that best supports the required workflow rather than automatically choosing the newest option.
Eleven Multilingual v2 is generally better suited to polished, stable prerecorded narration. Flash v2.5 prioritizes faster generation, while Eleven v3 provides more expressive control through audio tags, punctuation, and dialogue features.
For example:
- Use Multilingual v2 for tutorials, courses, audiobooks, and professional narration.
- Use Eleven v3 for dramatic speech, character performances, and multi-speaker conversations.
- Use Flash v2.5 for fast previews, real-time applications, and lower-latency workflows.
Test two likely models with the same short passage before committing to a long project.
Write the Script for Speech
Text copied from an article may contain long sentences, complicated wording, headings, links, citations, and formatting that do not sound natural when spoken.
Rewrite the content using:
- Shorter sentences
- Clear transitions
- Familiar wording
- Natural punctuation
- Paragraph breaks
- Spoken versions of numbers and symbols
For example:
Written version:
“Users are advised to review the platform’s settings prior to initiating the audio-generation process.”
Spoken version:
“Review the settings before generating the audio.”
Natural sentence structure and punctuation help ElevenLabs interpret pacing, emphasis, and emotion more effectively.
Provide Enough Context
Avoid generating only one or two isolated words unless that is absolutely necessary. The model may not have enough information to determine the correct emotion, language, emphasis, or pronunciation.
Instead of generating:
“Perfect.”
Use:
“Perfect. Your first voiceover is now ready to download.”
The longer version explains the intended meaning and makes the delivery more predictable.
ElevenLabs also notes that very short Eleven v3 prompts are more likely to produce inconsistent output.
Divide Long Scripts Into Sections
Do not generate an entire article, course, or audiobook chapter as one continuous block simply because it fits within the available limit.
Divide it into:
- Introduction
- Main sections
- Examples
- Important warnings
- Conclusion
Shorter sections are easier to review, replace, and organize. They also reduce the chance that one pronunciation mistake or unexpected voice change will require the complete narration to be regenerated.
ElevenLabs recommends using Studio for longer content because individual paragraphs can be regenerated without recreating the entire project.
Test the Most Difficult Passage First
Do not begin with the easiest sentence. Test the part of the script that contains the greatest risk of problems.
This might include:
- Brand names
- People’s names
- Foreign words
- Prices
- Phone numbers
- Acronyms
- Emotional dialogue
- Complicated instructions
When the selected voice cannot handle the difficult passage, you can choose another voice or model before generating the rest of the project.
Change Only One Setting at a Time
When a result sounds wrong, avoid changing the voice, model, stability, similarity, style, speed, and script simultaneously.
A better order is:
- Correct the script.
- Confirm the voice.
- Confirm the model.
- Adjust stability.
- Review similarity.
- Adjust speed.
- Test style or expressive instructions last.
Changing one control at a time makes it easier to determine which adjustment actually improved the result.
Begin Near the Default Voice Settings
The default settings provide a useful starting point for most voices. Extreme values can make the narration less predictable.
When the voice sounds too chaotic, raise stability slightly.
When it sounds too flat, lower stability slightly or choose a more expressive voice.
When a clone sounds insufficiently similar, increase similarity gradually.
When the narration is too fast or slow, adjust speed in small increments rather than moving immediately to the minimum or maximum. ElevenLabs currently supports speed settings between 0.7 and 1.2 and warns that extreme values may reduce audio quality.
Keep Style Exaggeration Low
Style exaggeration may amplify the selected voice’s emotional character, but it can also increase instability, unusual pacing, mispronunciations, and unwanted sounds.
Leave it at or near zero for:
- Tutorials
- Business narration
- Educational content
- Training videos
- Product demonstrations
- Informational podcasts
Increase it only when the project genuinely needs a more exaggerated performance.
Use Natural Punctuation
Punctuation can guide the voice without requiring complicated settings.
Use:
- Periods for complete stops
- Commas for brief pauses
- Question marks for questions
- Exclamation points for stronger energy
- Dashes for interruptions
- Ellipses for hesitation
For example:
“Wait — did you save the file?”
or:
“I thought the recording was gone… but it was still in History.”
Ellipses can produce hesitant or uncertain delivery, so they should not replace ordinary punctuation throughout the script.
Use Break Tags Sparingly
Supported models can use SSML break tags for more precise pauses:
<break time="1.0s" />
Breaks can currently be set for up to three seconds. However, placing too many break tags in one generation may cause unstable pacing, additional noise, or audio artifacts. Eleven v3 does not use SSML break tags and instead relies on punctuation, text structure, and audio tags.
Use a break tag only when a period or paragraph break does not create enough separation.
Match Audio Tags to the Voice
Eleven v3 supports expressive instructions such as:
[whispering][excited][sad][laughs][sighs][sarcastic]
The selected voice still needs to support that type of performance. A restrained professional voice may not convincingly respond to [shouting], while a highly energetic character may struggle to sound calm and neutral.
ElevenLabs recommends matching tags to the voice’s natural character and training style.
Do Not Overuse Audio Tags
Use tags when the emotion changes or an audible reaction is important. Adding several instructions to every sentence can make the result sound exaggerated or inconsistent.
For a normal tutorial, clear wording and punctuation may be enough.
Use expressive tags mainly for:
- Character dialogue
- Storytelling
- Advertisements
- Reactions
- Important warnings
- Emotional introductions
- Dramatic transitions
Write Numbers and Symbols as Spoken Words
Numbers, dates, currencies, phone numbers, URLs, and abbreviations can be interpreted in several ways.
When accuracy matters, write the exact spoken form.
For example:
$1,500→one thousand five hundred dollars2026→twenty twenty-six7/28→July twenty-eighthAI→A.I.vaxali.com→vaxali dot com
ElevenLabs enables text normalization by default, but smaller and faster models may interpret complex numbers differently from larger narration-focused models.
Always preview financial figures, dates, measurements, and contact information.
Use Pronunciation Tools for Repeated Terms
When the same name, acronym, or technical term appears throughout the project, create a reusable pronunciation rule rather than correcting it manually every time.
Depending on the model and workflow, you can use:
- Phonetic spelling
- Alias rules
- Pronunciation dictionaries
- IPA notation with Eleven v3
- Supported phoneme tags
Different models support different pronunciation methods, so confirm compatibility before repeatedly generating the same line. Eleven v3 supports IPA written between forward slashes, while alias substitutions are useful for models that do not support the same phoneme controls.
Generate More Than One Version
AI speech is not completely deterministic. Two generations using the same script, voice, and settings may contain small differences in emphasis, timing, or emotional delivery.
When the first result is close but not perfect:
- Generate it again.
- Compare the two versions.
- Keep the stronger performance.
- Avoid changing the settings unless the same problem repeats.
Variation is particularly noticeable at lower stability settings.
Save Strong Generations Immediately
Do not assume that a strong performance can be recreated exactly later.
After approving a version:
- Download it.
- Rename it clearly.
- Save the script.
- Record the voice and model.
- Record the settings.
- Keep it until the complete project is finished.
History can help retrieve previous generations, but maintaining a local copy provides additional protection.
Keep the Voice and Settings Consistent
For related videos, course lessons, audiobook chapters, or podcast sections, reuse the same:
- Voice
- Model
- Stability
- Similarity
- Style
- Speed
- Pronunciation rules
A narrator can sound noticeably different when the model or settings change, even when the same voice remains selected.
Create a simple production note containing the approved settings and store it with the project files.
Use Clean Recordings for Voice Cloning
A cloned voice reflects the source recording. Background noise, echo, inconsistent volume, exaggerated performance, and changing microphone quality can all affect the clone.
Use:
- One speaker
- One microphone
- One recording environment
- Consistent volume
- Clear pronunciation
- A delivery style that matches the intended project
Do not expect similarity settings to completely repair poor cloning material.
Use Longer, Continuous Voice Samples
Very short or disconnected cloning samples may create unnatural pacing. ElevenLabs recommends using longer, continuous samples when creating a voice to reduce problems such as unnaturally fast speech.
Remove long silence and failed takes, but preserve enough natural speech for the system to understand the speaker’s rhythm and delivery.
Clean Source Audio Before Processing It
Before using Voice Changer, Speech to Text, dubbing, or voice cloning, remove unnecessary noise when possible.
A cleaner source helps the system identify:
- The intended speaker
- The correct words
- Emotional delivery
- Accent
- Pauses
- Speaker changes
Voice Isolator may help with moderate noise, but rerecording the passage is often better when the original speech is severely damaged.
Use Studio for Longer Content
Studio makes it easier to manage chapters, speakers, paragraph-level generations, pronunciation rules, media, and exports.
Use Studio when creating:
- Audiobooks
- Online courses
- Long videos
- Podcasts
- Narrated articles
- Projects with several speakers
- Productions containing music and effects
ElevenLabs specifically recommends Studio for content longer than a few thousand characters.
Review Imported Documents
When importing a webpage, PDF, DOCX, EPUB, or another document into Studio, inspect the text before generating audio.
Remove:
- Page numbers
- Headers and footers
- Citations
- Navigation menus
- Image captions
- Tables
- Footnotes
- Production notes
- Text that should remain visual
ElevenLabs warns that document imports can require manual adjustments because file formatting varies.
Regenerate Glitches Instead of Rebuilding Everything
Occasional generations may contain:
- Whispering
- Accent changes
- Muffled speech
- Volume shifts
- Unexpected sounds
- Abrupt paragraph transitions
When this occurs:
- Review the script.
- Generate the affected paragraph again.
- Increase stability slightly when necessary.
- Return style exaggeration toward zero.
- Confirm the voice matches the language.
- Try another model when the problem continues.
ElevenLabs states that regenerating the affected passage often resolves rare corrupted or inconsistent speech.
Review Audio in Context
A clip may sound good alone but feel inconsistent when placed between two other sections.
Listen to:
- The previous sentence
- The new generation
- The following sentence
Check whether the voice’s tone, speed, emotion, volume, and accent remain consistent.
This is especially important when replacing one paragraph inside a longer Studio project.
Add Music and Effects Last
Approve the narration before adding background music or sound effects.
A practical order is:
- Finalize the script.
- Generate the narration.
- Correct pronunciation.
- Arrange the voice clips.
- Add sound effects.
- Add music.
- Adjust the final mix.
- Export the project.
Music can hide speech problems and make it harder to compare alternate generations accurately.
Listen on Different Devices
Review the finished production through:
- Headphones
- Phone speakers
- Computer speakers
- A television when appropriate
- A car audio system for podcasts or audiobooks
A voice that sounds clear through studio headphones may sound too quiet or harsh through a phone speaker.
Pay attention to speech clarity, music volume, sudden changes in loudness, and distracting background effects.
Ask Another Person to Review It
After listening to the same project repeatedly, it becomes easy to overlook mistakes.
Ask another person to check:
- Pronunciation
- Clarity
- Naturalness
- Speaking speed
- Repeated sentences
- Missing sections
- Volume balance
- Whether the narrator fits the subject
When the project uses another language, ask a fluent speaker to review the pronunciation and wording before publishing it.
Keep the First Project Simple
Do not begin by combining voice cloning, dialogue, dubbing, music, sound effects, transcription, and a long Studio timeline.
A better first project is:
- Write a 30-second script.
- Select one ready-made voice.
- Use Multilingual v2.
- Generate the narration.
- Correct one pronunciation issue.
- Download the audio.
- Add it to a simple video.
Once that workflow feels comfortable, introduce one additional ElevenLabs feature at a time.
Complete a Final Quality Check
Before publishing the project, confirm that:
- The voice suits the audience.
- The model matches the use case.
- The script sounds natural when spoken.
- Names, numbers, and acronyms are correct.
- The accent remains consistent.
- The speaking speed is comfortable.
- No generation contains glitches or strange sounds.
- Emotional instructions are not overused.
- Separate clips transition smoothly.
- Music and effects do not overpower the voice.
- The strongest versions have been downloaded.
- Voice and model settings have been recorded.
- The final audio has been reviewed on more than one device.
- The content and voices are authorized for the intended use.
The best ElevenLabs results usually come from several small improvements rather than one complicated prompt or extreme setting. Start with a suitable voice and clear script, generate a short test, and refine only the specific parts that need correction.
Frequently Asked Questions
Is ElevenLabs free to use?
Yes. ElevenLabs automatically places new users on its Free plan, which can be used to test Text to Speech and other basic creative tools. Free-plan content does not include commercial usage rights and requires attribution when it is shared publicly for noncommercial purposes.
Can I use ElevenLabs voiceovers commercially?
Content generated during an eligible paid subscription can generally be used commercially, provided you have the necessary rights to the script, voice, and other material and follow ElevenLabs’ terms and prohibited-use rules. Content produced while subscribed retains its commercial license even after the subscription ends.
Free-plan generations cannot be used commercially.
Which speech model should beginners use?
Eleven Multilingual v2 is usually the safest starting point for tutorials, courses, audiobooks, and longer prerecorded narration because it prioritizes stable, natural speech.
Eleven v3 is better suited to emotional performances, audio tags, characters, and multi-speaker dialogue. Flash v2.5 prioritizes faster and less expensive generation for supported workflows.
Test the same short paragraph with two models when you are unsure which one fits the project.
Can one ElevenLabs voice speak several languages?
Yes. A voice can speak any language supported by the selected speech model. ElevenLabs determines the language from the written text, while the selected voice strongly influences the accent and pronunciation.
For the most natural result, choose or clone a voice that already speaks the target language with the correct regional accent.
Why does the voice have the wrong accent?
The written text controls the language, but the voice itself influences the accent. An English-trained voice reading Spanish may retain an English accent or move between pronunciations.
Try a voice trained in the required language and region. Also avoid mixing several languages inside one short generation because that can make automatic language detection less reliable.
Can I clone another person’s voice?
Instant Voice Cloning requires permission from the voice owner.
Professional Voice Cloning is more restrictive. ElevenLabs allows users to create a Professional Voice Clone only of their own voice, even when another person has given permission. The other person must create and verify the clone through their own account before sharing it privately.
Never use voice cloning to impersonate someone deceptively or make them appear to say something they did not authorize.
Can I download or export a cloned voice?
You can generate and download audio using an authorized cloned voice, but you cannot export the underlying voice model from ElevenLabs. Voice clones remain usable only through the ElevenLabs platform.
Keep the original source recordings securely stored. They will be needed when the clone must be recreated later.
What formats can I download?
Text to Speech and Voice Changer generations can be downloaded in formats including MP3 and WAV. Additional download options may include M4A and FLAC.
Use MP3 for smaller finished files and WAV when the audio will receive more editing or professional mixing.
Can I recover an earlier generation?
Previous Text to Speech and Voice Changer generations can normally be reopened and downloaded through History. Studio also includes paragraph-level Generation History, which allows users to listen to, download, and restore previous versions.
Important files should still be downloaded and backed up locally because ElevenLabs does not guarantee that every generation will remain accessible indefinitely after a subscription ends.
Does every generation use credits?
Pressing Generate normally deducts credits because ElevenLabs must create the audio before it can be played. Some Text to Speech, Voice Changer, and Studio workflows may provide limited free regenerations when the text, voice, model, and other qualifying conditions remain unchanged.
Check whether the button says Generate or Regenerate and review the displayed cost before confirming.
Should I use Text to Speech or Studio?
Use the standard Text to Speech Playground for short scripts, voice tests, individual clips, and quick narration.
Use Studio for longer scripts, audiobooks, courses, narrated articles, and projects requiring several paragraphs, speakers, visuals, music, captions, or sound effects. ElevenLabs specifically recommends Studio for content longer than a few thousand characters.
Studio also makes it easier to regenerate individual paragraphs or selected phrases instead of recreating the complete narration.
Can I use ElevenLabs for dubbing?
Yes. Dubbing is available on all ElevenLabs plans, including the Free plan. It can translate audio or video into supported languages while attempting to retain the original speakers’ vocal characteristics and delivery.
Important translations should still be reviewed by a fluent speaker before publication.
What should I do when a generation sounds wrong?
First identify whether the problem comes from:
- The script
- The selected voice
- The speech model
- Pronunciation
- Stability or similarity settings
- Speaking speed
- Insufficient context
Correct the script first, then regenerate the complete phrase or sentence. ElevenLabs recommends regenerating a full phrase or sentence instead of replacing only one isolated word when possible.
Change only one setting at a time so you can identify what improves the result.
Do unused credits roll over?
Unused paid-plan credits can roll into the next billing cycle, up to the permitted rollover limit, when the account remains on the same eligible subscription. Free-plan credits do not receive the same rollover benefit. Canceling or downgrading causes unused subscription credits to be lost when the change takes effect.
Review the renewal date and remaining balance before changing the subscription.
Start Creating With ElevenLabs
ElevenLabs gives beginners a simple way to create realistic AI narration without recording every script manually or learning complicated audio-production software. Its browser-based creative tools do not require coding, and the basic Text to Speech process can be completed by selecting a voice, entering a script, choosing a model, and generating the audio.
The platform includes many additional features, but beginners should not try to learn everything at once. ElevenCreative now combines Text to Speech, voice customization, Voice Changer, Studio, dubbing, music, sound effects, and other creative workflows within the same platform. It also provides access to more than 10,000 community voices.
A simple first workflow is:
- Create a free ElevenLabs account.
- Open Text to Speech.
- Select a ready-made voice.
- Enter a short script.
- Start with Multilingual v2 for stable narration.
- Generate and review the audio.
- Correct any pronunciation or pacing problems.
- Download the strongest version.
- Save the voice, model, and settings.
- Move into Studio only when the project becomes longer or more complicated.
ElevenLabs currently recommends Multilingual v2 for stable multilingual narration, Eleven v3 for emotional and expressive speech, and Flash v2.5 for fast, low-latency applications. Choosing the model according to the project is usually more effective than automatically selecting the newest option.
For a first project, create a voiceover lasting approximately 30 seconds. Use one narrator, one language, and a simple script. Do not begin with voice cloning, several speakers, music, dubbing, and a complete video production at the same time.
After the basic workflow feels comfortable, add one feature at a time:
- Explore the Voice Library when the default voices do not fit the project.
- Use Voice Design when an original synthetic narrator is needed.
- Clone only your own voice or a voice you are authorized to use.
- Use Voice Changer when you want to preserve a recorded performance.
- Move longer narration into ElevenCreative Studio.
- Add dialogue when the script requires multiple speakers.
- Generate music and sound effects after approving the narration.
- Use Speech to Text for transcripts and subtitles.
- Clean noisy recordings with Voice Isolator.
- Use Dubbing to create carefully reviewed language versions.
Studio is the better choice for audiobooks, podcasts, narrated articles, courses, and other long-form projects because it organizes narration and media on a timeline and allows individual sections to be regenerated.
The quality of the finished audio depends on more than the selected voice. A strong result normally requires:
- A script written for listening
- A voice suited to the language and audience
- The correct speech model
- Clear punctuation
- Accurate pronunciation
- Consistent settings
- Short test generations
- Careful review before publishing
When the audio sounds wrong, identify the specific problem before changing anything. Rewrite an unnatural sentence before adjusting several settings. Try a different voice when the accent is unsuitable. Change the model when the project requires greater stability, expression, or speed. Regenerate only the affected section rather than recreating the entire project.
ElevenLabs can accelerate voice production, but human review remains important. Listen carefully to names, dates, prices, technical terminology, emotional lines, and translated content. AI-generated speech can sound convincing while still pronouncing a word incorrectly or changing the meaning of a sentence.
Users should also confirm that they have permission to use every script, recording, voice, image, video, and piece of music included in a project. Voice cloning and transformation should never be used to deceptively impersonate another person or make someone appear to say something they did not authorize.
The Free plan is suitable for learning and personal testing, while commercial use of generated audio generally requires an eligible paid plan. Review the current subscription details and applicable terms before publishing monetized videos, client work, advertisements, courses, or other commercial material.
For most beginners, ElevenLabs is easiest to learn by completing a real but manageable project. Create one short voiceover, correct it, download it, and place it into a simple video. That process teaches the most important parts of the platform before more advanced tools are introduced.
Once that first project is complete, the same basic skills can be expanded into longer narration, multilingual videos, podcasts, audiobooks, character dialogue, sound design, and other professional audio productions.
Related Articles
- How to Use Pictory: A Beginner’s Guide (2026)
- How to Use Synthesia: A Beginner’s Guide (2026)
- How to Use InVideo AI: A Beginner’s Guide (2026)
- How to Use Descript: A Beginner’s Guide (2026)
- How to Use Murf AI: A Beginner’s Guide (2026)
- How to Use Speechify: A Beginner’s Guide (2026)
- How to Use Canva: A Beginner’s Guide (2026)
- Best AI Voice Generators for Beginners (2026)
