Blog
Updated Feature7 min read

Give Your Characters a Voice — Literally

Assign unique voices to your characters and let Paintbrush automatically read their dialog. Write narration and dialogue in your scenes and every character speaks in their own voice.


AI video tools can generate stunning visuals, but they've always been silent films. You'd generate your scenes, export them, and then spend as long on audio as you did on video — recording voiceover, syncing timing, juggling text-to-speech tools in separate tabs. The gap between visuals and audio has been one of the biggest friction points in the entire workflow.

That changes today. Paintbrush now lets you assign a unique voice to each character and write dialog directly into your scenes. When you generate, every character speaks in their own voice — automatically. No external TTS tools, no manual syncing, no post-production audio editing.

How it works

The system has two parts: voice assignment and scene dialog.

When you edit a character, you'll see a new Voice dropdown. Choose from 20 ElevenLabs voices spanning different genders, accents, and tones — from a deep, confident narrator to a warm British storyteller to a casual Australian. Each voice has a preview so you can hear exactly what it sounds like before committing.

Then, in any scene, write narration or dialog in the audio text field. Paintbrush generates the speech with ElevenLabs v4, times the shot around it, and lip-syncs any character who speaks on screen. Your visuals and voiceover are in sync from the start.

Voices that match your characters

Every voice is labeled with the attributes that matter most when casting: gender, accent, and tone, such as "Male · British · authoritative", "Female · American · upbeat", or "Non-binary · American · calm". Need a British voice with gravitas for your villain? He's in there.

Accents include American, British, Australian, Transatlantic, and Swedish. The preview button lets you hear a sample instantly, right in the character editor. Set a separate narrator voice for the project if your story has narration.

Once a voice is assigned, it sticks. Every scene where that character has dialog uses the same voice automatically. Recast a character's voice at any point and future generations pick up the change.

Write dialog, not prompts

Each scene has an audio text field where you write what should be spoken during that scene. This can be straight narration, character dialog, or a mix of both. Write naturally — the way you'd write a script. Put dialogue in quotes and assign each quote to a speaker; anything outside quotes is read by the narrator.

The audio text is separate from the scene's visual description. Your visual prompt tells the AI what to show; the audio text tells it what to say. This separation means you can describe complex visual action in the prompt while keeping the spoken dialog simple and natural, or vice versa.

Automatic timing

One of the most tedious parts of video production is matching audio to video length. Paintbrush handles this automatically. When you generate a scene with audio text, the system first generates each line, lays them out with a short lead-in and natural pauses between speakers, and measures the result. The video is then generated at a length that holds all of the speech. On Auto, the shot length also accounts for the action, so a line followed by a slow reaction gets room to breathe instead of being cut off.

If the speech runs longer than the longest clip the model can make (15 seconds on Standard, 30 on Pro), Paintbrush splits it at a sentence boundary into continuation scenes. Each continuation starts from the last frame of the one before, so the visuals flow across the split, and every line stays with its speaker. You write one block of narration and the system handles the segmentation — no manual splitting required.

Lip sync and delivery

When a character speaks on screen, the video model animates their mouth from the recorded line as it generates the clip. This works on cartoon, anime, animal, and creature faces as well as realistic ones. Narration and lines marked offscreen are heard but not mouthed.

You can direct the performance with delivery tags inside a quote, like "[whispers] Don't move." or "[laughs] You're kidding." The tag is performed, not spoken, and can describe as much as you like, such as "[barely containing laughter] You're kidding." The creation agent tags every line for you. For a full walkthrough, see how to make AI characters talk.

Works with the creation agent

This is where things get powerful. When you paste a Reddit post, book passage, or script into the creation agent, it carries the narration into each scene, assigns dialogue in quotes to the right speakers, and gives characters voices from the built-in library. Those lines are ready to be spoken the moment you generate.

The workflow becomes: paste a story, approve the character and setting previews and the scene list, and generate. Any voice or line can be changed afterward in the scene editor. You go from a wall of text to a fully voiced, visually consistent animated video in minutes. The narration that used to require a separate recording session is now part of the generation pipeline.

Choosing the right voice

A few tips for getting the most out of character voices:

  • Contrast your cast. If you have two male characters, give them noticeably different voices — one deep and measured, one younger and energetic. Distinct voices help viewers follow dialog without visual cues
  • Match tone to genre. A horror story benefits from a calm, understated narrator. Comedy works better with expressive, dynamic voices. The voice sets emotional expectations before the visuals even register
  • Preview in context. A voice that sounds great in isolation might not fit your character. Listen to the preview while looking at your character's reference sheet — your instinct for the match is usually right
  • Keep narration concise. Short, punchy lines produce tighter scenes. If a narration line runs long, consider splitting it into two scenes for better pacing

What this unlocks

Character voices turn Paintbrush from a visual generation tool into a complete production pipeline. The use cases that open up are significant:

  • YouTube story channels — Reddit stories, creepypasta, and drama channels can go from text to fully voiced animated video without any external tools
  • Children's content — Storybook narration with distinct character voices, perfectly synced to animated scenes
  • Explainer videos — A narrator walks viewers through concepts while a mascot character appears on screen, both with consistent voices
  • Social media series — Episodic content where characters speak in the same voice across every episode, building audience familiarity
  • Prototyping and pitches — Quickly produce voiced storyboards for pitch meetings or concept validation

Audio was the last missing piece in the AI video production chain. Now it's built in.

To prepare the scenes and spoken lines before casting voices, follow the script-to-video workflow.

Give your characters a voice.

Create characters, plan scenes, and bring your story to life in one workspace.

Try Paintbrush