Blog
Updated Guide9 min read

How to Make AI Characters Talk: Voices, Dialogue, and Lip Sync

Give AI-generated characters consistent voices and lip-synced dialogue. Compare native model audio, avatar tools, and a character voice pipeline, and learn to write lines that sound natural.


To make an AI character talk, you need three things: a character who looks the same in every shot, a voice that belongs to that character, and lip sync that matches the mouth to the speech. In Paintbrush, you assign a voice to each character, write the line in quotes in a scene, and assign the quote to the speaker. Paintbrush generates the speech, times the shot around it, and lip-syncs every on-screen line, including on cartoon, anime, and animal faces.

This guide compares the three ways to get talking characters, explains how the Paintbrush pipeline works, and shows how to write dialogue that sounds natural when a machine performs it.

Three ways to make AI characters talk

1. Native audio from a video model

Some video models, including Google Veo, can generate speech along with the clip. That works for a single shot. The catch is that the voice is generated fresh each time, so a character can sound different in the next clip. You also have little control over exactly which line is spoken and how it is timed.

2. Avatar and talking-head tools

Tools like HeyGen and Synthesia animate a presenter reading a script. They are great for training videos and announcements. They are not built for narrative scenes, where characters move through locations and talk to each other.

3. A character voice pipeline

The voice is part of the character, just like the face. Each line is generated with that character's voice, placed on a timeline, and synced to the mouth. This is the approach Paintbrush uses, and it is the only one of the three that keeps a cast sounding consistent across a whole video.

How talking characters work in Paintbrush

Assign voices

Pick a voice for each character from a library of 20 ElevenLabs voices with American, British, Australian, and other accents, and tones from calm to intense. Each voice has a preview. Set a narrator voice for the project if your story has narration. Once assigned, a voice follows the character into every scene.

Write the lines

In a scene's audio text, put dialogue in quotes and assign each quote to a character. Text outside quotes is read by the narrator. For example:

The storm hit just after midnight. "[whispers] Did you hear that?" "It's only the wind. Go back to sleep."

The first sentence goes to the narrator. The first quote goes to one character and the second to another. If a character is talking but not visible, mark them offscreen. They keep their voice but are not lip-synced.

Add emotion with delivery tags

Speech uses ElevenLabs v4, which performs audio tags written in square brackets inside a quote. Tags are performed, not spoken aloud. Open a line with one, then add a new one whenever the feeling shifts, roughly every sentence or two:

"[hushed] Eighteenth hole. One putt to win the championship, and the crowd has gone completely silent. [barely audible] He's lining it up now."

Tags can be single emotions like [smug] or [jittery], reactions like [sighs] or [laughs], or short descriptions: [whispering, fearful], [like a sports commentator, speeding up], [a tired detective who has heard it all before], [starting calm, then losing patience]. The creation agent tags every line for you, and before recording, each scene's lines are directed again with the whole scene in view, keeping any tags you wrote.

Timing and shot length

Each line is generated separately and laid out on the scene's timeline with a short lead-in and natural pauses between speakers. The clip is always at least long enough to hold the speech. On Auto, the shot length is chosen for the action around the dialogue, so a line does not get crammed into a rushed clip. If a passage is too long for one clip (15 seconds on Standard, 30 on Pro), Paintbrush splits it at a sentence boundary into continuation scenes and keeps each line with the right speaker.

Lip sync

The recorded lines go into video generation, and the video model animates the speaker's mouth from the recording itself, so there is no separate lip-sync pass afterward. Narration and offscreen lines are heard but not mouthed by anyone on screen. This works on stylized, cartoon, 3D, animal, and creature faces as well as realistic ones, as long as the speaker's mouth is visible.

Conversations between two characters

For a back-and-forth, the classic approach is a shot and reverse shot: one shot on the speaker, a cut to the listener's reply. Ask the creation agent for "a shot/reverse-shot conversation between @Mara and @Theo" and it plans each spoken beat as its own shot, keeping positions and eyelines consistent. You can also keep both characters in one shared shot when the staging calls for it.

How to write dialogue that sounds natural

  • Keep lines short. One or two sentences per line. Long speeches sound flat and are harder to cut.
  • Write how people talk. Use contractions, fragments, and interruptions. "You're late." is better than "You have arrived later than expected."
  • Use punctuation for pacing. Ellipses add hesitation. Dashes cut a line off. Question marks change intonation.
  • Tag the shifts. A tag carries until the next one, so add a new tag when the feeling changes, not on every phrase. Precise beats broad: [let down] says more than [sad].
  • Contrast your voices. Two similar voices make it hard for viewers to follow who is speaking. Pair a deep, calm voice with a bright, quick one.
  • Leave room for silence. A reaction shot with no dialogue often lands harder than another line.

Troubleshooting talking scenes

The mouth is not moving

The video model can only animate a mouth it can see, so lip sync will not show when the character faces away or something covers the face. Check that the line is not marked offscreen, and describe a framing where the speaker faces the camera, like "medium shot, @Mara facing camera."

The line sounds flat

Make the tags more specific, such as [nervous, trying to sound confident] instead of [nervous], add one where the feeling turns, and rewrite the line the way a person would say it. Shorter lines with clear punctuation usually sound more natural than long sentences.

The wrong character is speaking

Each quote is assigned to a speaker. Check the assignment for each line, and give your characters clearly different voices so a mistake is easy to hear.

What it costs

Speech costs 10 credits per scene, covering every line in that scene. Lip sync costs 13.4 credits per second of on-screen dialogue. A scene with a 4-second spoken line costs about 64 credits for voice and lip sync, plus the video itself (16.8 credits per second on Standard). Rates are as of October 2026.

Ideas for talking-character videos

FAQ

How do I make an AI character's mouth move with the audio?

Use lip sync. In Paintbrush, the video model animates the speaker's mouth from the recorded line while it generates the clip, so you do not need a separate tool.

Can animals and cartoon characters talk?

Yes. Lip sync works on stylized, cartoon, anime, animal, and creature faces as long as the mouth is visible.

Will my character sound the same in every scene?

Yes. The voice is assigned to the character, not generated fresh per clip, so it stays the same for the whole project. You can recast a voice at any time and later generations use the new one.

Can I use my own voice?

Not currently. Characters use voices from the built-in library.

Can two AI characters have a conversation?

Yes. Write each line in quotes, assign each to a character, and use either one shared shot or a shot and reverse shot per line.

Give your characters a voice.

Create characters, plan scenes, and bring your story to life in one workspace.

Try Paintbrush