Blog
Updated Tutorial9 min read

How to Write Dialogue for AI Video: Scripts, Audio Tags, and Timing

How to write dialogue for AI video: short speaker lines, one bracketed delivery tag each, silent beats, and one-take performance so voices react. With samples.


To write dialogue for AI video, write it as a script of short lines. Give every line a named speaker and open it with one delivery tag in brackets, like [whispers] or [dry, quietly pleased]. Put what the audience watches without speech into silent beats. Then perform the whole scene as one take, so the voices hear and react to each other. In Paintbrush, each chapter of a story holds its script, an AI voice director adds tags, pauses and stress to new or changed lines, and ElevenLabs Eleven v4 performs the whole chapter as one take. Speech costs 8 credits per 1,000 characters as of October 2026.

What does a dialogue script look like in Paintbrush?

Dialogue lives in chapters. A chapter is one continuous stretch of story in one setting, written as a script and performed as one piece of audio. Every shot in your video is a cut of a chapter. The script has two kinds of items:

  • Lines. A speaker (a character or the Narrator) and the text: the words plus audio tags in brackets, with no quotation marks. A line can be marked offscreen, for a voice from another room or out of frame.
  • Silent beats. An action the audience watches with no speech: a look, a door opening. Give it a length in seconds, or leave it empty and it is chosen for you, between 0.5 and 10 seconds.

You can also paste a whole story or screenplay, up to 50,000 characters, into the creation agent. It writes each chapter's script with speakers, tags and silent beats, and shows it for approval. See how to turn a script into an AI video.

What are audio tags?

Audio tags are short instructions in square brackets that tell the voice how to perform. Eleven v4 reads them as direction, so they are not spoken. [whispers] makes a whisper. [sighs] adds a sigh. A tag carries forward until the next one, so never repeat it.

How do you write a good audio tag?

Brief the voice the way you would brief an actor:

  1. Open each line with one short tag. One to five words that describe how the voice sounds: [bitter], [dry, quietly pleased], [low, shaking voice], [barely holding it together].
  2. Add a second tag only where the delivery turns. A realisation, a joke landing, rising panic. Most lines need one tag. Never stack more than two in a row.
  3. Be specific. [let down] or [bitter] performs better than [sad]. [anxious] performs better than [scared].
  4. Describe the voice, not a thing. A tag that names an object, a scene or a type of person can be performed as a sound effect. Write [low, gravelly voice], not [a tired detective].
  5. Never use sound effects as tags. No [footsteps] or [door slams]. The video model supplies ambience and sound effects.
  6. Size the tag to the moment. Save shouting and sobbing for the scenes that earn them.

What else shapes the delivery besides tags?

  • Ellipses for hesitation or trailing off: I thought... never mind.
  • A dash for a break or a self-interruption.
  • Commas for a breath.
  • [pause] or [long pause] inside a line for a held beat.
  • One stressed word in capitals, where a person would lean on it: You did WHAT? Use it once or twice in a line at most.

Before and after: fixing weak tags

BeforeAfterWhy
[sad] I miss him.[let down, quietly] I miss him.A precise feeling performs better than a broad one.
[a grumpy old sailor] Get off my boat.[low, gravelly voice] Get off my boat.Describe the voice, not the person.
[thunder crashes] Get inside![shouting, urgent] Get inside!Sound effects come from the video, not the voice.
[angry] [furious] [shouting] Give it back.[angry] Give it back. [shouting] NOW.One tag to open, one where it turns.

Why perform the whole scene as one take?

Most text-to-speech tools read one line at a time. Each line sounds fine alone, but nobody is listening to anybody. Every reply lands with the same gap, and a shout gets a calm answer.

Paintbrush performs the whole chapter with ElevenLabs Eleven v4 Text to Dialogue as one take. Every voice hears the others. A comeback comes quickly. A reply to a confession waits. Because the timing between lines comes from the performance, a beat you care about belongs in the line as a [pause], an ellipsis or a dash.

What does the voice director add?

When you press Generate audio, an AI voice director reads the whole chapter plus a briefing: the story, each speaker's character and voice, and how the previous chapter ended. It directs only new or changed lines with audio tags, pacing punctuation, [pause] beats and a stressed word in capitals. It keeps your words exactly as written, keeps every tag you wrote, and fits the delivery to the voice: a gentle voice will not sell a scream. So you can tag every line yourself, or none.

What happens when I change one line?

Editing a line later re-performs only the changed lines. Unchanged takes are kept and free. Through the creation agent, only the affected shots are re-cut, and you can ask in plain words: "change line 3" or "give her a quieter voice".

How do you write lines that work in AI video?

  • Keep lines short. A dialogue shot usually carries 2 to 6 seconds of speech. A line you cannot say in one breath makes a long shot, or forces a cut in the middle of a thought.
  • One idea per line. If a line makes a point and then changes the subject, make it two lines.
  • Give reactions their own silent beat. A look, a pause, a hand on a shoulder. The cut director gives weighty moments like a touch or a reaction their own shot, and a silent beat makes the moment exist.
  • Separate narration from character lines. Narration goes to the Narrator speaker, read by the project narrator voice. Do not put "she said" or stage directions inside a character's line.
  • Use offscreen lines on purpose. A voice from the hallway, a phone call. Narration counts as offscreen too. Offscreen lines move no mouth, so a line can play over a listener's face without the listener mouthing it.
  • Give every character a distinct voice. Voices come from the ElevenLabs Voice Library, searchable by text, gender, age and accent. The creation agent casts from about 50 hand-picked character voices and gives characters who share a scene clearly different ones. More in character voices and narration.
  • Read it aloud. If you trip over a line, so will the voice.

For three or more speakers, see multiple characters in one AI video scene.

How does timing work?

In Paintbrush, the audio sets the length of the video. You never set a shot duration by hand.

  1. The chapter is performed first. Its audio fixes when every line starts and ends, and how long each silent beat lasts.
  2. The cut director splits the chapter into shots. It opens on an establishing shot, moves closer as emotion rises, and uses shot and reverse shot for conversations. Dialogue shots are usually 2 to 6 seconds. Cutting is free.
  3. Each clip is the shortest whole-second clip that holds its slice of audio, then trimmed to fit. As of October 2026, every project, in every style, uses Seedance 2.5 with clips of 4 to 30 seconds.
  4. You can adjust the cuts. Split a shot, merge it with a neighbour, or drag a boundary. A dragged boundary snaps to the nearest silence between lines, so a cut never lands mid-word.

Long speeches make long clips, and every second of video is generated and billed, at 30.3 credits per second as of October 2026. For each shot, Seedance hears the chapter's recording and performs the on-camera lines itself, voice and mouth together, so there is no separate lip-sync step. Narration and offscreen lines move no mouth. Details in how to make AI characters talk.

What does AI dialogue cost?

As of October 2026, speech costs 0.008 credits per performed character, or 8 credits per 1,000 characters. Tags and punctuation count as characters. A one-minute conversation is roughly 800 to 1,000 characters, so under 10 credits. Unchanged lines are free when you re-generate.

The voice is the cheap part. A 5-second shot with on-camera dialogue costs 151.5 credits of video, lip sync included, so tight dialogue saves far more on video seconds than on speech. See what AI video costs per minute.

Sample chapter: before and after direction

Setting: a kitchen at night. Staging: @Mara sits at the table under one lamp. @Theo, her younger brother, comes in from the hallway. Summary: a letter from their estranged father arrived today. Quiet, tense, tender.

As written

  1. Silent beat (3s): Mara sits alone at the table, an unopened envelope in front of her.
  2. Theo (offscreen): You're still up?
  3. Mara: Couldn't sleep.
  4. Silent beat: Theo steps into the doorway and sees the envelope.
  5. Theo: Is that from Dad?
  6. Mara: It came this morning. I haven't opened it.
  7. Theo: Then open it. Or I will.
  8. Silent beat: Mara slides the envelope across the table to him.
  9. Mara: You do it. I can't.

After the voice director

  1. Silent beat (3s): Mara sits alone at the table, an unopened envelope in front of her.
  2. Theo (offscreen): [sleepy, surprised] You're still up?
  3. Mara: [flat, tired] Couldn't sleep.
  4. Silent beat: Theo steps into the doorway and sees the envelope.
  5. Theo: [quietly] Is that... from Dad?
  6. Mara: [guarded] It came this morning. [pause] I haven't opened it.
  7. Theo: [impatient] Then open it. [half joking] Or I will.
  8. Silent beat: Mara slides the envelope across the table to him.
  9. Mara: [voice breaking] You do it. I CAN'T.

The words did not change. The ellipsis gives Theo a moment before he names their father. The [pause] lets Mara's admission land. The capitals put the weight of the scene on one word. The silent beats give the cut director three visual moments to cut to.

The directed lines total about 240 characters, so about 2 credits to perform. The first line is offscreen, so it moves no mouth: the opening shot can stay on Mara alone while Theo's voice comes from the hallway.

FAQ

Do I have to write audio tags myself?

No. The voice director tags new or changed lines when you generate audio, and the creation agent tags the scripts it drafts. Tags you write yourself are kept.

Are audio tags spoken out loud?

No. Eleven v4 reads bracketed tags as performance direction. Tags that name a sound or an object instead of a voice can be performed as a sound effect, which is why tags should describe the voice: [low, gravelly voice], not [thunder].

Can I use quotation marks in AI video dialogue?

Not in a Paintbrush script. Each line already has a speaker, so the text is just the words and tags. Narration goes on its own Narrator line rather than around a character's quote.

How long should a line of AI dialogue be?

Short enough to say in one breath. Dialogue shots usually carry 2 to 6 seconds of speech. As of October 2026, clips run 4 to 30 seconds, but a long speech plays better as several lines with reactions between them.

Why does my AI dialogue sound robotic?

Usually because each line was generated alone. Perform the whole scene as one take, give each line a specific tag, and add pauses and ellipses where a person would hesitate.

Write your first scene.

Create characters, plan scenes, and bring your story to life in one workspace.

Try Paintbrush