Blog
Tutorial10 min read

How to Turn a Script Into an AI-Generated Video

A practical workflow for going from a written script to a multi-scene AI video: break it into shots, lock a cast, generate clips, and add narration without a separate audio stack.


A script is already a production plan. It has characters, locations, dialogue, and a sequence of events. The reason "text to video" still feels hard is that most AI tools ignore that structure. They want one prompt for one clip. You end up rewriting your screenplay as a pile of disconnected descriptions, then wondering why the lead looks different in scene four.

In Paintbrush, open a project and paste your source text into the creation agent. It plans the characters, settings, and a shot list, assigns quoted dialogue to its speakers with voices from the built-in library, and asks you to approve the character and setting previews and the scene list, with a credit estimate, before it generates any video.

This is a working pipeline for turning a script — original, adapted, or pulled from a story post — into a generated video without throwing the structure away.

Step 1: Treat the script as a shot list, not a prompt

Don't paste the whole script into a single generation box. Read it once and mark three things: who appears, where they are, and where the camera would cut. A page of dialogue in one location is usually two or three shots, not one. A location change is always a new scene. Emotional peaks (the reveal, the argument, the punchline) deserve their own clip so you can control pacing later.

If you'd rather not do that markup by hand, paste the script into the creation agent (up to 50,000 characters). It reuses any characters and settings already in the project, plans the rest, and lays out a shot list with narration and dialogue pulled from the source. Read the scene list before you approve — the AI will over-split some beats and merge others. Ask the agent to fix those now. It's cheaper than regenerating video. One plan covers up to 24 shots; for a longer script the agent makes the first part, then offers to continue with the next part using the same cast and settings.

Step 2: Lock the cast before you write a single scene prompt

Create every speaking character (and any silent one who appears more than once) as a saved asset with a specific visual description: hair, face, wardrobe, one distinctive accessory. Generate their multi-angle reference sheet and look at all four views. If the back of the jacket or the profile doesn't match the front, rewrite the prompt and regenerate the character — not the later scenes. Every future clip will inherit this sheet.

Assign a voice while you're here. Contrast the cast: two similar-sounding leads make dialogue hard to follow once faces are stylized or the viewer is on mute with captions.

Step 3: Build settings as reusable locations

A script that returns to "the kitchen" five times needs one kitchen, not five interpretations of the word kitchen. Create each recurring location as a setting: time of day, materials, a few furniture anchors, mood. Scene prompts should then say what happens in that place, not re-describe the room.

Step 4: Write scene prompts as action and camera

Once characters and settings exist, the scene prompt is only the new information: who is doing what, which way the camera faces, the emotional beat. "@Maya slams the @Kitchen table and leans toward @Eli, close-up, warm under-cabinet light" is a shot. "Maya is angry in a kitchen" is a suggestion the model will invent from scratch.

Keep clips short — 3 to 6 seconds for dialogue beats, a little longer for establishing shots. Shorter clips hold likeness better and give you more edit points. Put the spoken line in the scene's audio field, not in the visual prompt. Visuals and speech are separate inputs on purpose.

Step 5: Generate in story order, chain when the camera stays

Generate the first scene, watch it, then decide whether the next shot should continue from the last frame or start fresh. Stay chained for a conversation or a walk through one room. Break the chain for a new location, a time jump, or a parallel storyline. If a clip drifts, regenerate that scene — the references are still attached — before you move on. Drift compounds if you ignore it and keep chaining.

Step 6: Edit for talk time, not for footage

Export the sequence and drop it into CapCut, Resolve, or iMovie. Your job in the editor is pacing: trim motion-blur tails, add captions (most people watch muted), and keep music under the voices. Don't rewrite the story in the timeline. If a beat is wrong, change the scene prompt and regenerate. The script is still the source of truth.

What to do with different kinds of scripts

  • Dialogue-heavy. One thought per clip. Split long monologues. Match duration to the line so you're not padding or cutting speech.
  • Action / set-piece. Chain short clips. Describe motion, not wardrobe. Two characters per shot is the consistency sweet spot.
  • Narrated story (no on-screen speaker). Create a narrator voice and locations; characters can appear without talking. The creation agent already carries narration from your text into each scene.
  • Adapted prose. Cut interiority that can't be shown. If a paragraph is all thought, turn it into a look, a pause, or a line of voiceover over a still-ish shot.

The mistake that wastes the most credits

Generating video before the cast is locked. A beautiful scene of the wrong-looking lead is footage you will throw away. Spend the first session on character sheets, settings, and a reviewed shot list. The second session is generation. That order is the entire difference between a script that becomes a video and a script that becomes a folder of almosts.

Before generating the shot list, establish your cast with how character references work.

Turn your script into a video.

Create characters, plan scenes, and bring your story to life in one workspace.

Try Paintbrush