How to Write AI Video Prompts That Actually Work
Vague prompts produce vague clips. Here's a practical prompt structure for AI video — action, camera, lighting, and what to leave out once you have character references.
Most bad AI video isn't a model failure. It's a prompt that describes a vibe and hopes the model invents a shot. "Cinematic, emotional, a woman in a city at dusk" is a mood board. The model will give you a plausible woman, a plausible city, and a plausible dusk — and a different trio every time you run it.
A working video prompt specifies the things a camera would have to choose: who is in frame, what they do, where they stand, how close we are, and what the light is doing. Everything else is noise.
The four-part prompt
Write every scene as four pieces, in this order:
- Subject. Who is in the shot. In Paintbrush that's an @mention so the reference sheet is attached. In a generic tool, this is the only place you spend words on appearance — and you'll still get drift.
- Action. A verb you can film in a few seconds. Walks, turns, slams, looks, sits. Not "experiences loss" or "embarks on a journey."
- Camera. Distance and angle: close-up, medium, wide, over-the-shoulder, tracking, locked off. If you don't say, the model picks a new framing every generation.
- Light and weather. One lighting condition. Golden hour, fluorescent office, neon rain, overcast noon. Don't stack three moods in one sentence.
Example: "@Aria sprints across the @Neon Rooftop, medium tracking shot from the side, rain streaks through neon, night." That's a clip. "Aria has a dramatic moment on a roof in a cyberpunk city, cinematic, 8k, masterpiece" is a lottery ticket.
Stop re-describing characters you already built
If the tool has a character system, repeating "red hair, green jacket, silver pendant" in every scene is how you introduce contradictions. The reference sheet already holds those details. The scene prompt should only mention appearance when something changes — a soaked coat, a removed hood, blood on a sleeve.
If you're stuck in a prompt-only tool, you don't have that luxury. Then you do need the full description every time, and you should expect to regenerate. That's the consistency problem in a sentence.
Words that help, words that waste credits
Helpful: facing camera, three-quarter view, looking off-screen left, slow push-in, handheld, locked tripod, shallow depth of field, backlight, practical lamps, overcast, hard noon sun.
Wasteful: cinematic, stunning, masterpiece, 8k, highly detailed, trending on ArtStation, award-winning, epic. Those tokens compete with the words that actually specify the shot. Quality adjectives don't make the model more careful. They make it more generic.
One idea per clip
A 5-second generation cannot do "she enters, argues, cries, then leaves." Pick the beat you would hold on in an edit. Entering is a shot. The argument is a shot. The exit is a shot. If you need a longer thought, write it as narration over two short clips, not as one overloaded prompt.
The same rule applies to crowds. Two named characters is the reliability line. A third person is a cameo, not a conversation. Split group scenes into cuts.
Match the prompt to the style
Anime and cartoon prompts should name silhouette and color, not skin pores. Photoreal prompts should name materials and light, not "looks real." A watercolor project wants atmosphere words (mist, paper texture, soft edges) and will fight you if you also ask for razor-sharp product photography.
Set style at the project level and keep scene prompts about staging. Mixing "Studio Ghibli" in one scene and "handheld documentary" in the next is how a sequence stops looking like one video.
A quick rewrite exercise
Weak: "A knight in a dark forest, cinematic lighting, epic mood, highly detailed armor."
Strong: "@Rowan in dented steel plate stands on the @Pine Path, medium shot, facing camera, overcast daylight, wet leaves underfoot, no other figures."
The second version tells the model what to match (Rowan, the path), what to do (stand, face us), and what the frame contains. That's enough. Generate, watch, then tweak one variable — camera or action, not both — if the first take misses.
The takeaway
Write prompts like a shot list, not like a novel and not like a hashtag pile. Subject, action, camera, light. Leave identity to the reference sheet. Leave speech to the audio field. If a clip is wrong, change one of the four parts and generate again. Prompting gets easier the moment you stop asking the model to invent the movie and start asking it to film the next shot.
To give those prompts a stable visual identity, read how character references work.
Put your next scene into motion.
Create characters, plan scenes, and bring your story to life in one workspace.
Try Paintbrush