How to Get Readable Text in AI Video: Phone Screens, Messages, and Signs
Why AI video garbles text on phones, screens and signs, and how to fix it: draw the frame as a still first, keep the text short, then animate it. Step by step.
To get readable text in AI video, do not ask the video model to write it. Draw the shot's first frame as a still with an image model, which renders short text well, put the exact words on the phone, screen or sign in that still, then animate the still with an instruction to keep the screen unchanged. Keep the text to a few words and give it a close insert of its own. In Paintbrush this is built in: when the audience must read a phone, a screen, a page or a sign, that shot gets an On screen field with the exact picture and words, and the shot is drawn as a keyframe first, then animated. As of October 2026 the keyframe adds 8 credits to the shot.
Below: why video models garble text, the workflow that fixes it, how to write what the screen shows, and what it costs.
Why does AI video garble text?
Video models are trained to make motion look right, not to spell. Text is a small, exact pattern that has to stay identical in every frame, and video models redraw every frame. The result is familiar: letters that almost read, words that change as the camera moves, a headline in a language that does not exist.
Screens are worse than signs. A phone in a story usually has to show something composed: a message from a named person, a social post with a photo, a news headline. In our tests, a video model asked for a phone showing a character's post pasted that character's reference portrait into the phone and garbled the caption. A character who appears only on a screen is easily rendered as a person in the room instead.
What is the fix?
Split the job between two kinds of model:
- An image model draws the first frame. Image models compose a picture and render short text far better than video models. The still holds the room, the hand holding the phone, and the screen with its exact words.
- The video model animates that frame. It gets the still as its opening frame and an instruction to keep the screen's picture and text exactly as they are, steady and legible, for the whole shot: no redrawing, scrolling or changing.
The motion in the shot should be small. A thumb resting on the glass, a slight push in, a hand lowering the letter. The more the screen moves, the more chances the model has to redraw it.
How does Paintbrush handle text on screens?
When a chapter's script calls for the audience to read something (a text message, a post, a headline, a letter, a sign on a door), the cut director gives it a close insert and writes what the screen shows into that shot's On screen field: the picture and its exact text, kept short, with an @mention for anyone pictured on it and what they are doing there. You can also add one yourself: open a shot and click Add a phone, screen or sign.
A shot with a screen is always drawn as a keyframe first. The keyframe is made by an image model (Nano Banana 2) from the shot's description, the screen text, the references of everyone it names, the setting, and the real last frame of the earlier shot it is staged from, so the hand, the room and the light match the shot before. Then Seedance 2.5 animates it with the keep-the-screen-steady instruction. Someone pictured only on the screen is not treated as in the room.
A screen shot is never a continuous shot: it is always a cut, which is how films show a phone anyway. Read more about how shots are staged from each other in scene chaining.
How should you write what the screen shows?
Write it like a prop note: what the picture is, then the exact words.
A text message thread. Top: contact name "Mom". One grey bubble: "Are you coming home for dinner?" Below it, one blue bubble: "Can't. Working late." Time 9:41.
A social post by @Jess: a photo of @Jess and @Theo laughing at a bonfire on the beach, @Jess on the left. Caption: "best night ever". 2,481 likes.
Rules that keep it legible:
- A few words, not a paragraph. A caption, a message, a headline. Long text gets small, and small text breaks first.
- One idea per insert. If the story needs three messages, give the reply its own insert after a reaction shot.
- Exact words in quotation marks. In a shot description, quotation marks mean visible text, so do not use them for anything else.
- @mention anyone pictured. Their reference images are what make the photo on the phone look like them.
- Say where the screen is. "In @Mara's hand, close, filling most of the frame" reads better than a phone on a distant table.
- Do not rely on text outside a screen shot. A sign in the background of a wide shot will not be legible. If the audience must read it, give it an insert.
Where does this matter most?
- Text-message and social-media stories. Reddit retellings and drama shorts often turn on a message. See how to turn Reddit stories into animated videos.
- Explainers. A board, a slide or a product screen with a short label. See how to make explainer videos with AI.
- Mysteries and thrillers. A note under the door, a headline, a name on a file.
- Short-form hooks. A notification in the first second. See story TikToks and Shorts.
Captions and subtitles are different: add those in your editor after export, not into the generated picture. Most short-form viewers watch muted, so burned-in captions are worth the five minutes.
What does a screen shot cost?
As of October 2026, video costs 30.3 credits per second in every style, and a screen shot adds an 8-credit keyframe. Inserts are usually silent or play a line from someone off camera, so they also pay 12.5 credits for voice cleanup:
| Shot | Credits |
|---|---|
| 4 s insert, silent or narrated | 141.7 (121.2 video + 12.5 cleanup + 8 keyframe) |
| 5 s insert, silent or narrated | 172 (151.5 + 12.5 + 8) |
| 5 s, the holder speaks on camera | 159.5 (151.5 + 8) |
Clips are at least 4 seconds, so a 2-second glance at a phone still pays for 4. Hold the insert long enough to read: about a second per three or four words, plus a beat. Full pricing is in what AI video costs per minute.
Can you do this in other AI video tools?
Yes, by hand. Generate the frame in an image model that renders text well, check the spelling, then use image-to-video with a prompt that tells the model to keep the screen unchanged and move little. Or composite the text in an editor afterwards, which is the most reliable route for long text and any text that must match a brand exactly. The trade-off is time: every insert becomes a two-tool job, and the hand and room in your still have to match the shots around it.
FAQ
Why can't AI video models write text?
They redraw every frame and are trained for plausible motion, not exact spelling, so text drifts and garbles. Image models render short text far better, which is why the reliable workflow draws the frame as a still first and then animates it.
How long can the text on a screen be?
Keep it to a caption, a message or a headline: a few words. Long text is rendered small and breaks first. For a page of text, use several inserts or add it in an editor.
Can a character appear in a photo on a phone?
Yes. @mention them in what the screen shows. Their reference images are used to draw them in the photo, and they are not placed in the room.
Does Paintbrush add subtitles?
No. Add captions in your editor after export. The On screen field is for text that exists in the story, like a message or a sign.
Put the message on screen.
Create characters, plan scenes, and bring your story to life in one workspace.
Try Paintbrush