AI Video Prompts: How to Direct a Shot Instead of Describing One
18 min read ยท Updated 2026-08-23
A good AI video prompt reads like a shot list, not a description: subject and action first, then camera framing and movement, then lighting and lens, then style. The most common mistake is writing a scene instead of a shot โ asking for a story that unfolds over a minute when the model generates a few seconds of continuous motion. Specify one action, one camera move, and one lighting condition, and the hit rate goes up sharply across every generator.
Why video prompting is not image prompting with movement added
People arriving from image generation write video prompts the same way they write image prompts โ a rich static description, a style reference, a mood โ and get clips that either barely move or move in ways that look wrong. The reason is that a still prompt describes a state, and a video prompt has to describe a change.
A video model is resolving two things at once: what is in the frame, and what happens to it over the clip's duration. If your prompt only specifies the first, the model invents the second, and its invention is usually a slow drift โ a lazy push-in, a subject shifting weight, a camera orbit nobody asked for. That is why so many first attempts look like a photograph being gently wobbled.
The fix is structural rather than stylistic. Say what moves. Say what the camera does. Say how long the action takes relative to the clip. Everything else in this article is elaboration on those three instructions.
The shot-list structure that works across generators
Different tools have different syntaxes and different strengths, but the ordering below works everywhere, because it puts the information the model weights most heavily at the front and the negotiable styling at the back.
Write it as one flowing sentence or as labelled lines โ both work. What matters is that every element is present and that they are in this order, because when a model has to compromise, it drops what came last.
- 1. Subject โ who or what, described concretely enough to be consistent: "a woman in her sixties in a grey wool coat", not "a person".
- 2. Action โ the single thing that happens: "turns to look over her left shoulder". One action per clip. Not two, not a sequence.
- 3. Setting โ where, with one or two anchoring details rather than a full environment essay.
- 4. Camera framing โ wide, medium, close-up, over-the-shoulder, low angle, top-down.
- 5. Camera movement โ static, slow push in, pull back, pan left, tilt up, handheld follow, orbit. Or explicitly "locked-off camera, no movement", which is often the most valuable instruction in the whole prompt.
- 6. Lighting โ the single biggest lever on whether a clip looks cheap or expensive: "hard late-afternoon sun from camera left", "single practical lamp, deep shadows", "flat overcast".
- 7. Lens and depth โ "35mm, deep focus" versus "85mm, shallow depth of field, background dissolved".
- 8. Style and grade โ film stock, era, colour treatment, animation style. Last, because it is the most negotiable.
The single biggest fix: one action per clip
Nearly every disappointing generation traces back to a prompt that asked for a sequence. "A chef chops vegetables, then slides them into a pan, then the pan flames up" is three shots. Given a few seconds of output, the model tries to compress all three, and produces something where hands blur through impossible motions and the pan appears mid-frame from nowhere.
Ask for the flame-up alone and you get a clean, usable clip. Ask for the chopping alone and you get another. Then cut them together in an editor, which is exactly how the shot you were imagining would be made anyway.
This reframing โ from "generate my video" to "generate my shots" โ is what separates people who get usable footage from people who conclude the tools do not work. The tools generate shots. Editing is still editing.
- Bad: "a cyclist rides through the city, stops at a cafe, orders a coffee and sits down."
- Good, shot 1: "medium tracking shot, cyclist riding past a row of shopfronts, camera moving alongside at matched speed, late afternoon sun flaring between buildings."
- Good, shot 2: "close-up, hands lifting a coffee cup from a marble counter, shallow depth of field, warm interior light, locked-off camera."
- Good, shot 3: "wide, cyclist sitting at an outdoor table, bike leaning against the railing behind, static camera, soft overcast light."
Camera movement: say it or the model chooses
Camera language is the highest-leverage vocabulary in video prompting, and it is the part most people leave out entirely. An unspecified camera gets an arbitrary one, and arbitrary camera movement is the strongest tell that a clip was generated rather than shot.
Two disciplines matter. First, name the move using real terminology โ models have seen far more footage labelled with proper terms than with descriptions like "the camera kind of moves in". Second, specify the speed, because "push in" without a qualifier tends to come back faster and more aggressive than you wanted.
- Static / locked-off โ the camera does not move. Underused, and the most reliable way to get a clean clip.
- Push in / dolly in โ the camera physically moves toward the subject. Add "slow" or "barely perceptible" unless you want drama.
- Pull back / dolly out โ reveals context. Works well as an opening shot.
- Pan / tilt โ the camera rotates on a fixed point. Different from dolly, and models do distinguish them.
- Tracking / following โ the camera moves with a moving subject, keeping framing constant. Say "matched speed" so the subject does not drift out of frame.
- Handheld โ adds organic instability. Specify "subtle handheld" unless you want a shaky documentary feel.
- Orbit / arc โ the camera circles the subject. Expensive-looking and reliably requested; also the move that most often produces geometry errors on complex subjects.
- Crane / boom โ vertical camera movement. Pairs well with a wide establishing shot.
Lighting is what makes a clip look expensive
If you change one thing about your prompts after reading this, make it lighting. Generic prompts get generic light โ a flat, evenly-lit, slightly grey look that reads as amateur regardless of how good the subject is.
Real lighting has direction, quality and a source. Naming all three costs you a clause and transforms the output.
Direction: where it comes from relative to camera. Quality: hard, with defined shadow edges, or soft and diffused. Source: sun, an overcast sky, a window, a practical lamp in the shot, a fire, a screen, neon signage. "Lit by the laptop screen in a dark room" is a complete lighting design in eight words.
- Golden hour, low sun from camera left, long shadows โ warm, cinematic, forgiving.
- Overcast, soft directionless light โ flat and honest; good for product and documentary looks.
- Single hard source, deep shadows, high contrast โ dramatic; use with a dark environment.
- Backlit, subject rimmed with light, haze in the air โ the most reliably beautiful setup in generated video.
- Practical sources only: lamps, neon, screens, candles โ grounded and specific, and the light source appearing in frame helps the model stay coherent.
- Night exterior, wet ground reflecting light โ high impact, and the reflections give the model motion cues that improve stability.
The failure modes every generator shares
These are not bugs in one tool; they are consequences of how video generation works, and knowing them lets you write around them instead of regenerating twenty times hoping for luck.
The common thread is that models struggle with anything requiring persistent physical logic across frames โ objects that must remain the same object, hands that must have the same number of fingers throughout, text that must stay the same text.
- Hands and fine manipulation โ anything where fingers grip, count or manipulate small objects. Frame it wider, cut away at the critical moment, or choose a different action.
- Text in frame โ signage, screens, labels, packaging. It will warp or morph mid-clip. Add text in your editor afterwards.
- Object permanence โ things that leave frame and return, or pass behind something, often come back changed. Avoid occlusion in the action you request.
- Multiple interacting people โ two subjects touching, handing something over, or shaking hands is significantly harder than one subject alone.
- Fast motion โ running, sports, anything with rapid limb movement. Slow the action down or shoot it wider.
- Long clips โ coherence degrades with duration. Generate short and cut, rather than generating long and trimming.
- Faces in close-up over time โ subtle identity drift across a clip. Shorter clips and less extreme close-ups reduce it.
Image-to-video: the reliability upgrade nobody uses enough
Most video tools accept a starting image as well as a text prompt, and this is the single largest improvement available to anyone struggling with consistency.
The reason is that half the generation problem โ what everything looks like โ is solved before the model starts. You have already fixed the subject, the wardrobe, the environment, the palette and the framing. The prompt then only needs to specify what moves.
This also solves the character consistency problem more cleanly than any prompt trick. Generate or shoot one strong reference frame, then use it as the first frame across multiple clips with different motion prompts. The subject stays the same subject because it started as the same pixels.
- With a start image, write the prompt as motion only: "the subject turns her head slowly to face camera; the camera holds still; leaves move in the background wind." Do not re-describe what is already in the image.
- Do not fight the image. Asking for a camera move the framing cannot support โ a pull-back from an already-wide shot โ produces artefacts as the model invents the space.
- For a sequence, generate your reference frames first as stills, get them right, and only then animate. Fixing a look is far cheaper in stills.
- Where the tool supports both a first and last frame, you are effectively specifying the whole shot arc and the results are dramatically more controllable.
Continuity across a multi-shot sequence
Once you accept that you are generating shots rather than videos, continuity becomes the real craft problem: shot two has to look like it belongs in the same world as shot one.
The technique is a locked block of text that never changes between prompts, holding everything that must stay constant, with only the action and camera lines varying. Treat it like a style bible.
- Write a fixed block: subject description with specific wardrobe and physical details, location description, time of day, lighting setup, lens, grade. Paste it verbatim into every prompt in the sequence.
- Change only the action line and the camera line between shots.
- Keep the time of day and light direction constant unless the story requires a jump โ inconsistent shadow direction between cuts is the most visible continuity error.
- Respect screen direction: if the subject moves left to right in one shot, keep it left to right in the next, or the cut will feel wrong even to viewers who cannot say why.
- Where possible, use a still frame from one clip as the start image for the next. It is the most reliable continuity mechanism available.
Prompt templates you can adapt
These follow the ordering above. Replace the bracketed parts and keep the structure โ the structure is doing most of the work.
- Product hero: "[PRODUCT] on a [SURFACE], slow 180-degree orbit around it, single soft key light from upper left with a subtle rim from behind, dark neutral background, 85mm, shallow depth of field, no text, no hands."
- Talking-head b-roll: "[PERSON DESCRIPTION] at a desk in a [ROOM], glancing down at their notes then back up, locked-off camera, medium shot from slightly below eye level, soft window light from camera right, 50mm, background gently out of focus."
- Establishing exterior: "wide shot of [PLACE], slow crane up revealing [WHAT IS BEYOND], early morning haze, low sun behind the buildings, 24mm, deep focus, no people in frame."
- Food: "close-up of [DISH] as steam rises, static camera, hard side light from camera left with deep shadows, dark wooden table, 100mm macro, very shallow depth of field."
- Atmosphere / mood: "[SUBJECT] standing still in [PLACE], only the wind moving fabric and hair, camera drifts almost imperceptibly forward, backlit with haze, long lens compression, muted desaturated grade."
- Animated / stylised: "[SUBJECT] [ACTION], [ANIMATION STYLE] with visible line work and limited frame rate, flat colour, camera pans slowly right, consistent character design throughout."
Negative prompting and what to exclude
Where the tool supports negative prompts, they are most useful for excluding the specific artefacts that plague video rather than for excluding subject matter. Where it does not, the same instructions work reasonably well phrased positively in the main prompt.
Be aware that in some models, naming a thing to exclude increases its likelihood of appearing โ a known quirk of instruction following. If "no text" produces text, try describing the surface as blank instead.
- Common useful exclusions: warping, morphing, extra limbs, distorted hands, text, watermark, logo, jitter, flicker, sudden cuts, camera shake, motion blur (when you want it clean).
- Positive phrasings that often work better: "clean blank surfaces", "the camera remains completely still", "the subject stays fully in frame throughout", "consistent lighting throughout the clip".
- Excluding a camera move you do not want is often more effective than describing the one you do: "the camera does not move" beats "static shot" in several tools.
Aspect ratio, duration and the platform you are actually publishing to
Decide the output format before you generate, not after. Cropping a landscape clip to vertical costs you the composition you asked for, and generating at the target ratio means the model composes for that frame.
Duration deserves the same up-front thought. A short clip generated cleanly and repeated or extended in an editor beats a long clip that degrades halfway through.
- Vertical for short-form social โ compose for the centre and expect the top and bottom to be covered by interface elements.
- Landscape for anything embedded on a website or intended for a larger screen.
- Square is a compromise that rarely looks intentional; pick one of the other two unless the platform demands it.
- Generate at the shortest duration that contains the action. Extend by cutting, not by asking for length.
- Plan for the cut: give yourself a beat of stillness at the start and end of each clip so you have somewhere to cut without landing mid-motion.
Sound, and why the visual prompt should account for it
Whether your tool generates audio or not, the clip will end up with sound, and the visual prompt determines how easy that is to add convincingly.
The specific trap is dialogue. Generated lip movement rarely matches added audio well, so prompts that put a talking subject in close-up create a problem you cannot fix in the edit. Prompts that keep the subject wide, turned away, or not speaking give you complete freedom to lay voiceover over the top.
- For voiceover-led content, prompt subjects who are doing something rather than speaking. The footage becomes b-roll and any narration fits.
- If you need someone talking on camera, keep it wide or at an angle rather than a front-on close-up, which is where lip sync failures are most visible.
- Ambient sound sells a generated clip more than almost anything else. Prompt for visible sound sources โ rain, traffic, a crowd, a fire โ so the audio you add has an on-screen cause.
- Cut on motion or on a sound hit rather than on silence. Generated footage holds up better when the edit gives the viewer something else to attend to.
A workflow that produces usable footage instead of a folder of near-misses
The people getting consistently good results are not writing better single prompts. They are running a process that concentrates effort where it pays.
The core insight is that iterating on stills is orders of magnitude cheaper than iterating on video, and that most of what makes a clip good โ subject, framing, lighting, palette โ can be resolved in a still.
- 1. Write the shot list first, in plain language, before touching any tool. How many shots, what happens in each, how they cut together.
- 2. Generate the key frames as still images. Iterate here until the look is right. This is where you fix wardrobe, palette and framing.
- 3. Animate from those stills, with motion-only prompts.
- 4. Generate three to four variations of each shot rather than one. Selection is part of the craft; expect to discard most of what you generate.
- 5. Cut in an editor. Add text, audio and grading there, not in the prompt.
- 6. Keep a file of the prompts that worked, with a note on the tool and settings. Video prompting rewards a personal library far more than image prompting does, because the variables are harder to hold in your head.
Disclosure, likeness and the rules that actually apply
Two practical constraints worth knowing before you publish, both of which are about people rather than technology.
The first is likeness. Generating a recognisable real person โ a celebrity, a public figure, or anyone identifiable โ in a scenario they did not participate in carries real legal exposure in most jurisdictions, and every major platform has rules against it independent of the law. This applies as much to a synthetic endorsement as to anything more obviously malicious.
The second is disclosure. Most major platforms now require or strongly encourage labelling AI-generated media, and some detect and label it automatically. Advertising standards bodies in several markets treat undisclosed synthetic content in ads as a compliance matter. Check the specific rules for the platform you are publishing to; they change often enough that any list here would go stale.
- Do not generate identifiable real people without consent, and be aware that "it is obviously fake" is not a defence in most frameworks.
- Label AI-generated content where the platform asks for it โ and assume detection will label it for you if you do not.
- For commercial work, check whether the tool's terms grant you commercial rights to the output. They vary considerably by provider and tier.
- Keep your source material clean. If a client asks where footage came from, "generated, here is the prompt and the tool" is a complete answer; "somewhere on the internet" is not.
How to get better at this quickly
Video prompting rewards vocabulary more than almost any other prompting skill, because the vocabulary is a hundred years old and the models have absorbed all of it. Every term a cinematographer uses is a lever you can pull, and every term you do not know is a lever you are leaving alone.
The fastest improvement available is not learning a tool. It is watching footage you admire and writing down, in shot-list language, exactly what it is doing โ the framing, the move, the light direction, the lens. Do that for twenty shots and your prompts will change more than they would from any amount of tool switching.
The second fastest is to keep everything. Save the prompts that worked, the seeds where the tool exposes them, the reference frames, and the notes about what failed. Generated video is a numbers game with a strong skill component, and the skill compounds only if you can remember what you did.
Frequently Asked Questions
What makes a good AI video prompt?
Structure, in this order: subject, one action, setting, camera framing, camera movement, lighting, lens, style. The most common failure is describing a scene that would take several shots to film when the model generates one continuous shot โ split it into separate clips and cut them together.
Why does my AI video barely move?
Because the prompt described a state rather than a change. Video models need an explicit action and an explicit camera instruction; without them they default to a slow drift. Name what moves and how the camera behaves, and specify the speed of both.
How do I keep the same character across multiple AI video clips?
Use a fixed block of text describing the character in specific detail and paste it unchanged into every prompt, varying only the action and camera lines. More reliably, generate one strong still and use it as the starting image for each clip, or use a frame from the previous clip as the next clip's start frame.
Why do hands and text look wrong in AI-generated video?
Both require the model to keep fine detail consistent across every frame, which is exactly what it is worst at. Frame hands wider or avoid fine manipulation in the action, and add any text in your editor afterwards rather than asking for it in frame.
Should I use text-to-video or image-to-video?
Image-to-video, whenever the tool supports it. Starting from a fixed image removes half the uncertainty โ the look is already decided โ so the prompt only has to specify motion. It also solves character consistency more reliably than any prompt wording.
Do I have to disclose that a video was AI-generated?
On most major platforms, yes, and several detect and label synthetic media automatically. Advertising regulators in a number of markets treat undisclosed synthetic content in ads as a compliance issue. Check the current rules for the specific platform, and never generate identifiable real people without consent.
Put this into practice
Generate a structured prompt or turn your workflow into a reusable Agent Skill โ both free.