Say you are trying to put together a 15-second product teaser for tomorrow’s social post, and you’ve got Runway or Pika open, typing something like “a sleek phone rotating on a table with dramatic lighting.” What comes back is a phone that warps halfway through the rotation, lighting that shifts for no reason, and a camera move that drifts somewhere you didn’t ask it to go. You reword it, run it again, get a different but equally unusable version, and now forty minutes are gone.

That’s the pattern that pushed me to actually treat video prompts as their own discipline instead of recycled image prompts with “video of” tacked on the front. Video models are juggling motion, time, and physics on top of everything a still image already has to get right, and a prompt that doesn’t account for that gets punished fast. Below is the split I’d have wanted on day one: what works if you’re just starting out, and what changes once you’re trying to produce something you’d put in front of a client.


Beginner: Describe a scene, not a story

If you’re new to Sora, Runway, or Pika, the first mistake is asking for too much narrative in one clip. These tools are not generating a mini-movie — they’re extending a single visual moment across a few seconds. A prompt like “a woman walks into a coffee shop, orders a latte, sits down, and starts reading” is five scenes’ worth of action stuffed into a clip that’s maybe 4-10 seconds long. The model has to compress all of that, and compression here means garbled transitions and objects that don’t behave consistently frame to frame.

What works instead: pick one moment and describe it fully. “A woman sits at a window table in a coffee shop, steam rising from a ceramic mug in front of her, soft afternoon light coming through the glass.” One scene, one camera relationship to that scene, and the motion the model needs to invent is limited to steam rising and maybe a subtle light shift — both things these models render reasonably well.

A second beginner habit worth building early: always specify the camera, even in one clause. “Static shot,” “slow pan left,” “handheld tracking shot following the subject” — video models default to an ambiguous, often wandering camera when you don’t pin it down, and an unstable camera is the single fastest way to make a clip look unusable regardless of how good the subject looks.

Third: name the shot length or pacing in plain terms. You can’t control exact seconds precisely across all three tools, but words like “slow motion,” “quick cut,” or “lingering shot” do measurably shift how the model paces the action within the clip it generates.

Here’s a before-and-after from my own testing on a product shot:

  • Before: “A watch on a marble surface with cinematic lighting”
  • After: “Static overhead shot of a silver wristwatch on white marble, soft studio light from the left casting a gentle shadow, no camera movement, no other objects in frame”

The second version came out consistent across four regenerations. The first one gave me a different lighting setup, a different watch angle, and once, an extra hand reaching into frame that nobody asked for.


Beginner: Keep negative space explicit

Text-to-image users are used to negative prompts as an optional refinement. In video, leaving out an exclusion is riskier, because the model has more surface area — more frames, more time — in which to introduce something you didn’t want. If you don’t say “no text overlays,” you might get stray, garbled text baked into the clip. If you don’t say “no people in the background,” a random extra might drift into frame around second three.

State exclusions the same way you’d state them in a written spec: on their own, plainly, not folded into a longer sentence about lighting and mood. It costs you five words and saves you a regeneration.


Advanced: Control motion as its own parameter, separate from the scene

Once you’ve got clean, single-scene prompts working, the next level is treating motion as something you author explicitly rather than something you imply through the scene description. This is where beginner and advanced prompts diverge the most, and where Sora, Runway, and Pika each reward slightly different phrasing.

Instead of writing “a car driving down a mountain road,” which leaves speed, camera relationship, and trajectory all undefined, an advanced prompt separates the static description from the motion instruction: “A red convertible on a winding mountain road, shot from a drone tracking alongside the car at a steady distance, moderate speed, camera height level with the car’s roofline.” Every element of motion — subject speed, camera type, camera-to-subject relationship — is now its own explicit clause instead of something inferred from a single adjective like “driving.”

In testing across Runway and Pika specifically, I’ve found that pairing a camera motion clause with a subject motion clause, kept grammatically separate, produces noticeably more stable results than blending them into one description. Something like: “Subject motion: the dancer spins slowly in place. Camera motion: slow dolly-in, ending in a close-up on her face.” Neither tool requires that literal labeled format to work, but structuring your sentence in that order — subject first, camera second — consistently outperforms the reverse or a merged version in side-by-side runs.


Advanced: Use reference frames and iterative extension instead of one long prompt

This is the biggest practical difference between beginner and advanced use, and it’s less about wording and more about workflow. Beginners tend to write one prompt and hope for a full clip. People producing usable output consistently break longer sequences into chained segments, using the last frame of one generation as the starting reference for the next.

If Runway or Pika supports image-to-video from a still frame, export the final frame of your best take and feed it back in as the starting point for the next clip, with a new motion prompt describing only what happens from that point forward. This keeps visual continuity — same subject, same lighting, same framing — while letting you author motion in short, controllable increments rather than gambling on one long generation nailing everything at once.

Sora’s longer native clip length changes the math slightly: you can ask for more contained motion within a single generation before you need to chain segments. But the same principle holds — the longer and more eventful your ask within one prompt, the more the output degrades toward incoherence. Chaining shorter, well-specified clips and stitching them in post consistently beats one ambitious single-shot prompt, even when the tool technically supports longer durations.


Advanced: Lock consistency with a repeated description block

If you’re generating multiple clips meant to sit together — a product from three angles, a character across three scenes — inconsistency across generations is the most common thing that kills a project at the editing stage. The face changes slightly, the outfit color shifts, the lighting mood drifts from clip to clip.

The fix that has held up in my own workflow: write a fixed description block for the subject — appearance, color, material, defining details — and paste that identical block into every prompt for that project, changing only the scene and motion clauses around it. It’s repetitive to write, and it feels redundant the third time you paste it, but it is the difference between three clips that look like they belong in the same video and three clips that look like they came from three different projects.


Quick Comparison: Beginner vs. Advanced Habits

AspectBeginner ApproachAdvanced Approach
ScopeOne full narrative in one promptOne scene, one moment, per generation
MotionImplied through a single verbExplicit, separate clause from the scene description
CameraLeft unspecified or vagueNamed explicitly every time (static, pan, dolly, tracking)
ContinuityEach generation treated independentlyReference frames or repeated description blocks carried across clips
LengthOne long ambitious clip attemptedShort segments chained together in post
ExclusionsAssumed, not statedStated plainly and separately

None of this requires knowing how diffusion or transformer-based video generation works under the hood. It requires treating motion, camera, and continuity as parameters you set on purpose instead of details you hope the model guesses correctly. Start with the beginner habits if a clip is coming out warped or wandering. Move to the advanced ones once your single clips are clean and the problem shifts to making several of them look like they belong together.

If you’ve been fighting a specific clip that keeps drifting or glitching partway through, try isolating whether it’s a scene problem, a motion problem, or a continuity problem first — the fix looks different for each one, and guessing at all three at once is usually what burns the most time.