AI video often looks convincing in a still frame but strange as soon as something moves. A person glides instead of walking, hair behaves like it is underwater, a product floats into place, or the entire shot plays with an accidental slow-motion feel.
The fix is rarely adding more cinematic language. Natural-looking motion comes from giving the model a simple physical event, a clear pace and enough environmental reaction to connect the movement to the scene.
Why AI video motion looks fake
A video model has to invent the frames after the starting image. Problems appear when the prompt leaves too much open to interpretation or asks several difficult events to happen at once.
Common causes include:
- A vague action such as "moves naturally" with no direction or pace
- Too many actions, camera moves and scene changes in one clip
- A starting pose that makes the requested movement awkward
- Motion with no effect on clothing, hair, shadows or nearby objects
- A long generation that gives errors time to accumulate
Runway's official image-to-video prompting guide recommends focusing the text prompt on motion rather than repeating the appearance already defined by the source image. It separates motion into subject action, environmental motion, camera motion, timing, direction and speed.
Use a simple motion prompt formula
Start with this structure:
The subject [performs one action] at [pace or timing]. [A visible part of the scene] responds to the movement. [Camera behaviour].
For example:
The woman takes three relaxed steps toward the window at normal walking speed. Her arms swing lightly and the hem of her coat moves with each step. The camera remains locked.
This is stronger than "the woman walks naturally in a cinematic scene" because it gives the model observable instructions. It names the distance, pace, secondary movement and camera state.
For a product shot:
The bottle rotates one quarter turn clockwise on the wooden surface, then settles. Its contact shadow moves with the base and the glass catches a narrow highlight. Locked camera, real-time motion.
The words "settles," "contact shadow" and "real-time motion" help communicate weight. They do more useful work than broad descriptions such as "premium movement."
Start from a frame that supports the action
The first image is the physical starting position of the shot.
If you want a person to walk, use a frame with visible floor and space in the direction of travel. A tightly cropped portrait does not provide the legs, ground contact or destination needed for a believable walk. If you want a hand to pick up a product, the hand should begin near the object in a plausible position.
Check that required hands, feet and joints are clear. Leave room for movement, keep scale and perspective consistent, and include a visible surface when contact shadows matter. A weak hand, distorted product or blurred edge in the first frame often becomes more obvious after animation.
Ask for one physical event per clip
A short clip does not need a complete story. It needs one successful shot.
"The creator stands, walks to the table, picks up the serum, applies it and smiles at the camera" contains several changes in pose, hand contact and gaze. Generate those as separate clips:
- The creator walks to the table.
- A close-up shows the hand lifting the serum.
- The creator applies one drop.
- The creator looks at the camera and smiles.
This gives you cleaner edit points and makes failed generations easier to replace. It also helps protect identity and product details. If facial consistency is the main problem, use the conservative workflow in turning an AI image into video without changing the face.
Define pace with observable language
Words such as "smoothly," "dynamically" and "naturally" are subjective. Pair them with a physical cue:
- At normal walking speed
- In one quick, continuous motion
- Slowly over the full shot
- Pauses briefly, then turns
- Takes two steps and comes to a stop
- Real-time motion, not stylised slow motion
If a clip keeps looking slow, reduce the amount of action before making the speed language more aggressive. A model may stretch a complex event across the available duration because it cannot complete every instruction cleanly.
Add secondary motion without overloading the shot
Real movement affects more than the main subject. A foot compresses against the ground. A sleeve follows an arm. Liquid changes shape when a bottle tilts. A moving object changes its shadow and reflection.
Add one or two reactions that matter:
The runner starts forward at a steady jogging pace. Her ponytail and loose shirt respond lightly to each stride. Small splashes form where her shoes meet the wet pavement.
Environmental motion can make the action feel grounded, but too many independent details compete for attention. Do not ask for rain, blowing leaves, moving traffic, dramatic fabric, several gestures and a camera orbit unless all are essential.
Runway's Gen-4 prompting guidance suggests beginning with the essential motion and adding one element at a time. It also recommends positive descriptions. "Locked camera" is clearer than repeatedly telling the model not to move.
Test subject motion before camera motion
When both the person and camera move, it becomes harder to see which instruction caused a failure.
First generate the action with a locked camera. Once the body movement or product interaction works, test a second version with a gentle push-in or tracking shot. Our guide to AI video camera movement prompts explains how to specify direction, speed and target without stacking conflicting moves.
A moving camera can improve a finished shot, but it should not disguise sliding feet, broken contact or unstable anatomy.
Prompt examples for common shots
Walking creator
Full-body creator takes three confident steps toward the camera at normal walking speed. Her feet make firm contact with the pavement, her arms swing gently and her jacket responds lightly to each step. Locked phone camera, real-time movement.
Person lifting a product
The seated creator reaches with her right hand, grips the bottle around its centre and lifts it a few centimetres from the table. The bottle keeps its shape and label orientation. Her wrist movement is small and controlled. Locked close-up.
Product rotation
The shoe rotates slowly from a side view to a three-quarter view on the platform, then stops. The sole stays in contact with the surface. The shadow and highlight move consistently with the shoe. Locked studio camera.
Each prompt limits the action, identifies the visible evidence of motion and controls the camera.
Troubleshoot the symptom, not the whole prompt
If the clip looks like slow motion, request less action and name a real-time pace. If feet slide, simplify the path, show the ground clearly and ask for firm foot contact. If arms bend unnaturally, reduce the gesture range and use a source frame with unobstructed joints.
For floating products, state the contact surface, describe how the shadow moves and include a settling point. If the background is frozen, add one relevant environmental reaction. If everything moves too much, remove secondary motion and rebuild one layer at a time.
Do not rewrite every part after one bad result. Change one variable, generate a short test and compare it with the previous version.
A reliable production order
Choose a clean starting frame, define one action, set its pace, add one physical reaction and keep the camera still for the first test. Approve the motion before adding camera movement, audio or a second event.
Rasgo can help creators build reference-led images and short AI video shots, but the same principle applies across tools: a controlled shot is easier to generate, diagnose and edit than an overloaded scene. Natural motion starts with a small event the viewer can physically believe.
Create your next Rasgo visual
Turn the ideas from this guide into generated images, creator content, product shots, or video-ready concepts.