Keeping a face stable for five seconds is one problem. Keeping the same character recognisable across six separately generated AI video clips is harder.
Each generation can reinterpret the face, age, proportions, hair, clothing and environment. Treat the character as a controlled production asset, not a description rewritten for every shot.
Build a video continuity pack first
Start with a small set of approved character images before generating any video. Include only views the planned sequence actually needs.
A useful pack might contain:
- front-facing head and shoulders
- three-quarter portrait from each required side
- profile view if the character turns sideways
- full-body view with clear proportions
- the exact outfit, hairstyle and accessories used in the scene
- close views of distinctive details such as glasses, tattoos or jewellery
Use neutral expressions and clean lighting. Avoid dramatic shadows, covered faces and extreme lenses.
This is the video version of a consistent AI character reference sheet. Runway’s current guide to creating longer AI videos also recommends character plates at different angles and framing to reduce subtle identity variation across generations.
Separate identity references from shot frames
An identity reference answers, “Who is this?” A starting frame answers, “What does this particular shot look like?”
Use the continuity pack to create an approved still for each shot with the correct person, outfit, pose, location, lighting and composition. The video model can then concentrate on movement.
Some current systems accept separate subject references. For example, Google’s official Veo reference-image documentation says supported Veo 3.1 workflows can use up to three images of one person, character or product to guide appearance. Availability differs by interface and model version, so check the tool you use.
When only one starting image is accepted, choose a face angle close to the intended motion. A straight-on portrait is a weak source for an immediate full-profile turn.
Plan the sequence before generating clips
Write a shot list and mark what changes between shots. A simple creator sequence could be:
- Medium shot, character speaks to camera at a desk.
- Close-up, character picks up a product.
- Side angle, character walks toward a window.
- Over-the-shoulder shot, character uses the product.
- Medium close-up, character returns to camera.
Next to each shot, record the fixed details:
- hair parting and length
- makeup and skin appearance
- top, sleeves and neckline
- product version
- room layout and time of day
- key light direction
- camera height and visual style
If the first clip uses a dark green jumper and window light from the left, do not reduce those details to “casual clothes” and “natural lighting” later.
Approve every starting still side by side
Generate the planned stills before animating any of them. Put the images in a grid and compare the face, body, hair and wardrobe across the full sequence.
Character drift often appears as:
- a narrower jaw or different cheek volume
- changing age or skin texture
- a new hairline or hair length
- different shoulder width or height
- altered glasses, earrings or clothing seams
- a face that matches closely but no longer feels like the same person
Reject weak stills now. Animation can amplify an uncertain identity, especially during head turns, speech and hand-to-face movement.
If a still is nearly correct, edit the affected area while locking the background and composition.
Keep the identity language stable
Use one short character block across all shot prompts. Do not keep adding new adjectives in an attempt to force consistency.
For example:
Same approved character, woman in her late twenties with an oval face, dark brown shoulder-length hair parted on the left, brown eyes, natural skin texture, small silver hoop earrings and a dark green crew-neck jumper.
Then add only the information required for the shot:
Medium close-up at the same desk, soft window light from camera left. She looks toward the lens and lifts the exact referenced bottle slightly. Subtle handheld phone framing.
The text block is a continuity reminder, not a replacement for visual references.
Prompt one manageable action per clip
Complex motion gives the model more chances to reinterpret the character. Split a long performance into short shots with one main action.
Avoid combining speech, a full turn, walking, a wardrobe reveal and a camera orbit in one clip.
For image-to-video, describe movement rather than restating the whole image. Runway’s current image-to-video prompting guide says the input image establishes composition, subject, lighting and style, while the prompt should focus mainly on motion.
A restrained prompt could be:
The character looks from the bottle back to the camera and smiles slightly. She raises the bottle a few centimetres. Minimal natural head movement. The handheld camera remains at the same height.
If the face changes during a single shot, use the more focused workflow for turning an AI image into video without changing the face.
Use the last frame only when the action continues
For a continuous action, extract a clean final frame from clip one and use it as the opening frame of clip two. This preserves position, framing and momentum. Runway documents this final-frame method for building longer sequences, with the shared frame trimmed during editing.
For a new camera setup, create a separately approved still from the same continuity pack. Carrying the last frame into a completely different angle can force an awkward transformation instead of a clean cut.
Choose the connection based on the edit:
- same action and camera position: continue from the last frame
- new angle or location: generate a new approved shot frame
- large time jump: use a deliberate transition or establishing shot
- uncertain face at the end: cut earlier and use a cleaner frame
Do not chain from a distorted final frame. Errors passed into the next clip often become the new baseline.
Match the filmmaking, not only the face
The same person can still look inconsistent when every shot uses different visual rules. Keep a compact style block for lens character, contrast, colour, camera movement and lighting.
A polished locked camera should not suddenly cut to aggressive handheld motion unless the story calls for it. Likewise, warm sunset light and cool overhead office light suggest different moments. If the change is intentional, show the transition.
Rasgo’s guide to AI video camera movement prompts can help define pans, push-ins, tracking shots and locked frames without overloading the action prompt.
Continuous room tone, voiceover or music can bind separate clips together. A hard sound change makes a small visual mismatch feel larger.
Review the sequence as a sequence
Watch each clip alone, then view every cut at normal speed and frame by frame. Compare the last clear face before a cut with the first clear face after it.
Check identity, age, hair, wardrobe, product, lighting, camera height and environment. Keep a simple approval grid with one row per shot. Mark a clip approved only when both the individual generation and its neighbouring cuts work.
The reliable production order is character pack, shot list, approved still grid, simple motion, cut assembly, audio, then final continuity review.
Rasgo can support the reference-led image and video stages by keeping a reusable character consistent while you develop different shot frames. Build the sequence clip by clip, but judge it as one piece. The character is only consistent when viewers believe every shot belongs to the same person and production.
Explore AI video generator for the workflow, then create a video in Rasgo.
Create your next Rasgo visual
Turn the ideas from this guide into generated images, creator content, product shots, or video-ready concepts.