H3 Max
Multimodal Video Tool

Reference to Video

Keep the same character when the scene changes

Upload one to three references and the subject in them carries into the shot: face, clothing, proportions. Then say where they are and what happens. The references handle who it is, so the prompt does not have to.

Creative workspace

Reference to Video

50

What is Reference to Video?

Reference to video is generating a clip from example images rather than from a prompt alone. You supply reference pictures of a character, a product, or a look, and every shot the model makes keeps them recognisable.

How the people who built the model describe it:

“fal's H3 Max is a post-trained variant of MiniMax H3, tuned for stronger prompt adherence and better aesthetics while co-optimized with our custom inference stack for higher throughput with no compromises on output quality”
MiniMax H3 Max on fal

Reference to Video at a glance

You bring
1-3 references
They fix
Who and what it looks like
The prompt fixes
Where and what happens

What you can make with the Reference to Video tool

These are the jobs people actually bring to Reference to Video, and the settings each one needs.

01

One character, many scenes

A host, mascot, or lead who has to be recognisably the same person in every clip of a series.

02

One product, many settings

The same bottle, shoe, or device in a kitchen, on a street, in a studio, without rebuilding the shot each time.

03

Several stills, one shot

Take the subject from one frame and the place from another, and let the motion between them be generated.

Reference to Video

A workspace shaped for the model

Reference to Video puts the handful of controls that change the result most in front of you, and keeps the rest out of the way.

01

More angles, better likeness

Two or three references of the same subject beat one. A single frame shows the model one side of a face, and it invents the rest. The extra views are what stop it.

02

Describe the scene, not the subject

The references already carry identity, so the whole prompt can go on place, action, camera and light.

03

Not the same as image to video

Image to video animates the picture you gave it, framing and all. This takes only the subject out of it and puts that somewhere else.

04

Sound arrives with the shot

Ambience, effects, and dialogue are generated alongside the picture rather than added to it afterwards.

05

Close, not identical

Likeness is strong but not guaranteed. It drifts first on faces far from camera, unusual angles, and fine detail like text on a logo. Plan on a second run rather than a surprise at delivery.

Reference to Video

Choose with context

Where Reference to Video is the right call, and where a different model does the job better.

Starts from

A subject you need to keep

The prompt decides

Everything except who

Pick this when

The same face or product has to appear again

How to keep a subject consistent across shots

01

Upload one to three references

Same subject, different angles, clearly lit. Three views of one face work better than three different faces.

02

Write the scene around them

Place, action, camera move, sound. Leave out what they look like; that is what the references are for.

03

Run it, then run it again

Likeness varies between runs. Keeping the best of two or three costs less than trying to describe your way there.

04

Reuse the same references

For a series, keep the reference set fixed and change only the prompt. That is what makes episode four look like episode one.

Questions about reference to video

Start Creating With Reference to Video

Set up the brief now, then continue in the shared Studio workspace.