390 lines
10 KiB
Markdown
390 lines
10 KiB
Markdown
# MiniMax H3 Ref2VA Prompt Guidelines
|
|
|
|
Source basis: Chris-provided `MiniMax H3 Ref2VA Reference & Scene Guidelines — V2`
|
|
|
|
## Purpose
|
|
|
|
Practical prompt and reference-structure guide for `MiniMax H3 Ref2VA`,
|
|
especially on an `8GB VRAM` workflow where the safe planning baseline is:
|
|
|
|
- `480p`
|
|
- maximum `10s` per generated shot
|
|
- one coherent shot per generation
|
|
- explicit roles for every reference image
|
|
|
|
## Core rule
|
|
|
|
Reduce the system to five responsibilities:
|
|
|
|
1. `Master scene`
|
|
Defines where the shot happens and how it initially looks.
|
|
2. `Full-body references`
|
|
Define who is present and their overall appearance.
|
|
3. `Headshot references`
|
|
Reinforce whose face it is.
|
|
4. `Specialist references`
|
|
Define clothing, props, vehicles, or other narrow details.
|
|
5. `Prompt`
|
|
Defines what happens during the next `10` seconds.
|
|
|
|
## Reference hierarchy
|
|
|
|
Use this authority order:
|
|
|
|
1. `master scene / environment`
|
|
2. `primary full-body character references`
|
|
3. `facial identity references`
|
|
4. `clothing references`
|
|
5. `props / vehicles / specialist references`
|
|
6. `text prompt`
|
|
|
|
Within a character:
|
|
|
|
- `full body` controls overall appearance, body proportions, hair, and clothing
|
|
- `headshots` refine facial identity
|
|
|
|
## Master scene rules
|
|
|
|
Prefer `<Picture 1>` as the master scene when possible.
|
|
|
|
It should control:
|
|
|
|
- environment
|
|
- composition
|
|
- camera position
|
|
- perspective
|
|
- field of view
|
|
- lighting
|
|
- initial subject placement
|
|
- important object placement
|
|
|
|
Preferred phrasing:
|
|
|
|
`<Picture 1> is the master scene, environment and composition reference.`
|
|
|
|
`Use <Picture 1> as the sole visual reference for the environment.`
|
|
|
|
`Preserve its initial camera position, framing, perspective, field of view, horizon, lighting, environment geometry, object positions, proportions and spatial layout.`
|
|
|
|
`Do not duplicate, offset, mirror, layer or reconstruct a second version of the environment.`
|
|
|
|
## Environment rules
|
|
|
|
If a reference already establishes the environment, do not re-describe every
|
|
visible object in text unless something changes.
|
|
|
|
Prefer:
|
|
|
|
`The environment, background, spatial layout and lighting remain as established by <Picture 1>.`
|
|
|
|
Over a full textual rebuild of the room.
|
|
|
|
## Character reference grouping
|
|
|
|
When multiple images represent the same person, explicitly say so.
|
|
|
|
Preferred phrasing:
|
|
|
|
`<Picture 2>, <Picture 3> and <Picture 4> all represent Subject 1.`
|
|
|
|
`Use their consistent identity information together to preserve Subject 1 throughout the video.`
|
|
|
|
`Do not describe them as separate subjects.`
|
|
|
|
## Full-body reference role
|
|
|
|
Primary use:
|
|
|
|
- overall appearance
|
|
- body proportions
|
|
- build
|
|
- hair
|
|
- clothing
|
|
- silhouette
|
|
|
|
Preferred phrasing:
|
|
|
|
`<Picture 2> is the primary full-body reference for Subject 1.`
|
|
|
|
`Use <Picture 2> to define Subject 1's overall appearance, body proportions, build, hair and clothing.`
|
|
|
|
`Ignore the background, environment, camera composition and location shown in <Picture 2>.`
|
|
|
|
## Headshot reference role
|
|
|
|
Primary use:
|
|
|
|
- facial identity
|
|
- facial structure
|
|
- eyes
|
|
- nose
|
|
- mouth
|
|
- skin appearance
|
|
|
|
Preferred phrasing:
|
|
|
|
`<Picture 3> is a facial identity reference for Subject 1.`
|
|
|
|
`Use <Picture 3> to reinforce Subject 1's face, facial structure, eyes, nose, mouth, skin appearance and other facial characteristics.`
|
|
|
|
`Do not use <Picture 3> to determine body proportions, clothing, environment or scene composition.`
|
|
|
|
## Full-body vs headshot conflict rule
|
|
|
|
Always resolve authority explicitly:
|
|
|
|
`<Picture 2> defines Subject 1's overall body, proportions, hair and clothing.`
|
|
|
|
`<Picture 3> and <Picture 4> provide additional facial identity information for the same Subject 1.`
|
|
|
|
`Use the headshot references to improve facial identity while preserving the body, clothing and overall appearance established by <Picture 2>.`
|
|
|
|
## Multi-character separation
|
|
|
|
Keep groups explicit to avoid identity leakage:
|
|
|
|
`<Picture 2> and <Picture 3> represent Subject 1.`
|
|
|
|
`<Picture 4> and <Picture 5> represent Subject 2.`
|
|
|
|
`Do not allow the identity or physical characteristics of one subject to influence another subject.`
|
|
|
|
## Background isolation rule
|
|
|
|
Character references often contain irrelevant backgrounds. State that those
|
|
backgrounds should be ignored.
|
|
|
|
Preferred phrasing:
|
|
|
|
`Use all character reference images only for their assigned character information.`
|
|
|
|
`Ignore their original backgrounds, locations, scene compositions, camera positions and environmental lighting.`
|
|
|
|
`The environment is defined exclusively by <Picture 1>.`
|
|
|
|
## Specialist references
|
|
|
|
Narrow responsibility prevents conflicts.
|
|
|
|
Example clothing phrasing:
|
|
|
|
`<Picture 5> defines the clothing worn by Subject 1.`
|
|
|
|
`Use <Picture 5> only as a clothing reference.`
|
|
|
|
`Ignore the person, face, body, pose, environment and background shown in <Picture 5>.`
|
|
|
|
Example vehicle phrasing:
|
|
|
|
`<Picture 6> defines the appearance of the vehicle.`
|
|
|
|
`Use <Picture 6> only as a visual reference for the vehicle.`
|
|
|
|
`Ignore its original environment, camera position and background.`
|
|
|
|
## 8GB VRAM constraints
|
|
|
|
Assume this baseline unless direct testing proves otherwise:
|
|
|
|
- resolution: `480p`
|
|
- maximum duration: `10s`
|
|
- one coherent action beat per shot
|
|
|
|
Do not design a single generation as a long sequence of unrelated events.
|
|
|
|
## Action design for 10 seconds
|
|
|
|
Good:
|
|
|
|
- `Subject 1 walks towards Subject 2.`
|
|
- `Subject 2 turns towards Subject 1.`
|
|
- `They briefly look at each other.`
|
|
- `The camera slowly tracks forward.`
|
|
|
|
Bad:
|
|
|
|
- entering a room
|
|
- crossing the room
|
|
- talking
|
|
- sitting
|
|
- picking something up
|
|
- finishing a conversation
|
|
- standing
|
|
- leaving
|
|
|
|
all in one `10s` shot
|
|
|
|
## Timing guidance
|
|
|
|
Simple timing is useful when sequencing matters:
|
|
|
|
- `0-3 seconds`: approach
|
|
- `3-7 seconds`: reaction
|
|
- `7-10 seconds`: settle / hold
|
|
|
|
Avoid over-micro-timing unless necessary.
|
|
|
|
## Dialogue guidance
|
|
|
|
Dialogue must fit within the same `10s` budget as motion and camera changes.
|
|
|
|
If the line is long, keep action simple.
|
|
|
|
Preferred pattern:
|
|
|
|
- character holds position or makes one simple move
|
|
- one spoken line
|
|
- one visible reaction
|
|
- one restrained camera movement
|
|
|
|
## Camera guidance
|
|
|
|
The master scene defines the opening frame.
|
|
|
|
Prompt text should describe how that camera changes, not replace it with a new
|
|
simultaneous framing.
|
|
|
|
Good camera phrasing:
|
|
|
|
- `The camera slowly pushes forward.`
|
|
- `The camera gently pans right to follow Subject 1.`
|
|
- `The camera slowly arcs around both subjects while maintaining their spatial relationship.`
|
|
|
|
## Expression guidance
|
|
|
|
Headshots lock identity, not frozen expression.
|
|
|
|
Expression belongs in performance instructions:
|
|
|
|
`Subject 1 initially appears relaxed.`
|
|
|
|
`As Subject 2 speaks, their expression gradually becomes concerned.`
|
|
|
|
## Recommended prompt architecture
|
|
|
|
Use this section order:
|
|
|
|
1. `Reference definitions`
|
|
2. `Scene anchor`
|
|
3. `Reference restrictions`
|
|
4. `Subject placement`
|
|
5. `Action / performance`
|
|
6. `Camera`
|
|
7. `Dialogue / audio`
|
|
|
|
## Single-character template
|
|
|
|
```text
|
|
<Picture 1> is the master scene, environment and composition reference.
|
|
|
|
<Picture 2>, <Picture 3> and <Picture 4> all represent Subject 1.
|
|
|
|
<Picture 2> is Subject 1's primary full-body reference and defines
|
|
their overall appearance, body proportions, build, hair and clothing.
|
|
|
|
<Picture 3> and <Picture 4> are complementary facial identity
|
|
references for Subject 1.
|
|
|
|
Use these headshot references to reinforce Subject 1's facial identity
|
|
while preserving the overall appearance established by <Picture 2>.
|
|
|
|
Use <Picture 1> as the sole visual reference for the environment.
|
|
|
|
Preserve its initial composition, camera position, perspective,
|
|
field of view, lighting, environment geometry and spatial layout.
|
|
|
|
Ignore the backgrounds, environments and compositions shown in
|
|
<Picture 2>, <Picture 3> and <Picture 4>.
|
|
|
|
[ACTION — MAXIMUM 10 SECONDS]
|
|
|
|
[CAMERA]
|
|
|
|
[DIALOGUE]
|
|
|
|
[AUDIO]
|
|
```
|
|
|
|
## Multi-character template
|
|
|
|
```text
|
|
<Picture 1> is the master scene, environment and composition reference.
|
|
|
|
<Picture 2>, <Picture 3> and <Picture 4> all represent Subject 1.
|
|
|
|
<Picture 2> is the primary full-body reference for Subject 1
|
|
and defines their overall appearance, body proportions, build,
|
|
hair and clothing.
|
|
|
|
<Picture 3> and <Picture 4> are complementary facial identity
|
|
references for Subject 1 and provide additional information about
|
|
their facial structure and appearance.
|
|
|
|
<Picture 5> and <Picture 6> both represent Subject 2.
|
|
|
|
<Picture 5> is the primary full-body reference for Subject 2
|
|
and defines their overall appearance, body proportions, build,
|
|
hair and clothing.
|
|
|
|
<Picture 6> is a facial identity reference for Subject 2.
|
|
|
|
The opening frame uses the environment and composition established
|
|
by <Picture 1>.
|
|
|
|
Use <Picture 1> as the sole visual reference for the environment.
|
|
|
|
Preserve its initial camera position, framing, perspective,
|
|
field of view, horizon, lighting, environment geometry,
|
|
object positions, proportions and spatial layout.
|
|
|
|
Do not duplicate, offset, mirror, layer or reconstruct a second
|
|
version of the environment.
|
|
|
|
Use the full-body character references to establish the overall
|
|
appearance of their respective subjects.
|
|
|
|
Use the headshot references to reinforce facial identity while
|
|
preserving the body and overall appearance established by the
|
|
corresponding full-body reference.
|
|
|
|
Ignore the backgrounds, environments, camera positions and scene
|
|
compositions shown in all character reference images.
|
|
|
|
Do not allow the identity or physical characteristics of one
|
|
subject to influence another subject.
|
|
|
|
[ACTION AND PERFORMANCE — MAXIMUM 10-SECOND SHOT]
|
|
|
|
[CAMERA MOVEMENT]
|
|
|
|
[DIALOGUE]
|
|
|
|
[AUDIO / ENVIRONMENTAL SOUND]
|
|
```
|
|
|
|
## Efficiency rule
|
|
|
|
Do not add references just because they exist.
|
|
|
|
Good default set for one important character:
|
|
|
|
- `1 x` full-body reference
|
|
- `1 x` frontal headshot
|
|
- `1 x` three-quarter headshot
|
|
|
|
The goal is the smallest reference set that clearly defines the required
|
|
information.
|
|
|
|
## Practical default checklist
|
|
|
|
- work at `480p`
|
|
- treat `10s` as the safe maximum shot duration
|
|
- use one coherent action or interaction per generation
|
|
- use a master scene whenever possible
|
|
- give the master scene authority over environment and composition
|
|
- give each important character one primary full-body reference
|
|
- add headshots only when facial identity needs reinforcement
|
|
- explicitly state which images belong to the same character
|
|
- stop character-reference backgrounds from influencing the master scene
|
|
- use the text prompt mainly for action, performance, camera, dialogue, and audio
|