Files
comfyui-workflows/optimized/notes/minimax-h3-ref2va-prompt-guidelines.md

390 lines
10 KiB
Markdown

# MiniMax H3 Ref2VA Prompt Guidelines
Source basis: Chris-provided `MiniMax H3 Ref2VA Reference & Scene Guidelines — V2`
## Purpose
Practical prompt and reference-structure guide for `MiniMax H3 Ref2VA`,
especially on an `8GB VRAM` workflow where the safe planning baseline is:
- `480p`
- maximum `10s` per generated shot
- one coherent shot per generation
- explicit roles for every reference image
## Core rule
Reduce the system to five responsibilities:
1. `Master scene`
Defines where the shot happens and how it initially looks.
2. `Full-body references`
Define who is present and their overall appearance.
3. `Headshot references`
Reinforce whose face it is.
4. `Specialist references`
Define clothing, props, vehicles, or other narrow details.
5. `Prompt`
Defines what happens during the next `10` seconds.
## Reference hierarchy
Use this authority order:
1. `master scene / environment`
2. `primary full-body character references`
3. `facial identity references`
4. `clothing references`
5. `props / vehicles / specialist references`
6. `text prompt`
Within a character:
- `full body` controls overall appearance, body proportions, hair, and clothing
- `headshots` refine facial identity
## Master scene rules
Prefer `<Picture 1>` as the master scene when possible.
It should control:
- environment
- composition
- camera position
- perspective
- field of view
- lighting
- initial subject placement
- important object placement
Preferred phrasing:
`<Picture 1> is the master scene, environment and composition reference.`
`Use <Picture 1> as the sole visual reference for the environment.`
`Preserve its initial camera position, framing, perspective, field of view, horizon, lighting, environment geometry, object positions, proportions and spatial layout.`
`Do not duplicate, offset, mirror, layer or reconstruct a second version of the environment.`
## Environment rules
If a reference already establishes the environment, do not re-describe every
visible object in text unless something changes.
Prefer:
`The environment, background, spatial layout and lighting remain as established by <Picture 1>.`
Over a full textual rebuild of the room.
## Character reference grouping
When multiple images represent the same person, explicitly say so.
Preferred phrasing:
`<Picture 2>, <Picture 3> and <Picture 4> all represent Subject 1.`
`Use their consistent identity information together to preserve Subject 1 throughout the video.`
`Do not describe them as separate subjects.`
## Full-body reference role
Primary use:
- overall appearance
- body proportions
- build
- hair
- clothing
- silhouette
Preferred phrasing:
`<Picture 2> is the primary full-body reference for Subject 1.`
`Use <Picture 2> to define Subject 1's overall appearance, body proportions, build, hair and clothing.`
`Ignore the background, environment, camera composition and location shown in <Picture 2>.`
## Headshot reference role
Primary use:
- facial identity
- facial structure
- eyes
- nose
- mouth
- skin appearance
Preferred phrasing:
`<Picture 3> is a facial identity reference for Subject 1.`
`Use <Picture 3> to reinforce Subject 1's face, facial structure, eyes, nose, mouth, skin appearance and other facial characteristics.`
`Do not use <Picture 3> to determine body proportions, clothing, environment or scene composition.`
## Full-body vs headshot conflict rule
Always resolve authority explicitly:
`<Picture 2> defines Subject 1's overall body, proportions, hair and clothing.`
`<Picture 3> and <Picture 4> provide additional facial identity information for the same Subject 1.`
`Use the headshot references to improve facial identity while preserving the body, clothing and overall appearance established by <Picture 2>.`
## Multi-character separation
Keep groups explicit to avoid identity leakage:
`<Picture 2> and <Picture 3> represent Subject 1.`
`<Picture 4> and <Picture 5> represent Subject 2.`
`Do not allow the identity or physical characteristics of one subject to influence another subject.`
## Background isolation rule
Character references often contain irrelevant backgrounds. State that those
backgrounds should be ignored.
Preferred phrasing:
`Use all character reference images only for their assigned character information.`
`Ignore their original backgrounds, locations, scene compositions, camera positions and environmental lighting.`
`The environment is defined exclusively by <Picture 1>.`
## Specialist references
Narrow responsibility prevents conflicts.
Example clothing phrasing:
`<Picture 5> defines the clothing worn by Subject 1.`
`Use <Picture 5> only as a clothing reference.`
`Ignore the person, face, body, pose, environment and background shown in <Picture 5>.`
Example vehicle phrasing:
`<Picture 6> defines the appearance of the vehicle.`
`Use <Picture 6> only as a visual reference for the vehicle.`
`Ignore its original environment, camera position and background.`
## 8GB VRAM constraints
Assume this baseline unless direct testing proves otherwise:
- resolution: `480p`
- maximum duration: `10s`
- one coherent action beat per shot
Do not design a single generation as a long sequence of unrelated events.
## Action design for 10 seconds
Good:
- `Subject 1 walks towards Subject 2.`
- `Subject 2 turns towards Subject 1.`
- `They briefly look at each other.`
- `The camera slowly tracks forward.`
Bad:
- entering a room
- crossing the room
- talking
- sitting
- picking something up
- finishing a conversation
- standing
- leaving
all in one `10s` shot
## Timing guidance
Simple timing is useful when sequencing matters:
- `0-3 seconds`: approach
- `3-7 seconds`: reaction
- `7-10 seconds`: settle / hold
Avoid over-micro-timing unless necessary.
## Dialogue guidance
Dialogue must fit within the same `10s` budget as motion and camera changes.
If the line is long, keep action simple.
Preferred pattern:
- character holds position or makes one simple move
- one spoken line
- one visible reaction
- one restrained camera movement
## Camera guidance
The master scene defines the opening frame.
Prompt text should describe how that camera changes, not replace it with a new
simultaneous framing.
Good camera phrasing:
- `The camera slowly pushes forward.`
- `The camera gently pans right to follow Subject 1.`
- `The camera slowly arcs around both subjects while maintaining their spatial relationship.`
## Expression guidance
Headshots lock identity, not frozen expression.
Expression belongs in performance instructions:
`Subject 1 initially appears relaxed.`
`As Subject 2 speaks, their expression gradually becomes concerned.`
## Recommended prompt architecture
Use this section order:
1. `Reference definitions`
2. `Scene anchor`
3. `Reference restrictions`
4. `Subject placement`
5. `Action / performance`
6. `Camera`
7. `Dialogue / audio`
## Single-character template
```text
<Picture 1> is the master scene, environment and composition reference.
<Picture 2>, <Picture 3> and <Picture 4> all represent Subject 1.
<Picture 2> is Subject 1's primary full-body reference and defines
their overall appearance, body proportions, build, hair and clothing.
<Picture 3> and <Picture 4> are complementary facial identity
references for Subject 1.
Use these headshot references to reinforce Subject 1's facial identity
while preserving the overall appearance established by <Picture 2>.
Use <Picture 1> as the sole visual reference for the environment.
Preserve its initial composition, camera position, perspective,
field of view, lighting, environment geometry and spatial layout.
Ignore the backgrounds, environments and compositions shown in
<Picture 2>, <Picture 3> and <Picture 4>.
[ACTION — MAXIMUM 10 SECONDS]
[CAMERA]
[DIALOGUE]
[AUDIO]
```
## Multi-character template
```text
<Picture 1> is the master scene, environment and composition reference.
<Picture 2>, <Picture 3> and <Picture 4> all represent Subject 1.
<Picture 2> is the primary full-body reference for Subject 1
and defines their overall appearance, body proportions, build,
hair and clothing.
<Picture 3> and <Picture 4> are complementary facial identity
references for Subject 1 and provide additional information about
their facial structure and appearance.
<Picture 5> and <Picture 6> both represent Subject 2.
<Picture 5> is the primary full-body reference for Subject 2
and defines their overall appearance, body proportions, build,
hair and clothing.
<Picture 6> is a facial identity reference for Subject 2.
The opening frame uses the environment and composition established
by <Picture 1>.
Use <Picture 1> as the sole visual reference for the environment.
Preserve its initial camera position, framing, perspective,
field of view, horizon, lighting, environment geometry,
object positions, proportions and spatial layout.
Do not duplicate, offset, mirror, layer or reconstruct a second
version of the environment.
Use the full-body character references to establish the overall
appearance of their respective subjects.
Use the headshot references to reinforce facial identity while
preserving the body and overall appearance established by the
corresponding full-body reference.
Ignore the backgrounds, environments, camera positions and scene
compositions shown in all character reference images.
Do not allow the identity or physical characteristics of one
subject to influence another subject.
[ACTION AND PERFORMANCE — MAXIMUM 10-SECOND SHOT]
[CAMERA MOVEMENT]
[DIALOGUE]
[AUDIO / ENVIRONMENTAL SOUND]
```
## Efficiency rule
Do not add references just because they exist.
Good default set for one important character:
- `1 x` full-body reference
- `1 x` frontal headshot
- `1 x` three-quarter headshot
The goal is the smallest reference set that clearly defines the required
information.
## Practical default checklist
- work at `480p`
- treat `10s` as the safe maximum shot duration
- use one coherent action or interaction per generation
- use a master scene whenever possible
- give the master scene authority over environment and composition
- give each important character one primary full-body reference
- add headshots only when facial identity needs reinforcement
- explicitly state which images belong to the same character
- stop character-reference backgrounds from influencing the master scene
- use the text prompt mainly for action, performance, camera, dialogue, and audio