Add MiniMax H3 Ref2VA prompt guide
This commit is contained in:
@@ -0,0 +1,389 @@
|
|||||||
|
# MiniMax H3 Ref2VA Prompt Guidelines
|
||||||
|
|
||||||
|
Source basis: Chris-provided `MiniMax H3 Ref2VA Reference & Scene Guidelines — V2`
|
||||||
|
|
||||||
|
## Purpose
|
||||||
|
|
||||||
|
Practical prompt and reference-structure guide for `MiniMax H3 Ref2VA`,
|
||||||
|
especially on an `8GB VRAM` workflow where the safe planning baseline is:
|
||||||
|
|
||||||
|
- `480p`
|
||||||
|
- maximum `10s` per generated shot
|
||||||
|
- one coherent shot per generation
|
||||||
|
- explicit roles for every reference image
|
||||||
|
|
||||||
|
## Core rule
|
||||||
|
|
||||||
|
Reduce the system to five responsibilities:
|
||||||
|
|
||||||
|
1. `Master scene`
|
||||||
|
Defines where the shot happens and how it initially looks.
|
||||||
|
2. `Full-body references`
|
||||||
|
Define who is present and their overall appearance.
|
||||||
|
3. `Headshot references`
|
||||||
|
Reinforce whose face it is.
|
||||||
|
4. `Specialist references`
|
||||||
|
Define clothing, props, vehicles, or other narrow details.
|
||||||
|
5. `Prompt`
|
||||||
|
Defines what happens during the next `10` seconds.
|
||||||
|
|
||||||
|
## Reference hierarchy
|
||||||
|
|
||||||
|
Use this authority order:
|
||||||
|
|
||||||
|
1. `master scene / environment`
|
||||||
|
2. `primary full-body character references`
|
||||||
|
3. `facial identity references`
|
||||||
|
4. `clothing references`
|
||||||
|
5. `props / vehicles / specialist references`
|
||||||
|
6. `text prompt`
|
||||||
|
|
||||||
|
Within a character:
|
||||||
|
|
||||||
|
- `full body` controls overall appearance, body proportions, hair, and clothing
|
||||||
|
- `headshots` refine facial identity
|
||||||
|
|
||||||
|
## Master scene rules
|
||||||
|
|
||||||
|
Prefer `<Picture 1>` as the master scene when possible.
|
||||||
|
|
||||||
|
It should control:
|
||||||
|
|
||||||
|
- environment
|
||||||
|
- composition
|
||||||
|
- camera position
|
||||||
|
- perspective
|
||||||
|
- field of view
|
||||||
|
- lighting
|
||||||
|
- initial subject placement
|
||||||
|
- important object placement
|
||||||
|
|
||||||
|
Preferred phrasing:
|
||||||
|
|
||||||
|
`<Picture 1> is the master scene, environment and composition reference.`
|
||||||
|
|
||||||
|
`Use <Picture 1> as the sole visual reference for the environment.`
|
||||||
|
|
||||||
|
`Preserve its initial camera position, framing, perspective, field of view, horizon, lighting, environment geometry, object positions, proportions and spatial layout.`
|
||||||
|
|
||||||
|
`Do not duplicate, offset, mirror, layer or reconstruct a second version of the environment.`
|
||||||
|
|
||||||
|
## Environment rules
|
||||||
|
|
||||||
|
If a reference already establishes the environment, do not re-describe every
|
||||||
|
visible object in text unless something changes.
|
||||||
|
|
||||||
|
Prefer:
|
||||||
|
|
||||||
|
`The environment, background, spatial layout and lighting remain as established by <Picture 1>.`
|
||||||
|
|
||||||
|
Over a full textual rebuild of the room.
|
||||||
|
|
||||||
|
## Character reference grouping
|
||||||
|
|
||||||
|
When multiple images represent the same person, explicitly say so.
|
||||||
|
|
||||||
|
Preferred phrasing:
|
||||||
|
|
||||||
|
`<Picture 2>, <Picture 3> and <Picture 4> all represent Subject 1.`
|
||||||
|
|
||||||
|
`Use their consistent identity information together to preserve Subject 1 throughout the video.`
|
||||||
|
|
||||||
|
`Do not describe them as separate subjects.`
|
||||||
|
|
||||||
|
## Full-body reference role
|
||||||
|
|
||||||
|
Primary use:
|
||||||
|
|
||||||
|
- overall appearance
|
||||||
|
- body proportions
|
||||||
|
- build
|
||||||
|
- hair
|
||||||
|
- clothing
|
||||||
|
- silhouette
|
||||||
|
|
||||||
|
Preferred phrasing:
|
||||||
|
|
||||||
|
`<Picture 2> is the primary full-body reference for Subject 1.`
|
||||||
|
|
||||||
|
`Use <Picture 2> to define Subject 1's overall appearance, body proportions, build, hair and clothing.`
|
||||||
|
|
||||||
|
`Ignore the background, environment, camera composition and location shown in <Picture 2>.`
|
||||||
|
|
||||||
|
## Headshot reference role
|
||||||
|
|
||||||
|
Primary use:
|
||||||
|
|
||||||
|
- facial identity
|
||||||
|
- facial structure
|
||||||
|
- eyes
|
||||||
|
- nose
|
||||||
|
- mouth
|
||||||
|
- skin appearance
|
||||||
|
|
||||||
|
Preferred phrasing:
|
||||||
|
|
||||||
|
`<Picture 3> is a facial identity reference for Subject 1.`
|
||||||
|
|
||||||
|
`Use <Picture 3> to reinforce Subject 1's face, facial structure, eyes, nose, mouth, skin appearance and other facial characteristics.`
|
||||||
|
|
||||||
|
`Do not use <Picture 3> to determine body proportions, clothing, environment or scene composition.`
|
||||||
|
|
||||||
|
## Full-body vs headshot conflict rule
|
||||||
|
|
||||||
|
Always resolve authority explicitly:
|
||||||
|
|
||||||
|
`<Picture 2> defines Subject 1's overall body, proportions, hair and clothing.`
|
||||||
|
|
||||||
|
`<Picture 3> and <Picture 4> provide additional facial identity information for the same Subject 1.`
|
||||||
|
|
||||||
|
`Use the headshot references to improve facial identity while preserving the body, clothing and overall appearance established by <Picture 2>.`
|
||||||
|
|
||||||
|
## Multi-character separation
|
||||||
|
|
||||||
|
Keep groups explicit to avoid identity leakage:
|
||||||
|
|
||||||
|
`<Picture 2> and <Picture 3> represent Subject 1.`
|
||||||
|
|
||||||
|
`<Picture 4> and <Picture 5> represent Subject 2.`
|
||||||
|
|
||||||
|
`Do not allow the identity or physical characteristics of one subject to influence another subject.`
|
||||||
|
|
||||||
|
## Background isolation rule
|
||||||
|
|
||||||
|
Character references often contain irrelevant backgrounds. State that those
|
||||||
|
backgrounds should be ignored.
|
||||||
|
|
||||||
|
Preferred phrasing:
|
||||||
|
|
||||||
|
`Use all character reference images only for their assigned character information.`
|
||||||
|
|
||||||
|
`Ignore their original backgrounds, locations, scene compositions, camera positions and environmental lighting.`
|
||||||
|
|
||||||
|
`The environment is defined exclusively by <Picture 1>.`
|
||||||
|
|
||||||
|
## Specialist references
|
||||||
|
|
||||||
|
Narrow responsibility prevents conflicts.
|
||||||
|
|
||||||
|
Example clothing phrasing:
|
||||||
|
|
||||||
|
`<Picture 5> defines the clothing worn by Subject 1.`
|
||||||
|
|
||||||
|
`Use <Picture 5> only as a clothing reference.`
|
||||||
|
|
||||||
|
`Ignore the person, face, body, pose, environment and background shown in <Picture 5>.`
|
||||||
|
|
||||||
|
Example vehicle phrasing:
|
||||||
|
|
||||||
|
`<Picture 6> defines the appearance of the vehicle.`
|
||||||
|
|
||||||
|
`Use <Picture 6> only as a visual reference for the vehicle.`
|
||||||
|
|
||||||
|
`Ignore its original environment, camera position and background.`
|
||||||
|
|
||||||
|
## 8GB VRAM constraints
|
||||||
|
|
||||||
|
Assume this baseline unless direct testing proves otherwise:
|
||||||
|
|
||||||
|
- resolution: `480p`
|
||||||
|
- maximum duration: `10s`
|
||||||
|
- one coherent action beat per shot
|
||||||
|
|
||||||
|
Do not design a single generation as a long sequence of unrelated events.
|
||||||
|
|
||||||
|
## Action design for 10 seconds
|
||||||
|
|
||||||
|
Good:
|
||||||
|
|
||||||
|
- `Subject 1 walks towards Subject 2.`
|
||||||
|
- `Subject 2 turns towards Subject 1.`
|
||||||
|
- `They briefly look at each other.`
|
||||||
|
- `The camera slowly tracks forward.`
|
||||||
|
|
||||||
|
Bad:
|
||||||
|
|
||||||
|
- entering a room
|
||||||
|
- crossing the room
|
||||||
|
- talking
|
||||||
|
- sitting
|
||||||
|
- picking something up
|
||||||
|
- finishing a conversation
|
||||||
|
- standing
|
||||||
|
- leaving
|
||||||
|
|
||||||
|
all in one `10s` shot
|
||||||
|
|
||||||
|
## Timing guidance
|
||||||
|
|
||||||
|
Simple timing is useful when sequencing matters:
|
||||||
|
|
||||||
|
- `0-3 seconds`: approach
|
||||||
|
- `3-7 seconds`: reaction
|
||||||
|
- `7-10 seconds`: settle / hold
|
||||||
|
|
||||||
|
Avoid over-micro-timing unless necessary.
|
||||||
|
|
||||||
|
## Dialogue guidance
|
||||||
|
|
||||||
|
Dialogue must fit within the same `10s` budget as motion and camera changes.
|
||||||
|
|
||||||
|
If the line is long, keep action simple.
|
||||||
|
|
||||||
|
Preferred pattern:
|
||||||
|
|
||||||
|
- character holds position or makes one simple move
|
||||||
|
- one spoken line
|
||||||
|
- one visible reaction
|
||||||
|
- one restrained camera movement
|
||||||
|
|
||||||
|
## Camera guidance
|
||||||
|
|
||||||
|
The master scene defines the opening frame.
|
||||||
|
|
||||||
|
Prompt text should describe how that camera changes, not replace it with a new
|
||||||
|
simultaneous framing.
|
||||||
|
|
||||||
|
Good camera phrasing:
|
||||||
|
|
||||||
|
- `The camera slowly pushes forward.`
|
||||||
|
- `The camera gently pans right to follow Subject 1.`
|
||||||
|
- `The camera slowly arcs around both subjects while maintaining their spatial relationship.`
|
||||||
|
|
||||||
|
## Expression guidance
|
||||||
|
|
||||||
|
Headshots lock identity, not frozen expression.
|
||||||
|
|
||||||
|
Expression belongs in performance instructions:
|
||||||
|
|
||||||
|
`Subject 1 initially appears relaxed.`
|
||||||
|
|
||||||
|
`As Subject 2 speaks, their expression gradually becomes concerned.`
|
||||||
|
|
||||||
|
## Recommended prompt architecture
|
||||||
|
|
||||||
|
Use this section order:
|
||||||
|
|
||||||
|
1. `Reference definitions`
|
||||||
|
2. `Scene anchor`
|
||||||
|
3. `Reference restrictions`
|
||||||
|
4. `Subject placement`
|
||||||
|
5. `Action / performance`
|
||||||
|
6. `Camera`
|
||||||
|
7. `Dialogue / audio`
|
||||||
|
|
||||||
|
## Single-character template
|
||||||
|
|
||||||
|
```text
|
||||||
|
<Picture 1> is the master scene, environment and composition reference.
|
||||||
|
|
||||||
|
<Picture 2>, <Picture 3> and <Picture 4> all represent Subject 1.
|
||||||
|
|
||||||
|
<Picture 2> is Subject 1's primary full-body reference and defines
|
||||||
|
their overall appearance, body proportions, build, hair and clothing.
|
||||||
|
|
||||||
|
<Picture 3> and <Picture 4> are complementary facial identity
|
||||||
|
references for Subject 1.
|
||||||
|
|
||||||
|
Use these headshot references to reinforce Subject 1's facial identity
|
||||||
|
while preserving the overall appearance established by <Picture 2>.
|
||||||
|
|
||||||
|
Use <Picture 1> as the sole visual reference for the environment.
|
||||||
|
|
||||||
|
Preserve its initial composition, camera position, perspective,
|
||||||
|
field of view, lighting, environment geometry and spatial layout.
|
||||||
|
|
||||||
|
Ignore the backgrounds, environments and compositions shown in
|
||||||
|
<Picture 2>, <Picture 3> and <Picture 4>.
|
||||||
|
|
||||||
|
[ACTION — MAXIMUM 10 SECONDS]
|
||||||
|
|
||||||
|
[CAMERA]
|
||||||
|
|
||||||
|
[DIALOGUE]
|
||||||
|
|
||||||
|
[AUDIO]
|
||||||
|
```
|
||||||
|
|
||||||
|
## Multi-character template
|
||||||
|
|
||||||
|
```text
|
||||||
|
<Picture 1> is the master scene, environment and composition reference.
|
||||||
|
|
||||||
|
<Picture 2>, <Picture 3> and <Picture 4> all represent Subject 1.
|
||||||
|
|
||||||
|
<Picture 2> is the primary full-body reference for Subject 1
|
||||||
|
and defines their overall appearance, body proportions, build,
|
||||||
|
hair and clothing.
|
||||||
|
|
||||||
|
<Picture 3> and <Picture 4> are complementary facial identity
|
||||||
|
references for Subject 1 and provide additional information about
|
||||||
|
their facial structure and appearance.
|
||||||
|
|
||||||
|
<Picture 5> and <Picture 6> both represent Subject 2.
|
||||||
|
|
||||||
|
<Picture 5> is the primary full-body reference for Subject 2
|
||||||
|
and defines their overall appearance, body proportions, build,
|
||||||
|
hair and clothing.
|
||||||
|
|
||||||
|
<Picture 6> is a facial identity reference for Subject 2.
|
||||||
|
|
||||||
|
The opening frame uses the environment and composition established
|
||||||
|
by <Picture 1>.
|
||||||
|
|
||||||
|
Use <Picture 1> as the sole visual reference for the environment.
|
||||||
|
|
||||||
|
Preserve its initial camera position, framing, perspective,
|
||||||
|
field of view, horizon, lighting, environment geometry,
|
||||||
|
object positions, proportions and spatial layout.
|
||||||
|
|
||||||
|
Do not duplicate, offset, mirror, layer or reconstruct a second
|
||||||
|
version of the environment.
|
||||||
|
|
||||||
|
Use the full-body character references to establish the overall
|
||||||
|
appearance of their respective subjects.
|
||||||
|
|
||||||
|
Use the headshot references to reinforce facial identity while
|
||||||
|
preserving the body and overall appearance established by the
|
||||||
|
corresponding full-body reference.
|
||||||
|
|
||||||
|
Ignore the backgrounds, environments, camera positions and scene
|
||||||
|
compositions shown in all character reference images.
|
||||||
|
|
||||||
|
Do not allow the identity or physical characteristics of one
|
||||||
|
subject to influence another subject.
|
||||||
|
|
||||||
|
[ACTION AND PERFORMANCE — MAXIMUM 10-SECOND SHOT]
|
||||||
|
|
||||||
|
[CAMERA MOVEMENT]
|
||||||
|
|
||||||
|
[DIALOGUE]
|
||||||
|
|
||||||
|
[AUDIO / ENVIRONMENTAL SOUND]
|
||||||
|
```
|
||||||
|
|
||||||
|
## Efficiency rule
|
||||||
|
|
||||||
|
Do not add references just because they exist.
|
||||||
|
|
||||||
|
Good default set for one important character:
|
||||||
|
|
||||||
|
- `1 x` full-body reference
|
||||||
|
- `1 x` frontal headshot
|
||||||
|
- `1 x` three-quarter headshot
|
||||||
|
|
||||||
|
The goal is the smallest reference set that clearly defines the required
|
||||||
|
information.
|
||||||
|
|
||||||
|
## Practical default checklist
|
||||||
|
|
||||||
|
- work at `480p`
|
||||||
|
- treat `10s` as the safe maximum shot duration
|
||||||
|
- use one coherent action or interaction per generation
|
||||||
|
- use a master scene whenever possible
|
||||||
|
- give the master scene authority over environment and composition
|
||||||
|
- give each important character one primary full-body reference
|
||||||
|
- add headshots only when facial identity needs reinforcement
|
||||||
|
- explicitly state which images belong to the same character
|
||||||
|
- stop character-reference backgrounds from influencing the master scene
|
||||||
|
- use the text prompt mainly for action, performance, camera, dialogue, and audio
|
||||||
@@ -28,6 +28,14 @@
|
|||||||
- Keep `optimized/text-to-video/minimax-h3-3070-turbo-master/` as the first MiniMax default.
|
- Keep `optimized/text-to-video/minimax-h3-3070-turbo-master/` as the first MiniMax default.
|
||||||
- Keep the INT8 reference-video workflow as the next donor for future low-VRAM optimization.
|
- Keep the INT8 reference-video workflow as the next donor for future low-VRAM optimization.
|
||||||
|
|
||||||
|
## Ref2VA prompt guidance
|
||||||
|
|
||||||
|
- For `MiniMax H3 Ref2VA`, treat `480p` and `10s` as the safe planning baseline on `8GB VRAM` unless direct tests prove otherwise.
|
||||||
|
- Use one master-scene reference to control environment/composition, one primary full-body reference per subject to control overall appearance, and headshots only to reinforce facial identity.
|
||||||
|
- Explicitly state when multiple images represent the same subject, and explicitly tell the model to ignore character-reference backgrounds so they do not fight the master environment.
|
||||||
|
- Use the prompt mainly for action, performance, camera movement, dialogue, and audio rather than re-describing scene details already visible in the references.
|
||||||
|
- Full guidance now lives in `optimized/notes/minimax-h3-ref2va-prompt-guidelines.md`.
|
||||||
|
|
||||||
## What needs Monday proof
|
## What needs Monday proof
|
||||||
|
|
||||||
- whether the Turbo LoRA workflow is just technically possible or actually usable
|
- whether the Turbo LoRA workflow is just technically possible or actually usable
|
||||||
|
|||||||
Reference in New Issue
Block a user