diff --git a/optimized/notes/minimax-h3-ref2va-prompt-guidelines.md b/optimized/notes/minimax-h3-ref2va-prompt-guidelines.md new file mode 100644 index 0000000..a46a598 --- /dev/null +++ b/optimized/notes/minimax-h3-ref2va-prompt-guidelines.md @@ -0,0 +1,389 @@ +# MiniMax H3 Ref2VA Prompt Guidelines + +Source basis: Chris-provided `MiniMax H3 Ref2VA Reference & Scene Guidelines — V2` + +## Purpose + +Practical prompt and reference-structure guide for `MiniMax H3 Ref2VA`, +especially on an `8GB VRAM` workflow where the safe planning baseline is: + +- `480p` +- maximum `10s` per generated shot +- one coherent shot per generation +- explicit roles for every reference image + +## Core rule + +Reduce the system to five responsibilities: + +1. `Master scene` + Defines where the shot happens and how it initially looks. +2. `Full-body references` + Define who is present and their overall appearance. +3. `Headshot references` + Reinforce whose face it is. +4. `Specialist references` + Define clothing, props, vehicles, or other narrow details. +5. `Prompt` + Defines what happens during the next `10` seconds. + +## Reference hierarchy + +Use this authority order: + +1. `master scene / environment` +2. `primary full-body character references` +3. `facial identity references` +4. `clothing references` +5. `props / vehicles / specialist references` +6. `text prompt` + +Within a character: + +- `full body` controls overall appearance, body proportions, hair, and clothing +- `headshots` refine facial identity + +## Master scene rules + +Prefer `` as the master scene when possible. + +It should control: + +- environment +- composition +- camera position +- perspective +- field of view +- lighting +- initial subject placement +- important object placement + +Preferred phrasing: + +` is the master scene, environment and composition reference.` + +`Use as the sole visual reference for the environment.` + +`Preserve its initial camera position, framing, perspective, field of view, horizon, lighting, environment geometry, object positions, proportions and spatial layout.` + +`Do not duplicate, offset, mirror, layer or reconstruct a second version of the environment.` + +## Environment rules + +If a reference already establishes the environment, do not re-describe every +visible object in text unless something changes. + +Prefer: + +`The environment, background, spatial layout and lighting remain as established by .` + +Over a full textual rebuild of the room. + +## Character reference grouping + +When multiple images represent the same person, explicitly say so. + +Preferred phrasing: + +`, and all represent Subject 1.` + +`Use their consistent identity information together to preserve Subject 1 throughout the video.` + +`Do not describe them as separate subjects.` + +## Full-body reference role + +Primary use: + +- overall appearance +- body proportions +- build +- hair +- clothing +- silhouette + +Preferred phrasing: + +` is the primary full-body reference for Subject 1.` + +`Use to define Subject 1's overall appearance, body proportions, build, hair and clothing.` + +`Ignore the background, environment, camera composition and location shown in .` + +## Headshot reference role + +Primary use: + +- facial identity +- facial structure +- eyes +- nose +- mouth +- skin appearance + +Preferred phrasing: + +` is a facial identity reference for Subject 1.` + +`Use to reinforce Subject 1's face, facial structure, eyes, nose, mouth, skin appearance and other facial characteristics.` + +`Do not use to determine body proportions, clothing, environment or scene composition.` + +## Full-body vs headshot conflict rule + +Always resolve authority explicitly: + +` defines Subject 1's overall body, proportions, hair and clothing.` + +` and provide additional facial identity information for the same Subject 1.` + +`Use the headshot references to improve facial identity while preserving the body, clothing and overall appearance established by .` + +## Multi-character separation + +Keep groups explicit to avoid identity leakage: + +` and represent Subject 1.` + +` and represent Subject 2.` + +`Do not allow the identity or physical characteristics of one subject to influence another subject.` + +## Background isolation rule + +Character references often contain irrelevant backgrounds. State that those +backgrounds should be ignored. + +Preferred phrasing: + +`Use all character reference images only for their assigned character information.` + +`Ignore their original backgrounds, locations, scene compositions, camera positions and environmental lighting.` + +`The environment is defined exclusively by .` + +## Specialist references + +Narrow responsibility prevents conflicts. + +Example clothing phrasing: + +` defines the clothing worn by Subject 1.` + +`Use only as a clothing reference.` + +`Ignore the person, face, body, pose, environment and background shown in .` + +Example vehicle phrasing: + +` defines the appearance of the vehicle.` + +`Use only as a visual reference for the vehicle.` + +`Ignore its original environment, camera position and background.` + +## 8GB VRAM constraints + +Assume this baseline unless direct testing proves otherwise: + +- resolution: `480p` +- maximum duration: `10s` +- one coherent action beat per shot + +Do not design a single generation as a long sequence of unrelated events. + +## Action design for 10 seconds + +Good: + +- `Subject 1 walks towards Subject 2.` +- `Subject 2 turns towards Subject 1.` +- `They briefly look at each other.` +- `The camera slowly tracks forward.` + +Bad: + +- entering a room +- crossing the room +- talking +- sitting +- picking something up +- finishing a conversation +- standing +- leaving + +all in one `10s` shot + +## Timing guidance + +Simple timing is useful when sequencing matters: + +- `0-3 seconds`: approach +- `3-7 seconds`: reaction +- `7-10 seconds`: settle / hold + +Avoid over-micro-timing unless necessary. + +## Dialogue guidance + +Dialogue must fit within the same `10s` budget as motion and camera changes. + +If the line is long, keep action simple. + +Preferred pattern: + +- character holds position or makes one simple move +- one spoken line +- one visible reaction +- one restrained camera movement + +## Camera guidance + +The master scene defines the opening frame. + +Prompt text should describe how that camera changes, not replace it with a new +simultaneous framing. + +Good camera phrasing: + +- `The camera slowly pushes forward.` +- `The camera gently pans right to follow Subject 1.` +- `The camera slowly arcs around both subjects while maintaining their spatial relationship.` + +## Expression guidance + +Headshots lock identity, not frozen expression. + +Expression belongs in performance instructions: + +`Subject 1 initially appears relaxed.` + +`As Subject 2 speaks, their expression gradually becomes concerned.` + +## Recommended prompt architecture + +Use this section order: + +1. `Reference definitions` +2. `Scene anchor` +3. `Reference restrictions` +4. `Subject placement` +5. `Action / performance` +6. `Camera` +7. `Dialogue / audio` + +## Single-character template + +```text + is the master scene, environment and composition reference. + +, and all represent Subject 1. + + is Subject 1's primary full-body reference and defines +their overall appearance, body proportions, build, hair and clothing. + + and are complementary facial identity +references for Subject 1. + +Use these headshot references to reinforce Subject 1's facial identity +while preserving the overall appearance established by . + +Use as the sole visual reference for the environment. + +Preserve its initial composition, camera position, perspective, +field of view, lighting, environment geometry and spatial layout. + +Ignore the backgrounds, environments and compositions shown in +, and . + +[ACTION — MAXIMUM 10 SECONDS] + +[CAMERA] + +[DIALOGUE] + +[AUDIO] +``` + +## Multi-character template + +```text + is the master scene, environment and composition reference. + +, and all represent Subject 1. + + is the primary full-body reference for Subject 1 +and defines their overall appearance, body proportions, build, +hair and clothing. + + and are complementary facial identity +references for Subject 1 and provide additional information about +their facial structure and appearance. + + and both represent Subject 2. + + is the primary full-body reference for Subject 2 +and defines their overall appearance, body proportions, build, +hair and clothing. + + is a facial identity reference for Subject 2. + +The opening frame uses the environment and composition established +by . + +Use as the sole visual reference for the environment. + +Preserve its initial camera position, framing, perspective, +field of view, horizon, lighting, environment geometry, +object positions, proportions and spatial layout. + +Do not duplicate, offset, mirror, layer or reconstruct a second +version of the environment. + +Use the full-body character references to establish the overall +appearance of their respective subjects. + +Use the headshot references to reinforce facial identity while +preserving the body and overall appearance established by the +corresponding full-body reference. + +Ignore the backgrounds, environments, camera positions and scene +compositions shown in all character reference images. + +Do not allow the identity or physical characteristics of one +subject to influence another subject. + +[ACTION AND PERFORMANCE — MAXIMUM 10-SECOND SHOT] + +[CAMERA MOVEMENT] + +[DIALOGUE] + +[AUDIO / ENVIRONMENTAL SOUND] +``` + +## Efficiency rule + +Do not add references just because they exist. + +Good default set for one important character: + +- `1 x` full-body reference +- `1 x` frontal headshot +- `1 x` three-quarter headshot + +The goal is the smallest reference set that clearly defines the required +information. + +## Practical default checklist + +- work at `480p` +- treat `10s` as the safe maximum shot duration +- use one coherent action or interaction per generation +- use a master scene whenever possible +- give the master scene authority over environment and composition +- give each important character one primary full-body reference +- add headshots only when facial identity needs reinforcement +- explicitly state which images belong to the same character +- stop character-reference backgrounds from influencing the master scene +- use the text prompt mainly for action, performance, camera, dialogue, and audio diff --git a/optimized/notes/minimax.md b/optimized/notes/minimax.md index 22815d9..b2329a3 100644 --- a/optimized/notes/minimax.md +++ b/optimized/notes/minimax.md @@ -28,6 +28,14 @@ - Keep `optimized/text-to-video/minimax-h3-3070-turbo-master/` as the first MiniMax default. - Keep the INT8 reference-video workflow as the next donor for future low-VRAM optimization. +## Ref2VA prompt guidance + +- For `MiniMax H3 Ref2VA`, treat `480p` and `10s` as the safe planning baseline on `8GB VRAM` unless direct tests prove otherwise. +- Use one master-scene reference to control environment/composition, one primary full-body reference per subject to control overall appearance, and headshots only to reinforce facial identity. +- Explicitly state when multiple images represent the same subject, and explicitly tell the model to ignore character-reference backgrounds so they do not fight the master environment. +- Use the prompt mainly for action, performance, camera movement, dialogue, and audio rather than re-describing scene details already visible in the references. +- Full guidance now lives in `optimized/notes/minimax-h3-ref2va-prompt-guidelines.md`. + ## What needs Monday proof - whether the Turbo LoRA workflow is just technically possible or actually usable