Implement structured H3 reference objects
This commit is contained in:
@@ -34,13 +34,14 @@
|
||||
- Reads back the nine optional images for a selected MiniMax H3 plan scene, for example by connecting the current `clip_index`.
|
||||
|
||||
- `Dumas H3 Long Videos (FL2VA + REF2VA)`
|
||||
- Inputs: H3 model stack, prompt socket, optional `first_frame`, optional `ref_image_1`..`ref_image_9`, optional `plan`, optional `plan_scene_index`, plus the upstream long-video control surface for pacing, continuity, audio, overlays, and guards
|
||||
- Inputs: H3 model stack, prompt socket, optional `first_frame`, optional `ref_1`..`ref_9`, optional `plan`, optional `plan_scene_index`, plus the upstream long-video control surface for pacing, continuity, audio, overlays, and guards
|
||||
- Outputs: `images`, `audio`, `info`, `script`, `frames_per_shot`, `total_frames`, `shots`, `video_seconds`, `fps`, `fps_int`, `latent`, `soundscape`
|
||||
- First-pass Dumas port of the `MiniMax-H3-Longvideos` sampler, brought in as a local starting point for long-form H3 chaining work.
|
||||
- Keeps the upstream split-beats / handoff / ref-routing behavior close to source so future Dumas-specific improvements can be compared against a known baseline.
|
||||
- Only the canonical `DumasH3LongVideos` node key is exposed now; the older FL2VA/REF2VA alias entries are no longer duplicated in the Add Node menu.
|
||||
- Prompt `<Picture N>` tags now map to the actual ref socket numbers you wire, even with gaps such as only `ref_image_2` and `ref_image_7` connected.
|
||||
- A connected H3 plan can now supply the current scene’s 9-image bundle directly; any directly-wired `ref_image_*` socket overrides the same numbered plan slot.
|
||||
- Prompt `<Picture N>` tags now map to the actual ref socket numbers you wire, even with gaps such as only `ref_2` and `ref_7` connected.
|
||||
- Character refs now contribute appearance and wardrobe context from the same structured object, while location refs contribute environment context from theirs.
|
||||
- A connected H3 plan can now supply the current scene’s 9-image bundle directly; any directly-wired `ref_*` socket overrides the same numbered plan slot.
|
||||
- The default ref2v bias is now stronger: `ref_mode` defaults to `auto ref2v` so untagged prompts condition every shot instead of only shot 1, and `ref_noise_aug` defaults to `0.95` rather than the upstream-literal `0.999`.
|
||||
|
||||
- `Dumas H3 Shot Length`
|
||||
@@ -53,11 +54,15 @@
|
||||
- Outputs: `format`, `report`
|
||||
- Reports the detected H3 base precision / quant format and the relevant compute-capability hints for the current card.
|
||||
|
||||
- `Dumas Character Helper`
|
||||
- Inputs: `image1`, `image2`, `image1_picture_id`, `image2_picture_id`, `character_id`, `name`, `alias`, `gender`, `age`, `nationality`, `occupation`, `height_feet`, `height_inches`, `accent`, `general`
|
||||
- Outputs: `image1`, `image2`, `character_text`
|
||||
- Passes both images through unchanged and builds a character reference string such as `<Picture 2> and <Picture 3> reference the same character who is called Dave.`
|
||||
- Collapses optional non-visual facts such as alias, gender, age, nationality, occupation, height, and accent into one comma-separated sentence, then ends with the freeform note as the final sentence when provided.
|
||||
- `Dumas Character Reference`
|
||||
- Inputs: `image`, `picture_id`, `character_id`, `name`, `alias`, `gender`, `age`, `nationality`, `occupation`, `height_feet`, `height_inches`, `accent`, `description`, `general`, `wardrobe`
|
||||
- Output: `reference`
|
||||
- Builds one structured `REFERENCE` object carrying the conditioning image, identity description, wardrobe, general notes, and simple facts together.
|
||||
|
||||
- `Dumas Location Reference`
|
||||
- Inputs: `image`, `picture_id`, `location_id`, `name`, `alias`, `description`, `general`
|
||||
- Output: `reference`
|
||||
- Builds one structured `REFERENCE` object for a location/environment so H3 can use the same socket type for both character and scenic refs.
|
||||
|
||||
- `Dumas JSON String to Object`
|
||||
- Input: `json_string`
|
||||
@@ -215,7 +220,7 @@ decr -> use index - 1
|
||||
|
||||
`Dumas H3 Plan Attach Scene Images` and `Dumas H3 Plan Extract Scene Images` are a companion pair for `ComfyUI-MiniMaxH3-Contex-Loop` and the local `ref2v` lane. The upstream H3 plan node cannot dynamically grow nine new image sockets for every JSON-defined scene, so Dumas stores scene image bindings beside the plan using a lightweight token and an in-memory registry. That keeps `plan.json` archiving intact while still letting you wire up nine IMAGE sockets per scene through chained helper nodes.
|
||||
|
||||
`Dumas Character Helper` lives in `Dumas/String`. Use the picture ID dropdowns to decide which `<Picture N>` tags get mentioned in the generated text, while the two IMAGE sockets continue downstream unchanged. The node is tuned for useful non-visible facts rather than visual descriptions already obvious from the reference images.
|
||||
`Dumas Character Reference` and `Dumas Location Reference` live in `Dumas/MiniMax`. Both output a structured `REFERENCE` object that carries the image plus its semantic payload. `Dumas H3 Long Videos` accepts those `REFERENCE` sockets directly on `ref_1`..`ref_9`, resolves `<Picture N>` against the wired slot positions, and can also pull character wardrobe context from the structured ref data when `character_memory` is left blank.
|
||||
|
||||
`Dumas Strip Iteration Suffix` keeps the part before the first underscore and drops the rest. Names like `char123_pose_final.png` become `char123.png`, while names with no underscore such as `char123.png` are left untouched.
|
||||
|
||||
|
||||
@@ -0,0 +1,386 @@
|
||||
# Dumas Reference System Spec
|
||||
|
||||
## Goal
|
||||
|
||||
Replace the current image-only / text-fragment character-reference flow with one
|
||||
structured reference system that works for both:
|
||||
|
||||
- characters
|
||||
- locations / environments
|
||||
|
||||
Compatibility with old graphs is **not** a goal. This is a clean redesign for a
|
||||
single-user workflow.
|
||||
|
||||
## Core Direction
|
||||
|
||||
The system should stop treating a reference as "just an image".
|
||||
|
||||
A reference should be one structured object that carries:
|
||||
|
||||
- the image used for model conditioning
|
||||
- the semantic identity of the reference
|
||||
- the descriptive text tied to that reference
|
||||
- the routing metadata needed by H3 beat / prompt logic
|
||||
|
||||
This object type is called `REFERENCE`.
|
||||
|
||||
## Main Outcomes
|
||||
|
||||
After this redesign:
|
||||
|
||||
- H3 ref sockets accept `REFERENCE`, not raw `IMAGE`
|
||||
- `<Picture N>` resolves to the bound `REFERENCE` object
|
||||
- character references can provide identity, description, and wardrobe from one source
|
||||
- location references can provide environment description from one source
|
||||
- beat prompt logic no longer has to guess where a name or picture tag should pull
|
||||
its descriptive payload from
|
||||
|
||||
## Reference Kinds
|
||||
|
||||
`REFERENCE` is one base type with a `kind` field.
|
||||
|
||||
Initial supported kinds:
|
||||
|
||||
- `character`
|
||||
- `location`
|
||||
|
||||
Future kinds could be added later if useful, but they are out of scope for the
|
||||
first pass.
|
||||
|
||||
## Reference Object Schema
|
||||
|
||||
The `REFERENCE` object should be a plain JSON-serializable structure plus an
|
||||
attached Comfy image payload where needed by downstream nodes.
|
||||
|
||||
Conceptual schema:
|
||||
|
||||
```json
|
||||
{
|
||||
"kind": "character",
|
||||
"id": "francine",
|
||||
"name": "Francine",
|
||||
"aliases": [],
|
||||
"picture_id": 1,
|
||||
"picture_label": "<Picture 1>",
|
||||
"image": "<IMAGE>",
|
||||
"summary": "Same character shown in <Picture 1>.",
|
||||
"description": "Face, hair, proportions, persistent visual identity.",
|
||||
"wardrobe": "Clothing / styling / persistent look notes.",
|
||||
"general": "Freeform notes.",
|
||||
"facts": {
|
||||
"gender": "",
|
||||
"age": "",
|
||||
"nationality": "",
|
||||
"occupation": "",
|
||||
"height_feet": "",
|
||||
"height_inches": "",
|
||||
"accent": ""
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
For locations:
|
||||
|
||||
```json
|
||||
{
|
||||
"kind": "location",
|
||||
"id": "coffee-shop-01",
|
||||
"name": "Coffee Shop",
|
||||
"aliases": [],
|
||||
"picture_id": 2,
|
||||
"picture_label": "<Picture 2>",
|
||||
"image": "<IMAGE>",
|
||||
"summary": "Environment shown in <Picture 2>.",
|
||||
"description": "Persistent environment / layout / atmosphere description.",
|
||||
"wardrobe": "",
|
||||
"general": "Freeform notes.",
|
||||
"facts": {}
|
||||
}
|
||||
```
|
||||
|
||||
## Field Meaning
|
||||
|
||||
### Shared fields
|
||||
|
||||
- `kind`
|
||||
- Required.
|
||||
- `character` or `location`.
|
||||
|
||||
- `id`
|
||||
- Required.
|
||||
- Stable, slug-like internal identifier.
|
||||
- Used for matching and future persistence.
|
||||
|
||||
- `name`
|
||||
- Optional but strongly recommended.
|
||||
- Human-readable label.
|
||||
|
||||
- `aliases`
|
||||
- Optional list of alternate match names.
|
||||
|
||||
- `picture_id`
|
||||
- Optional integer representing the intended `<Picture N>` identity.
|
||||
- This is authoring metadata, not the final socket position.
|
||||
|
||||
- `picture_label`
|
||||
- Derived convenience text like `<Picture 1>`.
|
||||
|
||||
- `image`
|
||||
- Required.
|
||||
- The actual Comfy `IMAGE` payload used for reference conditioning.
|
||||
|
||||
- `summary`
|
||||
- Short purpose statement for prompt assembly and debugging.
|
||||
|
||||
- `description`
|
||||
- Primary descriptive payload tied to the reference.
|
||||
- For characters this is the physical identity description.
|
||||
- For locations this is the environment/layout/atmosphere description.
|
||||
|
||||
- `general`
|
||||
- Optional freeform notes.
|
||||
|
||||
- `facts`
|
||||
- Optional structured detail map.
|
||||
|
||||
### Character-only field
|
||||
|
||||
- `wardrobe`
|
||||
- Optional but expected for character references.
|
||||
- Persistent clothing / styling / accessories / visual state notes.
|
||||
|
||||
### Location usage note
|
||||
|
||||
For `kind = "location"`, `wardrobe` should be empty.
|
||||
|
||||
## Producer Nodes
|
||||
|
||||
### 1. Character helper replacement
|
||||
|
||||
Current `Dumas Character Helper` should be redesigned into a character reference
|
||||
builder.
|
||||
|
||||
Recommended name:
|
||||
|
||||
- `Dumas Character Reference`
|
||||
|
||||
Inputs:
|
||||
|
||||
- `image`
|
||||
- `picture_id`
|
||||
- `character_id`
|
||||
- `name`
|
||||
- `alias`
|
||||
- `gender`
|
||||
- `age`
|
||||
- `nationality`
|
||||
- `occupation`
|
||||
- `height_feet`
|
||||
- `height_inches`
|
||||
- `accent`
|
||||
- `description`
|
||||
- `wardrobe`
|
||||
- `general`
|
||||
|
||||
Outputs:
|
||||
|
||||
- `reference` (`REFERENCE`)
|
||||
|
||||
Notes:
|
||||
|
||||
- This node should stop emitting the current mixed output pattern.
|
||||
- It should build one clean `REFERENCE` object with `kind = "character"`.
|
||||
|
||||
### 2. Location / environment helper
|
||||
|
||||
Add a new builder node.
|
||||
|
||||
Recommended name:
|
||||
|
||||
- `Dumas Location Reference`
|
||||
|
||||
Inputs:
|
||||
|
||||
- `image`
|
||||
- `picture_id`
|
||||
- `location_id`
|
||||
- `name`
|
||||
- `alias`
|
||||
- `description`
|
||||
- `general`
|
||||
|
||||
Outputs:
|
||||
|
||||
- `reference` (`REFERENCE`)
|
||||
|
||||
Notes:
|
||||
|
||||
- This node builds one `REFERENCE` object with `kind = "location"`.
|
||||
- It should not attempt to mimic character-specific fields.
|
||||
|
||||
## Consumer Changes
|
||||
|
||||
## H3 Long Videos
|
||||
|
||||
Current numbered sockets:
|
||||
|
||||
- `ref_image_1` .. `ref_image_9`
|
||||
|
||||
should become:
|
||||
|
||||
- `ref_1` .. `ref_9`
|
||||
|
||||
Type:
|
||||
|
||||
- `REFERENCE`
|
||||
|
||||
Each socket should provide access to:
|
||||
|
||||
- the underlying `image`
|
||||
- the structured metadata
|
||||
|
||||
### Internal handling
|
||||
|
||||
The H3 node should internally derive two parallel lanes:
|
||||
|
||||
1. reference images for model conditioning
|
||||
2. reference definitions for semantic beat / prompt resolution
|
||||
|
||||
It must no longer depend on detached text sources to know what a reference means.
|
||||
|
||||
## Beat / prompt resolution rules
|
||||
|
||||
### Resolution by `<Picture N>`
|
||||
|
||||
If a beat contains `<Picture N>`:
|
||||
|
||||
- map `N` to the bound socket position in that shot
|
||||
- resolve the bound `REFERENCE`
|
||||
- use the resolved object's data when generating the shot-level descriptive payload
|
||||
|
||||
### Resolution by character name
|
||||
|
||||
For `kind = "character"`:
|
||||
|
||||
- match against `name`
|
||||
- match against `aliases`
|
||||
- optionally match `id`
|
||||
|
||||
When a character is matched:
|
||||
|
||||
- pull identity text from `description`
|
||||
- pull clothing / styling from `wardrobe`
|
||||
- pull additional context from `general` and `facts` if needed
|
||||
|
||||
### Resolution by location
|
||||
|
||||
For `kind = "location"`:
|
||||
|
||||
- `<Picture N>` should resolve directly
|
||||
- later name-based location matching can be added if useful, but the first pass
|
||||
does not need aggressive free-text environment matching
|
||||
|
||||
When a location is matched:
|
||||
|
||||
- pull environment text from `description`
|
||||
- optionally use `general`
|
||||
|
||||
## Prompt Assembly Expectations
|
||||
|
||||
Character refs should contribute:
|
||||
|
||||
- who the subject is
|
||||
- what persistent visual identity should be preserved
|
||||
- what clothing / styling should persist
|
||||
|
||||
Location refs should contribute:
|
||||
|
||||
- where the scene is
|
||||
- what environment layout / mood / persistent scenic identity should be preserved
|
||||
|
||||
The node should keep these conceptually separate.
|
||||
|
||||
Character reference text should not be mistaken for environment text.
|
||||
Location reference text should not be mistaken for wardrobe or identity text.
|
||||
|
||||
## Multi-reference Behavior
|
||||
|
||||
The system must support multiple references bound at once.
|
||||
|
||||
Typical cases:
|
||||
|
||||
- one character + one location
|
||||
- multiple character references
|
||||
- one character sheet reused across many beats
|
||||
|
||||
The H3 node should preserve current numbered socket behavior for explicit routing,
|
||||
but with `REFERENCE` objects replacing plain images.
|
||||
|
||||
## Plan / scene integration
|
||||
|
||||
The current H3 plan scene image system should evolve in parallel.
|
||||
|
||||
Where plan scenes currently carry image bundles, they should eventually carry
|
||||
reference bundles instead.
|
||||
|
||||
This is not required for the first code pass if it would enlarge the change too
|
||||
much, but the implementation should avoid boxing us into image-only assumptions.
|
||||
|
||||
## Non-goals
|
||||
|
||||
Not required for first pass:
|
||||
|
||||
- backward compatibility with old image-only graphs
|
||||
- migration helpers
|
||||
- automatic graph conversion
|
||||
- advanced location name matching
|
||||
- persistent storage outside the graph object itself
|
||||
|
||||
## Recommended Implementation Order
|
||||
|
||||
1. Define the `REFERENCE` type contract in code.
|
||||
2. Replace `Dumas Character Helper` with a character-reference builder node.
|
||||
3. Add `Dumas Location Reference`.
|
||||
4. Update H3 long videos sockets from `IMAGE` to `REFERENCE`.
|
||||
5. Update internal ref collection logic to extract both image and metadata.
|
||||
6. Update beat / prompt resolution to use `REFERENCE` objects.
|
||||
7. Update tests.
|
||||
8. Update README and workflow docs.
|
||||
|
||||
## Testing Requirements
|
||||
|
||||
Add tests for:
|
||||
|
||||
- character reference object construction
|
||||
- location reference object construction
|
||||
- `<Picture N>` resolving to the correct `REFERENCE`
|
||||
- name-based character lookup resolving to the correct `REFERENCE`
|
||||
- wardrobe text being pulled only from character refs
|
||||
- location description being pulled only from location refs
|
||||
- mixed character + location ref usage in one prompt
|
||||
- sparse numbered sockets still resolving correctly
|
||||
|
||||
## Open Design Choices
|
||||
|
||||
These should be decided during implementation, but the default recommendation is:
|
||||
|
||||
- one `REFERENCE` output per helper node
|
||||
- numbered H3 sockets remain explicit
|
||||
- no separate cast object in first pass
|
||||
- no compatibility fallback path
|
||||
|
||||
## Summary
|
||||
|
||||
This redesign turns references into first-class structured inputs.
|
||||
|
||||
The key rule is:
|
||||
|
||||
> A reference socket should carry both the image and the meaning of that image.
|
||||
|
||||
That is the change that removes the current ambiguity between:
|
||||
|
||||
- image conditioning
|
||||
- character identity
|
||||
- wardrobe data
|
||||
- location / environment description
|
||||
|
||||
+171
-24
@@ -8,7 +8,7 @@ One node covering both of H3's conditioning tasks:
|
||||
* REF2VA -- reference images condition the shot on what a character LOOKS like,
|
||||
independent of any frame.
|
||||
|
||||
Connect nothing to ref_image_* and it behaves exactly as the FL2VA node always
|
||||
Connect nothing to ref_* and it behaves exactly as the FL2VA node always
|
||||
did. Connect a reference and `ref_mode` decides which shots use it.
|
||||
|
||||
THE ONE RULE: a shot carries EITHER references or the last-frame handoff, never
|
||||
@@ -275,7 +275,7 @@ ADDED_WIDGETS = (
|
||||
"exposed_terms", "anatomy_guard", "lock_restraints", "solidity_guard",
|
||||
"motion_guard", "contact_guard",
|
||||
"auto_soundscape", "allow_nonspeech_vocals",
|
||||
"ref_image_5", "ref_image_6", "ref_image_7", "ref_image_8", "ref_image_9",
|
||||
"ref_5", "ref_6", "ref_7", "ref_8", "ref_9",
|
||||
"plan", "plan_scene_index",
|
||||
)
|
||||
|
||||
@@ -3869,13 +3869,14 @@ def ref_image_canvas(w, h, gen_w, gen_h, mode="match"):
|
||||
|
||||
|
||||
def _build_ref_images(vae, images, gen_w, gen_h, mode="match"):
|
||||
"""(tokenizer items, DiT blocks) for a list of reference IMAGE tensors.
|
||||
"""(tokenizer items, DiT blocks) for a list of reference images.
|
||||
|
||||
The tokenizer labels each one `<Picture N>:` itself, in the order given here --
|
||||
so the roster the prompt refers to is decided by input order, not by anything
|
||||
written in the prompt."""
|
||||
items, blocks = [], []
|
||||
for img in images:
|
||||
for source in images:
|
||||
img = _reference_image(source)
|
||||
if img is None:
|
||||
continue
|
||||
h, w = int(img.shape[1]), int(img.shape[2])
|
||||
@@ -3891,7 +3892,7 @@ def _build_shot_conditioning(clip, vae, prompt, width, height, length, fps, hand
|
||||
ref_images=None, ref_image_size="match", ref_noise_aug=None,
|
||||
audio_vae=None, silent=False):
|
||||
latent, fc = _empty_av_latent(width, height, length, fps)
|
||||
refs = [r for r in (ref_images or []) if r is not None]
|
||||
refs = [r for r in (ref_images or []) if _reference_image(r) is not None]
|
||||
if refs:
|
||||
# ref2va: this shot is reference-conditioned rather than keyframe-conditioned,
|
||||
# and run() decides which per shot. A tagged shot is handed the previous
|
||||
@@ -4044,9 +4045,135 @@ def picture_tags(text):
|
||||
return sorted({int(m.group(1)) for m in _PICTURE_TAG.finditer(text or "")})
|
||||
|
||||
|
||||
def _reference_slot(ref, slot_index=None):
|
||||
return _image_nodes.normalize_reference(ref, picture_id=slot_index, allow_image_fallback=True)
|
||||
|
||||
|
||||
def _reference_image(ref):
|
||||
try:
|
||||
normalized = _reference_slot(ref)
|
||||
except Exception:
|
||||
return None
|
||||
return normalized.get("image")
|
||||
|
||||
|
||||
def _reference_text(value):
|
||||
return " ".join(str(value or "").split()).strip()
|
||||
|
||||
|
||||
def _reference_sentence(value):
|
||||
text = _reference_text(value)
|
||||
if text and text[-1] not in ".!?":
|
||||
text += "."
|
||||
return text
|
||||
|
||||
|
||||
def _reference_name_keys(ref):
|
||||
names = []
|
||||
for key in ("name", "id"):
|
||||
value = _reference_text(ref.get(key))
|
||||
if value:
|
||||
names.append(value)
|
||||
for alias in ref.get("aliases") or []:
|
||||
value = _reference_text(alias)
|
||||
if value:
|
||||
names.append(value)
|
||||
seen = set()
|
||||
out = []
|
||||
for name in names:
|
||||
key = name.lower()
|
||||
if key in seen:
|
||||
continue
|
||||
seen.add(key)
|
||||
out.append(name)
|
||||
return out
|
||||
|
||||
|
||||
def _slot_refs_for_text(text, ref_slots):
|
||||
refs = []
|
||||
for slot_number in picture_tags(text):
|
||||
if not (1 <= slot_number <= len(ref_slots or [])):
|
||||
continue
|
||||
ref = ref_slots[slot_number - 1]
|
||||
if _reference_image(ref) is None:
|
||||
continue
|
||||
refs.append((slot_number, _reference_slot(ref, slot_number)))
|
||||
return refs
|
||||
|
||||
|
||||
def _named_character_refs_for_text(text, ref_slots):
|
||||
haystack = str(text or "")
|
||||
matched = []
|
||||
for slot_number, raw in enumerate(ref_slots or [], 1):
|
||||
ref = _reference_slot(raw, slot_number)
|
||||
if ref.get("kind") != "character" or _reference_image(ref) is None:
|
||||
continue
|
||||
for name in _reference_name_keys(ref):
|
||||
if re.search(r"\b" + re.escape(name) + r"\b", haystack, re.I):
|
||||
matched.append((slot_number, ref))
|
||||
break
|
||||
return matched
|
||||
|
||||
|
||||
def _matched_reference_slots(text, ref_slots):
|
||||
matched = []
|
||||
seen = set()
|
||||
for slot_number, ref in _slot_refs_for_text(text, ref_slots) + _named_character_refs_for_text(text, ref_slots):
|
||||
if slot_number in seen:
|
||||
continue
|
||||
seen.add(slot_number)
|
||||
matched.append((slot_number, ref))
|
||||
return matched
|
||||
|
||||
|
||||
def _reference_context_for_text(text, ref_slots):
|
||||
parts = []
|
||||
for slot_number, ref in _matched_reference_slots(text, ref_slots):
|
||||
label = _reference_text(ref.get("name")) or _reference_text(ref.get("id")) or f"reference {slot_number}"
|
||||
description = _reference_sentence(ref.get("description"))
|
||||
wardrobe = _reference_sentence(ref.get("wardrobe"))
|
||||
general = _reference_sentence(ref.get("general"))
|
||||
if ref.get("kind") == "location":
|
||||
if description:
|
||||
parts.append(f"Location context for {label}: {description}")
|
||||
if general:
|
||||
parts.append(f"Location notes for {label}: {general}")
|
||||
continue
|
||||
if description:
|
||||
parts.append(f"Persistent appearance for {label}: {description}")
|
||||
if wardrobe:
|
||||
parts.append(f"Persistent wardrobe/style for {label}: {wardrobe}")
|
||||
if general:
|
||||
parts.append(f"Character notes for {label}: {general}")
|
||||
return " ".join(parts).strip()
|
||||
|
||||
|
||||
def _reference_character_memory(ref_slots):
|
||||
lines = []
|
||||
seen = set()
|
||||
for slot_number, raw in enumerate(ref_slots or [], 1):
|
||||
ref = _reference_slot(raw, slot_number)
|
||||
if ref.get("kind") != "character":
|
||||
continue
|
||||
wardrobe = _reference_text(ref.get("wardrobe"))
|
||||
if not wardrobe:
|
||||
continue
|
||||
label = (
|
||||
_reference_text(ref.get("name"))
|
||||
or (_reference_name_keys(ref)[0] if _reference_name_keys(ref) else "")
|
||||
)
|
||||
line = f"{label} = {wardrobe}" if label else wardrobe
|
||||
key = line.lower()
|
||||
if key in seen:
|
||||
continue
|
||||
seen.add(key)
|
||||
lines.append(line)
|
||||
return "\n".join(lines)
|
||||
|
||||
|
||||
def _connected_refs(ref_slots):
|
||||
"""Connected refs only, preserving slot order and skipping empty sockets."""
|
||||
return [ref for ref in (ref_slots or []) if ref is not None]
|
||||
return [ref for ref in (ref_slots or []) if _reference_image(ref) is not None]
|
||||
|
||||
|
||||
def _plan_scene_refs(plan, scene_index):
|
||||
@@ -4054,7 +4181,13 @@ def _plan_scene_refs(plan, scene_index):
|
||||
if plan is None:
|
||||
return (None,) * _image_nodes._H3_PLAN_IMAGE_SLOTS
|
||||
extracted = _image_nodes.DumasH3PlanExtractSceneImagesNode().extract(plan, scene_index)
|
||||
return tuple(extracted[1:1 + _image_nodes._H3_PLAN_IMAGE_SLOTS])
|
||||
refs = []
|
||||
for slot_number, image in enumerate(
|
||||
extracted[1:1 + _image_nodes._H3_PLAN_IMAGE_SLOTS],
|
||||
1,
|
||||
):
|
||||
refs.append(_reference_slot(image, slot_number) if image is not None else None)
|
||||
return tuple(refs)
|
||||
|
||||
|
||||
def _merge_ref_slots(direct_slots, plan_slots):
|
||||
@@ -5328,18 +5461,18 @@ class H3LongVideos:
|
||||
# the order they are handed that shot.
|
||||
# Refer to socket tags in the prompt if you want a reference bound to
|
||||
# a named character ("Kristy, <Picture 7>, walks in").
|
||||
"ref_image_1": ("IMAGE", {"tooltip": "Reference image <Picture 1> -- identity/appearance "
|
||||
"ref_1": ("REFERENCE", {"tooltip": "Reference object for <Picture 1> -- image plus identity/environment metadata "
|
||||
"carried into the shots. Which shots receive it is set by ref_mode (or <Picture N> "
|
||||
"tags in the beats); a referenced shot ALSO carries the previous frame as its "
|
||||
"keyframe, so taking a reference never costs continuity."}),
|
||||
"ref_image_2": ("IMAGE", {"tooltip": "Reference image <Picture 2>."}),
|
||||
"ref_image_3": ("IMAGE", {"tooltip": "Reference image <Picture 3>."}),
|
||||
"ref_image_4": ("IMAGE", {"tooltip": "Reference image <Picture 4>."}),
|
||||
"ref_image_5": ("IMAGE", {"tooltip": "Reference image <Picture 5>."}),
|
||||
"ref_image_6": ("IMAGE", {"tooltip": "Reference image <Picture 6>."}),
|
||||
"ref_image_7": ("IMAGE", {"tooltip": "Reference image <Picture 7>."}),
|
||||
"ref_image_8": ("IMAGE", {"tooltip": "Reference image <Picture 8>."}),
|
||||
"ref_image_9": ("IMAGE", {"tooltip": "Reference image <Picture 9>."}),
|
||||
"ref_2": ("REFERENCE", {"tooltip": "Reference object for <Picture 2>."}),
|
||||
"ref_3": ("REFERENCE", {"tooltip": "Reference object for <Picture 3>."}),
|
||||
"ref_4": ("REFERENCE", {"tooltip": "Reference object for <Picture 4>."}),
|
||||
"ref_5": ("REFERENCE", {"tooltip": "Reference object for <Picture 5>."}),
|
||||
"ref_6": ("REFERENCE", {"tooltip": "Reference object for <Picture 6>."}),
|
||||
"ref_7": ("REFERENCE", {"tooltip": "Reference object for <Picture 7>."}),
|
||||
"ref_8": ("REFERENCE", {"tooltip": "Reference object for <Picture 8>."}),
|
||||
"ref_9": ("REFERENCE", {"tooltip": "Reference object for <Picture 9>."}),
|
||||
"plan": (
|
||||
"H3_CHAIN_PLAN",
|
||||
{
|
||||
@@ -5356,7 +5489,7 @@ class H3LongVideos:
|
||||
"max": 9999,
|
||||
"step": 1,
|
||||
"tooltip": "1-based plan scene index to read from `plan`. "
|
||||
"Any directly-wired ref_image socket overrides the same slot "
|
||||
"Any directly-wired ref socket overrides the same slot "
|
||||
"from the plan scene."
|
||||
},
|
||||
),
|
||||
@@ -5853,9 +5986,9 @@ class H3LongVideos:
|
||||
watermark_opacity=0.75, watermark_margin=3.0,
|
||||
intro_text="", intro_position="center", intro_seconds=3.0, intro_fade=0.6,
|
||||
intro_size=9.0, overlay_font="arial.ttf", overlay_stroke=0,
|
||||
ref_image_1=None, ref_image_2=None, ref_image_3=None, ref_image_4=None,
|
||||
ref_image_5=None, ref_image_6=None, ref_image_7=None, ref_image_8=None,
|
||||
ref_image_9=None,
|
||||
ref_1=None, ref_2=None, ref_3=None, ref_4=None,
|
||||
ref_5=None, ref_6=None, ref_7=None, ref_8=None,
|
||||
ref_9=None,
|
||||
plan=None, plan_scene_index=1,
|
||||
ref_mode="where tagged", ref_image_size="match", ref_noise_aug=0.999,
|
||||
graph=None, node_id=None):
|
||||
@@ -5878,13 +6011,15 @@ class H3LongVideos:
|
||||
fps = H3_FPS
|
||||
w, h = parse_resolution(resolution)
|
||||
direct_ref_slots = (
|
||||
ref_image_1, ref_image_2, ref_image_3, ref_image_4, ref_image_5,
|
||||
ref_image_6, ref_image_7, ref_image_8, ref_image_9,
|
||||
ref_1, ref_2, ref_3, ref_4, ref_5,
|
||||
ref_6, ref_7, ref_8, ref_9,
|
||||
)
|
||||
plan_ref_slots = _plan_scene_refs(plan, plan_scene_index)
|
||||
ref_slots = _merge_ref_slots(direct_ref_slots, plan_ref_slots)
|
||||
plan_ref_count = len(_connected_refs(plan_ref_slots))
|
||||
direct_ref_count = len(_connected_refs(direct_ref_slots))
|
||||
derived_character_memory = _reference_character_memory(ref_slots)
|
||||
effective_character_memory = (character_memory or "").strip() or derived_character_memory
|
||||
# A pixel budget overrides the preset's SIZE while keeping its aspect ratio,
|
||||
# so the dropdown chooses the shape and this chooses how big. Scaling from
|
||||
# the preset's own dimensions is what makes 1.00MP reproduce each native
|
||||
@@ -5933,7 +6068,7 @@ class H3LongVideos:
|
||||
# rather than losing a shot and the scene text along with it.
|
||||
anchor_note = ""
|
||||
if (not anchor_override.strip()) and paras and \
|
||||
(anchor_contributes_nothing(anchor, character_memory.strip())
|
||||
(anchor_contributes_nothing(anchor, effective_character_memory)
|
||||
or anchor_is_action_beat(anchor, paras[1:])):
|
||||
preview = " ".join(anchor.split())[:60]
|
||||
anchor, beat_paras = "", paras
|
||||
@@ -6080,7 +6215,7 @@ class H3LongVideos:
|
||||
wardrobe_notes = []
|
||||
strip_shots = [] # shots that newly bared a zone -> the NEXT shot starts fresh
|
||||
gens = distribute_generations(anchor, beats, global_soundscape.strip(),
|
||||
non_diegetic_music.strip(), character_memory.strip(),
|
||||
non_diegetic_music.strip(), effective_character_memory,
|
||||
auto_wardrobe, auto_silence_nonspeech, allow_nonspeech_vocals, count_subjects,
|
||||
lora_on, notes_out=wardrobe_notes, auto_props=auto_props,
|
||||
prevent_nudity=prevent_nudity,
|
||||
@@ -6092,6 +6227,18 @@ class H3LongVideos:
|
||||
motion_guard=motion_guard,
|
||||
contact_guard=contact_guard,
|
||||
count_auto=(subject_count_guard == "auto"))
|
||||
enriched_gens = []
|
||||
for block in gens:
|
||||
context = _reference_context_for_text(block, ref_slots)
|
||||
if context:
|
||||
block = re.sub(
|
||||
r"^(\[Generation \d+\]\s*)",
|
||||
lambda m: m.group(1) + context + " ",
|
||||
block,
|
||||
count=1,
|
||||
)
|
||||
enriched_gens.append(block)
|
||||
gens = enriched_gens
|
||||
|
||||
# A scenery beat mid-chain hands the next shot a frame with no people in
|
||||
# it. Both prompts are individually correct, so this is invisible without
|
||||
|
||||
+265
-53
@@ -20,6 +20,7 @@ _DATE_TOKEN_RE = re.compile(r"%date:([^%]+)%")
|
||||
_SERVE_TOKENS = OrderedDict()
|
||||
_SERVE_CAP = 256
|
||||
_H3_PLAN_TYPE = "H3_CHAIN_PLAN"
|
||||
_REFERENCE_TYPE = "REFERENCE"
|
||||
_H3_PLAN_IMAGE_BINDINGS_KEY = "_dumas_scene_image_bindings"
|
||||
_H3_PLAN_IMAGE_BINDINGS = OrderedDict()
|
||||
_H3_PLAN_IMAGE_BINDINGS_CAP = 128
|
||||
@@ -454,6 +455,135 @@ def _ensure_sentence(value):
|
||||
return text
|
||||
|
||||
|
||||
def _slug_like(value):
|
||||
text = _normalize_free_text(value).lower()
|
||||
text = re.sub(r"[^a-z0-9]+", "-", text).strip("-")
|
||||
return text
|
||||
|
||||
|
||||
def _parse_aliases(value):
|
||||
if isinstance(value, (list, tuple)):
|
||||
raw_items = value
|
||||
else:
|
||||
raw_items = re.split(r"[,;\n\r]+", str(value or ""))
|
||||
aliases = []
|
||||
seen = set()
|
||||
for item in raw_items:
|
||||
alias = _normalize_free_text(item)
|
||||
if not alias:
|
||||
continue
|
||||
key = alias.lower()
|
||||
if key in seen:
|
||||
continue
|
||||
seen.add(key)
|
||||
aliases.append(alias)
|
||||
return aliases
|
||||
|
||||
|
||||
def _coerce_picture_id(value):
|
||||
try:
|
||||
number = int(value)
|
||||
except (TypeError, ValueError):
|
||||
return None
|
||||
return number if number > 0 else None
|
||||
|
||||
|
||||
def _reference_id(explicit_id, name, fallback_prefix):
|
||||
explicit = _slug_like(explicit_id)
|
||||
if explicit:
|
||||
return explicit
|
||||
derived = _slug_like(name)
|
||||
if derived:
|
||||
return derived
|
||||
return f"{fallback_prefix}-{uuid.uuid4().hex[:8]}"
|
||||
|
||||
|
||||
def _reference_label(picture_id):
|
||||
return f"<Picture {picture_id}>" if picture_id else ""
|
||||
|
||||
|
||||
def _reference_summary(kind, name, picture_id):
|
||||
label = _reference_label(picture_id)
|
||||
subject = _normalize_free_text(name) or ("character" if kind == "character" else "location")
|
||||
if label:
|
||||
return f"{subject} shown in {label}."
|
||||
return f"{subject} reference."
|
||||
|
||||
|
||||
def make_reference(
|
||||
*,
|
||||
kind,
|
||||
image,
|
||||
explicit_id="",
|
||||
name="",
|
||||
aliases="",
|
||||
picture_id=None,
|
||||
description="",
|
||||
wardrobe="",
|
||||
general="",
|
||||
facts=None,
|
||||
summary="",
|
||||
):
|
||||
normalized_name = _normalize_free_text(name)
|
||||
normalized_aliases = _parse_aliases(aliases)
|
||||
normalized_picture_id = _coerce_picture_id(picture_id)
|
||||
normalized_kind = "location" if str(kind or "").strip().lower() == "location" else "character"
|
||||
normalized_description = _normalize_free_text(description)
|
||||
normalized_wardrobe = _normalize_free_text(wardrobe)
|
||||
normalized_general = _normalize_free_text(general)
|
||||
normalized_facts = dict(facts or {})
|
||||
normalized_summary = _ensure_sentence(
|
||||
summary or _reference_summary(normalized_kind, normalized_name, normalized_picture_id)
|
||||
)
|
||||
return {
|
||||
"kind": normalized_kind,
|
||||
"id": _reference_id(explicit_id, normalized_name, normalized_kind),
|
||||
"name": normalized_name,
|
||||
"aliases": normalized_aliases,
|
||||
"picture_id": normalized_picture_id,
|
||||
"picture_label": _reference_label(normalized_picture_id),
|
||||
"image": image,
|
||||
"summary": normalized_summary,
|
||||
"description": normalized_description,
|
||||
"wardrobe": normalized_wardrobe if normalized_kind == "character" else "",
|
||||
"general": normalized_general,
|
||||
"facts": normalized_facts,
|
||||
}
|
||||
|
||||
|
||||
def normalize_reference(value, picture_id=None, allow_image_fallback=True):
|
||||
if isinstance(value, dict):
|
||||
reference = dict(value)
|
||||
image = reference.get("image")
|
||||
if image is None and allow_image_fallback:
|
||||
image = value
|
||||
reference["image"] = image
|
||||
if picture_id is not None and not reference.get("picture_id"):
|
||||
reference["picture_id"] = _coerce_picture_id(picture_id)
|
||||
reference["picture_label"] = _reference_label(reference.get("picture_id"))
|
||||
reference.setdefault("kind", "character")
|
||||
reference.setdefault("id", _reference_id("", reference.get("name"), reference["kind"]))
|
||||
reference.setdefault("name", "")
|
||||
reference["aliases"] = _parse_aliases(reference.get("aliases"))
|
||||
reference["summary"] = _ensure_sentence(
|
||||
reference.get("summary")
|
||||
or _reference_summary(reference["kind"], reference.get("name"), reference.get("picture_id"))
|
||||
)
|
||||
reference["description"] = _normalize_free_text(reference.get("description"))
|
||||
reference["wardrobe"] = _normalize_free_text(reference.get("wardrobe"))
|
||||
reference["general"] = _normalize_free_text(reference.get("general"))
|
||||
reference["facts"] = dict(reference.get("facts") or {})
|
||||
return reference
|
||||
if not allow_image_fallback:
|
||||
raise TypeError("Expected a REFERENCE object.")
|
||||
return make_reference(
|
||||
kind="character",
|
||||
image=value,
|
||||
picture_id=picture_id,
|
||||
summary="Plan-bound fallback reference.",
|
||||
)
|
||||
|
||||
|
||||
def _parse_positive_int(value):
|
||||
text = str(value or "").strip()
|
||||
if not text:
|
||||
@@ -1114,35 +1244,26 @@ class DumasH3PlanExtractSceneImagesNode:
|
||||
return (passthrough_plan, *images, _connected_image_count(images))
|
||||
|
||||
|
||||
class DumasCharacterHelperNode:
|
||||
class DumasCharacterReferenceNode:
|
||||
DESCRIPTION = (
|
||||
"Build a MiniMax H3-ready character reference prompt and wardrobe sheet "
|
||||
"from two IMAGE sockets plus simple identity fields, while passing both "
|
||||
"images through unchanged."
|
||||
"Build one structured REFERENCE object for a character so H3 can carry "
|
||||
"the image, identity description, wardrobe, and facts through one socket."
|
||||
)
|
||||
RETURN_TYPES = ("IMAGE", "IMAGE", "STRING", "STRING")
|
||||
RETURN_NAMES = ("image1", "image2", "reference_prompt", "wardrobe")
|
||||
FUNCTION = "build_character_text"
|
||||
RETURN_TYPES = (_REFERENCE_TYPE,)
|
||||
RETURN_NAMES = ("reference",)
|
||||
FUNCTION = "build_reference"
|
||||
CATEGORY = "Dumas/MiniMax"
|
||||
|
||||
@classmethod
|
||||
def INPUT_TYPES(cls):
|
||||
return {
|
||||
"required": {
|
||||
"image1": ("IMAGE", {"tooltip": "Primary image to pass through and describe."}),
|
||||
"image2": ("IMAGE", {"tooltip": "Secondary image to pass through and describe."}),
|
||||
"image1_picture_id": (
|
||||
"image": ("IMAGE", {"tooltip": "Character reference image."}),
|
||||
"picture_id": (
|
||||
["1", "2", "3", "4", "5", "6", "7", "8", "9"],
|
||||
{
|
||||
"default": "1",
|
||||
"tooltip": "Picture number to mention for image1 in the H3 reference prompt.",
|
||||
},
|
||||
),
|
||||
"image2_picture_id": (
|
||||
["1", "2", "3", "4", "5", "6", "7", "8", "9"],
|
||||
{
|
||||
"default": "2",
|
||||
"tooltip": "Picture number to mention for image2 in the H3 reference prompt.",
|
||||
"tooltip": "Authoring picture number for this reference.",
|
||||
},
|
||||
),
|
||||
"character_id": (
|
||||
@@ -1150,7 +1271,7 @@ class DumasCharacterHelperNode:
|
||||
{
|
||||
"default": "",
|
||||
"multiline": False,
|
||||
"tooltip": "Optional character ID string to include in the output text.",
|
||||
"tooltip": "Stable internal identifier for the character.",
|
||||
},
|
||||
),
|
||||
"name": (
|
||||
@@ -1158,7 +1279,7 @@ class DumasCharacterHelperNode:
|
||||
{
|
||||
"default": "",
|
||||
"multiline": False,
|
||||
"tooltip": "Character name used in the main reference sentences.",
|
||||
"tooltip": "Human-readable character name.",
|
||||
},
|
||||
),
|
||||
"alias": (
|
||||
@@ -1166,7 +1287,7 @@ class DumasCharacterHelperNode:
|
||||
{
|
||||
"default": "",
|
||||
"multiline": False,
|
||||
"tooltip": "Optional alternate name, codename, or nickname.",
|
||||
"tooltip": "Comma- or newline-separated aliases for name matching.",
|
||||
},
|
||||
),
|
||||
"gender": (
|
||||
@@ -1220,7 +1341,15 @@ class DumasCharacterHelperNode:
|
||||
{
|
||||
"default": "",
|
||||
"multiline": False,
|
||||
"tooltip": "Optional short accent description.",
|
||||
"tooltip": "Optional accent or speaking-style fact.",
|
||||
},
|
||||
),
|
||||
"description": (
|
||||
"STRING",
|
||||
{
|
||||
"default": "",
|
||||
"multiline": True,
|
||||
"tooltip": "Persistent physical identity description for the character.",
|
||||
},
|
||||
),
|
||||
"general": (
|
||||
@@ -1228,7 +1357,7 @@ class DumasCharacterHelperNode:
|
||||
{
|
||||
"default": "",
|
||||
"multiline": True,
|
||||
"tooltip": "Optional non-clothing details appended as the last sentence of the H3 reference prompt.",
|
||||
"tooltip": "Optional freeform notes or extra context.",
|
||||
},
|
||||
),
|
||||
"wardrobe": (
|
||||
@@ -1236,18 +1365,16 @@ class DumasCharacterHelperNode:
|
||||
{
|
||||
"default": "",
|
||||
"multiline": True,
|
||||
"tooltip": "Optional H3 wardrobe/channel text. Plain clothing lists are auto-wrapped as 'Name = ...' when a name, alias, or character ID is present.",
|
||||
"tooltip": "Persistent clothing, styling, accessories, or look notes.",
|
||||
},
|
||||
),
|
||||
}
|
||||
}
|
||||
|
||||
def build_character_text(
|
||||
def build_reference(
|
||||
self,
|
||||
image1,
|
||||
image2,
|
||||
image1_picture_id,
|
||||
image2_picture_id,
|
||||
image,
|
||||
picture_id,
|
||||
character_id,
|
||||
name,
|
||||
alias,
|
||||
@@ -1258,31 +1385,112 @@ class DumasCharacterHelperNode:
|
||||
height_feet,
|
||||
height_inches,
|
||||
accent,
|
||||
description,
|
||||
general,
|
||||
wardrobe,
|
||||
):
|
||||
text = _build_character_helper_text(
|
||||
image1_picture_id,
|
||||
image2_picture_id,
|
||||
character_id,
|
||||
name,
|
||||
alias,
|
||||
gender,
|
||||
age,
|
||||
nationality,
|
||||
occupation,
|
||||
height_feet,
|
||||
height_inches,
|
||||
accent,
|
||||
general,
|
||||
reference = make_reference(
|
||||
kind="character",
|
||||
image=image,
|
||||
explicit_id=character_id,
|
||||
name=name,
|
||||
aliases=alias,
|
||||
picture_id=picture_id,
|
||||
description=description,
|
||||
wardrobe=wardrobe,
|
||||
general=general,
|
||||
facts={
|
||||
"gender": _normalize_free_text(gender),
|
||||
"age": str(_parse_positive_int(age) or ""),
|
||||
"nationality": _normalize_free_text(nationality),
|
||||
"occupation": _normalize_free_text(occupation),
|
||||
"height_feet": str(height_feet or "").strip(),
|
||||
"height_inches": str(height_inches or "").strip(),
|
||||
"accent": _normalize_free_text(accent),
|
||||
},
|
||||
)
|
||||
wardrobe_text = _build_character_wardrobe_text(
|
||||
wardrobe,
|
||||
character_id,
|
||||
name,
|
||||
alias,
|
||||
return (reference,)
|
||||
|
||||
|
||||
class DumasLocationReferenceNode:
|
||||
DESCRIPTION = (
|
||||
"Build one structured REFERENCE object for a location or environment so "
|
||||
"H3 can carry the image and environment description through one socket."
|
||||
)
|
||||
RETURN_TYPES = (_REFERENCE_TYPE,)
|
||||
RETURN_NAMES = ("reference",)
|
||||
FUNCTION = "build_reference"
|
||||
CATEGORY = "Dumas/MiniMax"
|
||||
|
||||
@classmethod
|
||||
def INPUT_TYPES(cls):
|
||||
return {
|
||||
"required": {
|
||||
"image": ("IMAGE", {"tooltip": "Location or environment reference image."}),
|
||||
"picture_id": (
|
||||
["1", "2", "3", "4", "5", "6", "7", "8", "9"],
|
||||
{
|
||||
"default": "1",
|
||||
"tooltip": "Authoring picture number for this reference.",
|
||||
},
|
||||
),
|
||||
"location_id": (
|
||||
"STRING",
|
||||
{
|
||||
"default": "",
|
||||
"multiline": False,
|
||||
"tooltip": "Stable internal identifier for the location.",
|
||||
},
|
||||
),
|
||||
"name": (
|
||||
"STRING",
|
||||
{
|
||||
"default": "",
|
||||
"multiline": False,
|
||||
"tooltip": "Human-readable location name.",
|
||||
},
|
||||
),
|
||||
"alias": (
|
||||
"STRING",
|
||||
{
|
||||
"default": "",
|
||||
"multiline": False,
|
||||
"tooltip": "Comma- or newline-separated alternate location names.",
|
||||
},
|
||||
),
|
||||
"description": (
|
||||
"STRING",
|
||||
{
|
||||
"default": "",
|
||||
"multiline": True,
|
||||
"tooltip": "Persistent environment, layout, and atmosphere description.",
|
||||
},
|
||||
),
|
||||
"general": (
|
||||
"STRING",
|
||||
{
|
||||
"default": "",
|
||||
"multiline": True,
|
||||
"tooltip": "Optional freeform location notes.",
|
||||
},
|
||||
),
|
||||
}
|
||||
}
|
||||
|
||||
def build_reference(self, image, picture_id, location_id, name, alias, description, general):
|
||||
return (
|
||||
make_reference(
|
||||
kind="location",
|
||||
image=image,
|
||||
explicit_id=location_id,
|
||||
name=name,
|
||||
aliases=alias,
|
||||
picture_id=picture_id,
|
||||
description=description,
|
||||
general=general,
|
||||
facts={},
|
||||
),
|
||||
)
|
||||
return (image1, image2, text, wardrobe_text)
|
||||
|
||||
|
||||
NODE_CLASS_MAPPINGS = {
|
||||
@@ -1291,8 +1499,10 @@ NODE_CLASS_MAPPINGS = {
|
||||
"DumasLoadImagesFolder": DumasLoadImagesFolderNode,
|
||||
"DumasH3PlanAttachSceneImages": DumasH3PlanAttachSceneImagesNode,
|
||||
"DumasH3PlanExtractSceneImages": DumasH3PlanExtractSceneImagesNode,
|
||||
"DumasCharacterHelper": DumasCharacterHelperNode,
|
||||
"DumasH3CharacterHelper": DumasCharacterHelperNode,
|
||||
"DumasCharacterReference": DumasCharacterReferenceNode,
|
||||
"DumasLocationReference": DumasLocationReferenceNode,
|
||||
"DumasCharacterHelper": DumasCharacterReferenceNode,
|
||||
"DumasH3CharacterHelper": DumasCharacterReferenceNode,
|
||||
}
|
||||
|
||||
NODE_DISPLAY_NAME_MAPPINGS = {
|
||||
@@ -1301,6 +1511,8 @@ NODE_DISPLAY_NAME_MAPPINGS = {
|
||||
"DumasLoadImagesFolder": "Load Images from Folder Dumas",
|
||||
"DumasH3PlanAttachSceneImages": "Dumas H3 Plan Attach Scene Images",
|
||||
"DumasH3PlanExtractSceneImages": "Dumas H3 Plan Extract Scene Images",
|
||||
"DumasCharacterHelper": "Dumas H3 Character Helper",
|
||||
"DumasH3CharacterHelper": "Dumas H3 Character Helper",
|
||||
"DumasCharacterReference": "Dumas Character Reference",
|
||||
"DumasLocationReference": "Dumas Location Reference",
|
||||
"DumasCharacterHelper": "Dumas Character Reference",
|
||||
"DumasH3CharacterHelper": "Dumas Character Reference",
|
||||
}
|
||||
|
||||
@@ -154,9 +154,19 @@ class DumasH3LongVideosHelperTests(unittest.TestCase):
|
||||
self.assertGreaterEqual(fn.cache_info().hits, 2)
|
||||
|
||||
def test_resolve_tagged_refs_preserves_sparse_socket_numbers(self):
|
||||
refs = [None, "img2", None, None, None, None, "img7", None, "img9"]
|
||||
refs = [
|
||||
None,
|
||||
{"kind": "character", "image": "img2", "name": "Jon"},
|
||||
None,
|
||||
None,
|
||||
None,
|
||||
None,
|
||||
{"kind": "character", "image": "img7", "name": "Mara"},
|
||||
None,
|
||||
{"kind": "location", "image": "img9", "name": "Watchtower"},
|
||||
]
|
||||
|
||||
text, images, dropped = self.module.resolve_tagged_refs(
|
||||
text, references, dropped = self.module.resolve_tagged_refs(
|
||||
"Mara <Picture 7> turns toward Jon <Picture 2> while <Picture 9> watches.",
|
||||
refs,
|
||||
)
|
||||
@@ -165,42 +175,61 @@ class DumasH3LongVideosHelperTests(unittest.TestCase):
|
||||
text,
|
||||
"Mara <Picture 2> turns toward Jon <Picture 1> while <Picture 3> watches.",
|
||||
)
|
||||
self.assertEqual(images, ["img2", "img7", "img9"])
|
||||
self.assertEqual(
|
||||
[self.module._reference_image(ref) for ref in references],
|
||||
["img2", "img7", "img9"],
|
||||
)
|
||||
self.assertEqual(dropped, [])
|
||||
|
||||
def test_resolve_tagged_refs_drops_unconnected_sparse_slots(self):
|
||||
refs = [None, "img2", None, None, None, None, "img7", None, None]
|
||||
refs = [
|
||||
None,
|
||||
{"kind": "character", "image": "img2", "name": "Jon"},
|
||||
None,
|
||||
None,
|
||||
None,
|
||||
None,
|
||||
{"kind": "character", "image": "img7", "name": "Mara"},
|
||||
None,
|
||||
None,
|
||||
]
|
||||
|
||||
text, images, dropped = self.module.resolve_tagged_refs(
|
||||
text, references, dropped = self.module.resolve_tagged_refs(
|
||||
"Use <Picture 7>, skip <Picture 4>, keep <Picture 2>.",
|
||||
refs,
|
||||
)
|
||||
|
||||
self.assertEqual(text, "Use <Picture 2>, skip, keep <Picture 1>.")
|
||||
self.assertEqual(images, ["img2", "img7"])
|
||||
self.assertEqual(
|
||||
[self.module._reference_image(ref) for ref in references],
|
||||
["img2", "img7"],
|
||||
)
|
||||
self.assertEqual(dropped, [4])
|
||||
|
||||
def test_shot_references_uses_all_connected_sparse_slots(self):
|
||||
refs = [None, "img2", None, "img4", None, None, "img7", None, None]
|
||||
refs = [
|
||||
None,
|
||||
{"kind": "character", "image": "img2"},
|
||||
None,
|
||||
{"kind": "character", "image": "img4"},
|
||||
None,
|
||||
None,
|
||||
{"kind": "location", "image": "img7"},
|
||||
None,
|
||||
None,
|
||||
]
|
||||
|
||||
self.assertEqual(
|
||||
self.module.shot_references(refs, "auto ref2v", 0, None),
|
||||
["img2", "img4", "img7"],
|
||||
)
|
||||
self.assertEqual(
|
||||
self.module.shot_references(refs, "first shot", 0, None),
|
||||
["img2", "img4", "img7"],
|
||||
)
|
||||
self.assertEqual(
|
||||
self.module.shot_references(refs, "every shot", 3, None),
|
||||
["img2", "img4", "img7"],
|
||||
)
|
||||
for mode, shot_index in (("auto ref2v", 0), ("first shot", 0), ("every shot", 3)):
|
||||
self.assertEqual(
|
||||
[self.module._reference_image(ref) for ref in self.module.shot_references(refs, mode, shot_index, None)],
|
||||
["img2", "img4", "img7"],
|
||||
)
|
||||
|
||||
def test_input_types_expose_nine_ref_slots(self):
|
||||
optional = self.module.H3LongVideos.INPUT_TYPES()["optional"]
|
||||
|
||||
for index in range(1, 10):
|
||||
self.assertIn(f"ref_image_{index}", optional)
|
||||
self.assertIn(f"ref_{index}", optional)
|
||||
self.assertIn("plan", optional)
|
||||
self.assertIn("plan_scene_index", optional)
|
||||
|
||||
@@ -215,18 +244,64 @@ class DumasH3LongVideosHelperTests(unittest.TestCase):
|
||||
|
||||
self.assertEqual(len(refs), 9)
|
||||
self.assertIsNone(refs[0])
|
||||
self.assertIs(refs[1], image2)
|
||||
self.assertIs(refs[6], image7)
|
||||
self.assertIs(self.module._reference_image(refs[1]), image2)
|
||||
self.assertIs(self.module._reference_image(refs[6]), image7)
|
||||
|
||||
def test_merge_ref_slots_prefers_direct_refs_over_plan_refs(self):
|
||||
merged = self.module._merge_ref_slots(
|
||||
(None, "direct2", None, None, "direct5", None, None, None, None),
|
||||
("plan1", "plan2", "plan3", None, "plan5", None, "plan7", None, None),
|
||||
(
|
||||
None,
|
||||
{"kind": "character", "image": "direct2"},
|
||||
None,
|
||||
None,
|
||||
{"kind": "character", "image": "direct5"},
|
||||
None,
|
||||
None,
|
||||
None,
|
||||
None,
|
||||
),
|
||||
(
|
||||
{"kind": "location", "image": "plan1"},
|
||||
{"kind": "location", "image": "plan2"},
|
||||
{"kind": "location", "image": "plan3"},
|
||||
None,
|
||||
{"kind": "location", "image": "plan5"},
|
||||
None,
|
||||
{"kind": "location", "image": "plan7"},
|
||||
None,
|
||||
None,
|
||||
),
|
||||
)
|
||||
|
||||
self.assertEqual(
|
||||
merged,
|
||||
("plan1", "direct2", "plan3", None, "direct5", None, "plan7", None, None),
|
||||
[self.module._reference_image(ref) if ref is not None else None for ref in merged],
|
||||
["plan1", "direct2", "plan3", None, "direct5", None, "plan7", None, None],
|
||||
)
|
||||
|
||||
def test_reference_context_matches_character_names_and_location_tags(self):
|
||||
refs = [
|
||||
{"kind": "character", "image": "img1", "name": "Mara", "description": "silver hair", "wardrobe": "red jacket"},
|
||||
{"kind": "location", "image": "img2", "name": "Hangar", "description": "wet concrete floor"},
|
||||
]
|
||||
|
||||
context = self.module._reference_context_for_text(
|
||||
"[Generation 1] Mara crosses the room toward <Picture 2>.",
|
||||
refs,
|
||||
)
|
||||
|
||||
self.assertIn("Persistent appearance for Mara: silver hair.", context)
|
||||
self.assertIn("Persistent wardrobe/style for Mara: red jacket.", context)
|
||||
self.assertIn("Location context for Hangar: wet concrete floor.", context)
|
||||
|
||||
def test_reference_character_memory_uses_character_wardrobe_only(self):
|
||||
refs = [
|
||||
{"kind": "character", "image": "img1", "name": "Mara", "wardrobe": "red jacket, black boots"},
|
||||
{"kind": "location", "image": "img2", "name": "Hangar", "description": "wet concrete floor", "wardrobe": "should be ignored"},
|
||||
]
|
||||
|
||||
self.assertEqual(
|
||||
self.module._reference_character_memory(refs),
|
||||
"Mara = red jacket, black boots",
|
||||
)
|
||||
|
||||
def test_ref_mode_defaults_are_ref2v_biased(self):
|
||||
|
||||
@@ -239,16 +239,13 @@ class DumasImageNodeTests(unittest.TestCase):
|
||||
saved_path = FakePILImage.saved_paths[0][0]
|
||||
self.assertTrue(os.path.isdir(os.path.dirname(saved_path)))
|
||||
|
||||
def test_character_helper_passes_through_images_and_formats_text(self):
|
||||
node = self.image_nodes.DumasCharacterHelperNode()
|
||||
image1 = FakeTensorBatch()
|
||||
image2 = FakeTensorBatch()
|
||||
def test_character_reference_builds_structured_reference(self):
|
||||
node = self.image_nodes.DumasCharacterReferenceNode()
|
||||
image = FakeTensorBatch()
|
||||
|
||||
result = node.build_character_text(
|
||||
image1=image1,
|
||||
image2=image2,
|
||||
image1_picture_id="2",
|
||||
image2_picture_id="3",
|
||||
result = node.build_reference(
|
||||
image=image,
|
||||
picture_id="2",
|
||||
character_id="char_dave",
|
||||
name="Dave",
|
||||
alias="The Locksmith",
|
||||
@@ -259,37 +256,46 @@ class DumasImageNodeTests(unittest.TestCase):
|
||||
height_feet="6",
|
||||
height_inches="2",
|
||||
accent="English",
|
||||
description="Square jaw, tired eyes, cropped brown hair.",
|
||||
general="wears a long grey coat",
|
||||
wardrobe="weathered red flight jacket, grey cargo shorts, black boots",
|
||||
)
|
||||
|
||||
self.assertIs(result[0], image1)
|
||||
self.assertIs(result[1], image2)
|
||||
reference = result[0]
|
||||
self.assertIs(reference["image"], image)
|
||||
self.assertEqual(
|
||||
result[2],
|
||||
(
|
||||
"<Picture 2> and <Picture 3> reference the same character who is called Dave.\n"
|
||||
"<Picture 2> is the primary full-body reference for Dave.\n"
|
||||
"<Picture 3> is a frontal facial reference for Dave.\n"
|
||||
"Dave is also known as The Locksmith, is male, is 41 years old, is English, works as a detective, is 6 feet 2 inches tall, speaks in an English accent.\n"
|
||||
"wears a long grey coat."
|
||||
),
|
||||
)
|
||||
self.assertEqual(
|
||||
result[3],
|
||||
"Dave = weathered red flight jacket, grey cargo shorts, black boots",
|
||||
reference,
|
||||
{
|
||||
"kind": "character",
|
||||
"id": "char-dave",
|
||||
"name": "Dave",
|
||||
"aliases": ["The Locksmith"],
|
||||
"picture_id": 2,
|
||||
"picture_label": "<Picture 2>",
|
||||
"image": image,
|
||||
"summary": "Dave shown in <Picture 2>.",
|
||||
"description": "Square jaw, tired eyes, cropped brown hair.",
|
||||
"wardrobe": "weathered red flight jacket, grey cargo shorts, black boots",
|
||||
"general": "wears a long grey coat",
|
||||
"facts": {
|
||||
"gender": "male",
|
||||
"age": "41",
|
||||
"nationality": "English",
|
||||
"occupation": "a detective",
|
||||
"height_feet": "6",
|
||||
"height_inches": "2",
|
||||
"accent": "English",
|
||||
},
|
||||
},
|
||||
)
|
||||
|
||||
def test_character_helper_handles_missing_optional_fields(self):
|
||||
node = self.image_nodes.DumasCharacterHelperNode()
|
||||
image1 = FakeTensorBatch()
|
||||
image2 = FakeTensorBatch()
|
||||
def test_character_reference_handles_missing_optional_fields(self):
|
||||
node = self.image_nodes.DumasCharacterReferenceNode()
|
||||
image = FakeTensorBatch()
|
||||
|
||||
result = node.build_character_text(
|
||||
image1=image1,
|
||||
image2=image2,
|
||||
image1_picture_id="4",
|
||||
image2_picture_id="6",
|
||||
result = node.build_reference(
|
||||
image=image,
|
||||
picture_id="4",
|
||||
character_id="",
|
||||
name="",
|
||||
alias="",
|
||||
@@ -300,53 +306,50 @@ class DumasImageNodeTests(unittest.TestCase):
|
||||
height_feet="",
|
||||
height_inches="",
|
||||
accent="",
|
||||
description="",
|
||||
general="",
|
||||
wardrobe="",
|
||||
)
|
||||
|
||||
self.assertEqual(
|
||||
result[2],
|
||||
(
|
||||
"<Picture 4> and <Picture 6> reference the same character.\n"
|
||||
"<Picture 4> is the primary full-body reference for the character.\n"
|
||||
"<Picture 6> is a frontal facial reference for the character."
|
||||
),
|
||||
)
|
||||
self.assertEqual(result[3], "")
|
||||
reference = result[0]
|
||||
self.assertEqual(reference["kind"], "character")
|
||||
self.assertEqual(reference["picture_id"], 4)
|
||||
self.assertEqual(reference["picture_label"], "<Picture 4>")
|
||||
self.assertEqual(reference["wardrobe"], "")
|
||||
self.assertEqual(reference["general"], "")
|
||||
self.assertEqual(reference["facts"]["age"], "")
|
||||
|
||||
def test_character_helper_allows_picture_one_and_preserves_full_sheet_wardrobe(self):
|
||||
node = self.image_nodes.DumasCharacterHelperNode()
|
||||
image1 = FakeTensorBatch()
|
||||
image2 = FakeTensorBatch()
|
||||
def test_location_reference_builds_structured_reference(self):
|
||||
node = self.image_nodes.DumasLocationReferenceNode()
|
||||
image = FakeTensorBatch()
|
||||
|
||||
result = node.build_character_text(
|
||||
image1=image1,
|
||||
image2=image2,
|
||||
image1_picture_id="1",
|
||||
image2_picture_id="9",
|
||||
character_id="char_kristy",
|
||||
name="Kristy",
|
||||
alias="",
|
||||
gender="",
|
||||
age="",
|
||||
nationality="",
|
||||
occupation="",
|
||||
height_feet="",
|
||||
height_inches="",
|
||||
accent="",
|
||||
general="",
|
||||
wardrobe="Kristy = black coat, silver boots",
|
||||
result = node.build_reference(
|
||||
image=image,
|
||||
picture_id="9",
|
||||
location_id="coffee-shop-01",
|
||||
name="Coffee Shop",
|
||||
alias="Cafe Interior",
|
||||
description="Warm tungsten lighting, narrow counter, rainy front window.",
|
||||
general="Evening ambience, cramped but cozy.",
|
||||
)
|
||||
|
||||
self.assertEqual(
|
||||
result[2],
|
||||
(
|
||||
"<Picture 1> and <Picture 9> reference the same character who is called Kristy.\n"
|
||||
"<Picture 1> is the primary full-body reference for Kristy.\n"
|
||||
"<Picture 9> is a frontal facial reference for Kristy."
|
||||
),
|
||||
result[0],
|
||||
{
|
||||
"kind": "location",
|
||||
"id": "coffee-shop-01",
|
||||
"name": "Coffee Shop",
|
||||
"aliases": ["Cafe Interior"],
|
||||
"picture_id": 9,
|
||||
"picture_label": "<Picture 9>",
|
||||
"image": image,
|
||||
"summary": "Coffee Shop shown in <Picture 9>.",
|
||||
"description": "Warm tungsten lighting, narrow counter, rainy front window.",
|
||||
"wardrobe": "",
|
||||
"general": "Evening ambience, cramped but cozy.",
|
||||
"facts": {},
|
||||
},
|
||||
)
|
||||
self.assertEqual(result[3], "Kristy = black coat, silver boots")
|
||||
|
||||
def test_save_image_returns_ui_entries_for_output_folder(self):
|
||||
node = self.image_nodes.DumasSaveImageNode()
|
||||
|
||||
Reference in New Issue
Block a user