8.9 KiB
Dumas Reference System Spec
Goal
Replace the current image-only / text-fragment character-reference flow with one structured reference system that works for both:
- characters
- locations / environments
Compatibility with old graphs is not a goal. This is a clean redesign for a single-user workflow.
Core Direction
The system should stop treating a reference as "just an image".
A reference should be one structured object that carries:
- the image used for model conditioning
- the semantic identity of the reference
- the descriptive text tied to that reference
- the routing metadata needed by H3 beat / prompt logic
This object type is called REFERENCE.
Main Outcomes
After this redesign:
- H3 ref sockets accept
REFERENCE, not rawIMAGE <Picture N>resolves to the boundREFERENCEobject- character references can provide identity, description, and wardrobe from one source
- location references can provide environment description from one source
- beat prompt logic no longer has to guess where a name or picture tag should pull its descriptive payload from
Reference Kinds
REFERENCE is one base type with a kind field.
Initial supported kinds:
characterlocation
Future kinds could be added later if useful, but they are out of scope for the first pass.
Reference Object Schema
The REFERENCE object should be a plain JSON-serializable structure plus an
attached Comfy image payload where needed by downstream nodes.
Conceptual schema:
{
"kind": "character",
"id": "francine",
"name": "Francine",
"aliases": [],
"picture_id": 1,
"picture_label": "<Picture 1>",
"image": "<IMAGE>",
"summary": "Same character shown in <Picture 1>.",
"description": "Face, hair, proportions, persistent visual identity.",
"wardrobe": "Clothing / styling / persistent look notes.",
"general": "Freeform notes.",
"facts": {
"gender": "",
"age": "",
"nationality": "",
"occupation": "",
"height_feet": "",
"height_inches": "",
"accent": ""
}
}
For locations:
{
"kind": "location",
"id": "coffee-shop-01",
"name": "Coffee Shop",
"aliases": [],
"picture_id": 2,
"picture_label": "<Picture 2>",
"image": "<IMAGE>",
"summary": "Environment shown in <Picture 2>.",
"description": "Persistent environment / layout / atmosphere description.",
"wardrobe": "",
"general": "Freeform notes.",
"facts": {}
}
Field Meaning
Shared fields
-
kind- Required.
characterorlocation.
-
id- Required.
- Stable, slug-like internal identifier.
- Used for matching and future persistence.
-
name- Optional but strongly recommended.
- Human-readable label.
-
aliases- Optional list of alternate match names.
-
picture_id- Optional integer representing the intended
<Picture N>identity. - This is authoring metadata, not the final socket position.
- Optional integer representing the intended
-
picture_label- Derived convenience text like
<Picture 1>.
- Derived convenience text like
-
image- Required.
- The actual Comfy
IMAGEpayload used for reference conditioning.
-
summary- Short purpose statement for prompt assembly and debugging.
-
description- Primary descriptive payload tied to the reference.
- For characters this is the physical identity description.
- For locations this is the environment/layout/atmosphere description.
-
general- Optional freeform notes.
-
facts- Optional structured detail map.
Character-only field
wardrobe- Optional but expected for character references.
- Persistent clothing / styling / accessories / visual state notes.
Location usage note
For kind = "location", wardrobe should be empty.
Producer Nodes
1. Character helper replacement
Current Dumas Character Helper should be redesigned into a character reference
builder.
Recommended name:
Dumas Character Reference
Inputs:
imagepicture_idcharacter_idnamealiasgenderagenationalityoccupationheight_feetheight_inchesaccentdescriptionwardrobegeneral
Outputs:
reference(REFERENCE)
Notes:
- This node should stop emitting the current mixed output pattern.
- It should build one clean
REFERENCEobject withkind = "character".
2. Location / environment helper
Add a new builder node.
Recommended name:
Dumas Location Reference
Inputs:
imagepicture_idlocation_idnamealiasdescriptiongeneral
Outputs:
reference(REFERENCE)
Notes:
- This node builds one
REFERENCEobject withkind = "location". - It should not attempt to mimic character-specific fields.
Consumer Changes
H3 Long Videos
Current numbered sockets:
ref_image_1..ref_image_9
should become:
ref_1..ref_9
Type:
REFERENCE
Each socket should provide access to:
- the underlying
image - the structured metadata
Internal handling
The H3 node should internally derive two parallel lanes:
- reference images for model conditioning
- reference definitions for semantic beat / prompt resolution
It must no longer depend on detached text sources to know what a reference means.
Beat / prompt resolution rules
Resolution by <Picture N>
If a beat contains <Picture N>:
- map
Nto the bound socket position in that shot - resolve the bound
REFERENCE - use the resolved object's data when generating the shot-level descriptive payload
Resolution by character name
For kind = "character":
- match against
name - match against
aliases - optionally match
id
When a character is matched:
- pull identity text from
description - pull clothing / styling from
wardrobe - pull additional context from
generalandfactsif needed
Resolution by location
For kind = "location":
<Picture N>should resolve directly- later name-based location matching can be added if useful, but the first pass does not need aggressive free-text environment matching
When a location is matched:
- pull environment text from
description - optionally use
general
Prompt Assembly Expectations
Character refs should contribute:
- who the subject is
- what persistent visual identity should be preserved
- what clothing / styling should persist
Location refs should contribute:
- where the scene is
- what environment layout / mood / persistent scenic identity should be preserved
The node should keep these conceptually separate.
Character reference text should not be mistaken for environment text. Location reference text should not be mistaken for wardrobe or identity text.
Multi-reference Behavior
The system must support multiple references bound at once.
Typical cases:
- one character + one location
- multiple character references
- one character sheet reused across many beats
The H3 node should preserve current numbered socket behavior for explicit routing,
but with REFERENCE objects replacing plain images.
Plan / scene integration
The current H3 plan scene image system should evolve in parallel.
Where plan scenes currently carry image bundles, they should eventually carry reference bundles instead.
This is not required for the first code pass if it would enlarge the change too much, but the implementation should avoid boxing us into image-only assumptions.
Non-goals
Not required for first pass:
- backward compatibility with old image-only graphs
- migration helpers
- automatic graph conversion
- advanced location name matching
- persistent storage outside the graph object itself
Recommended Implementation Order
- Define the
REFERENCEtype contract in code. - Replace
Dumas Character Helperwith a character-reference builder node. - Add
Dumas Location Reference. - Update H3 long videos sockets from
IMAGEtoREFERENCE. - Update internal ref collection logic to extract both image and metadata.
- Update beat / prompt resolution to use
REFERENCEobjects. - Update tests.
- Update README and workflow docs.
Testing Requirements
Add tests for:
- character reference object construction
- location reference object construction
<Picture N>resolving to the correctREFERENCE- name-based character lookup resolving to the correct
REFERENCE - wardrobe text being pulled only from character refs
- location description being pulled only from location refs
- mixed character + location ref usage in one prompt
- sparse numbered sockets still resolving correctly
Open Design Choices
These should be decided during implementation, but the default recommendation is:
- one
REFERENCEoutput per helper node - numbered H3 sockets remain explicit
- no separate cast object in first pass
- no compatibility fallback path
Summary
This redesign turns references into first-class structured inputs.
The key rule is:
A reference socket should carry both the image and the meaning of that image.
That is the change that removes the current ambiguity between:
- image conditioning
- character identity
- wardrobe data
- location / environment description