-
Notifications
You must be signed in to change notification settings - Fork 0
Expand file tree
/
Copy pathfeed.json
More file actions
166 lines (166 loc) · 353 KB
/
Copy pathfeed.json
File metadata and controls
166 lines (166 loc) · 353 KB
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
{
"version": "https://jsonfeed.org/version/1",
"title": "graphicdesigngeek.com",
"description": "",
"home_page_url": "https://graphicdesigngeek.com",
"feed_url": "https://graphicdesigngeek.com/feed.json",
"user_comment": "",
"author": {
"name": "admin"
},
"items": [
{
"id": "https://graphicdesigngeek.com/minimax-h3-study.html",
"url": "https://graphicdesigngeek.com/minimax-h3-study.html",
"title": "MiniMax H3 study",
"summary": "MiniMax H3 on a Budget: What Actually Works on 4070/5070/5080 (community input consolidation) Discussion A Note on Sources This article is built entirely from community…",
"content_html": "<div class=\"flex\">\n<div class=\"flex flex-col grow max-w-full\">\n<h1 id=\"post-title-t3_1vmprjh\" class=\"text-neutral-content-strong m-0 font-semibold text-18 xs:text-24 mb-xs px-md xs:px-0 xs:mb-md overflow-hidden\" dir=\"auto\" aria-label=\"Post Title: MiniMax H3 on a Budget: What Actually Works on 4070/5070/5080 (community input consolidation)\" aria-describedby=\"feed-post-credit-bar-t3_1vmprjh\">MiniMax H3 on a Budget: What Actually Works on 4070/5070/5080 (community input consolidation)</h1>\n<div class=\"mb-2xs px-md xs:px-0\"><a href=\"https://www.reddit.com/r/StableDiffusion/?f=flair_name%3A%22Discussion%22\" class=\"no-decoration\" actioned=\"\">\n<div class=\"flair-content [&_.flair-image]:align-bottom max-w-full overflow-hidden whitespace-nowrap text-ellipsis\" dir=\"auto\" aria-label=\"Flair: Discussion\">Discussion</div>\n</a></div>\n</div>\n<div> </div>\n</div>\n<div class=\"text-neutral-content\">\n<div class=\"mb-xs px-md xs:px-0 overflow-hidden\" data-post-click-location=\"text-body\">\n<div id=\"t3_1vmprjh-post-rtjson-content\" class=\"md text-14-scalable pb-2xs [--emote-size:20px]\" dir=\"auto\">\n<h1 class=\"text-24-scalable xs:text-20-scalable\">A Note on Sources</h1>\n<p>This article is built entirely from community feedback — Reddit threads, forum comments, and one independent comparison site (jo-nike.github.io/h3-turbo-eval). None of it comes from official documentation or controlled lab testing. Thank you to everyone whose posts, benchmarks, and hard-won troubleshooting notes made this possible, including GrayingGamer, Tystros, Chemical-Painter-485, katsura_otoko, infearia, JoNike, Sixhaunt, dtdisapointingresult, Snoo_64233, mellowanon, Just1Dev, smereces, DefloN92, StuffProfessional587, Creative_Finger_69, backworld_nograv, V4nKw15h, True_Protection6842, clex55, Maskwi2, Perfect-Campaign9551, and many others whose usernames didn't make it into these notes but whose comments shaped the consensus (and disagreements) captured here.</p>\n<p>Where the community disagreed with itself, that's presented as an open question rather than resolved — and where direct data for a specific card was simply missing, that gap is called out rather than papered over.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">Why This Is Confusing</h1>\n<p>Most of the detailed benchmarking in the MiniMax H3 community comes from people with RTX 3090s, 4090s, and 5090s — cards with 24GB+ VRAM that can afford to just try everything and report back. If you're on a 4070, 5070, or 5080, you're stuck reverse-engineering advice that wasn't written with your VRAM ceiling in mind. This piece pulls together what budget-card owners actually reported, plus what reasonably carries over from adjacent cards where direct data doesn't exist.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">The Three (and a Half) Speed Levers</h1>\n<p>Every thread assumes you already know these, so here's the plain version:</p>\n<ul>\n<li>\n<p><strong>Turbo LoRAs</strong> — swap-in models trained to produce good results in far fewer steps (4-8 instead of 20-32). Fastest option, but quality cost varies a lot depending on which checkpoint version you use.</p>\n</li>\n<li>\n<p><strong>Spectrum</strong> — a node that mathematically forecasts/predicts future denoising steps instead of computing them. Counterintuitively, it needs <em>more</em> steps to work well — it's not a low-step tool.</p>\n</li>\n<li>\n<p><strong>Sage Attention</strong> — an attention backend swap. Broad community agreement that this is close to \"free\" speed with minimal quality loss, and it's the one piece almost nobody argues against.</p>\n</li>\n<li>\n<p><strong>EasyCache</strong> — a quieter fourth option that came up as a serious alternative to Turbo LoRAs for drafting, not just a bonus add-on.</p>\n</li>\n</ul>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">What \"Budget\" Card Owners Actually Reported</h1>\n<p>This is the thin part of the record, so treat it as ground truth before anything else:</p>\n<ul>\n<li>\n<p><strong>RTX 4070 (12GB, 32GB RAM):</strong> did quick 0.3MP draft passes in a couple of minutes to tweak prompts and hunt for seeds, reserving longer ~40-minute runs for higher resolution/duration finals. VRAM was sufficient for T2V-style work specifically.</p>\n</li>\n<li>\n<p><strong>RTX 4070 Ti Super (16GB, 32GB RAM):</strong> reported working well, no further detail given.</p>\n</li>\n<li>\n<p><strong>RTX 5070 Ti (16GB, 32GB DDR4):</strong> upgrading from an RTX 2060 (6GB) described the speed difference as \"night and day\" — notably, <em>without</em> any Sage Attention or acceleration nodes running yet. This suggests raw generational/VRAM gains matter a lot on their own, before you even add speed tricks.</p>\n</li>\n<li>\n<p><strong>Warning flag for all of the above:</strong> reference-heavy Ref2V generation was specifically called \"brutal\" on modest VRAM cards, compared to plain T2V. If your workflow uses multiple reference images/videos, expect more friction than these numbers suggest.</p>\n</li>\n</ul>\n<p><strong>Gap, named honestly:</strong> there's no direct plain-5070 or 5080 speed benchmark in any of the source threads. The one 5080 comment that exists is qualitative (\"still great,\" runs the BF16 pruned model fine) with no timing numbers.</p>\n<p><strong>Extrapolation (clearly labeled):</strong> Since the 5070 Ti (16GB) and 4070 Ti Super (16GB) both reported comfortable results, and RTX-series cards were noted to benefit meaningfully from tensor cores over older architectures, a plain 5070 (12GB) likely lands closer to the 4070's experience — fine for T2V and quick low-res drafts, tighter on Ref2V with multiple references. A 5080 (16GB) likely performs at least as well as the 4070 Ti Super, probably closer to the low end of what 3090 owners report, given the VRAM parity and newer architecture. <strong>This is inference from adjacent data, not a report anyone actually made</strong> — treat it as a starting assumption to test, not a promise.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">The Draft → Final Two-Stage Workflow</h1>\n<p>This is the one thing nearly every thread converges on independently, and it's probably the most actionable takeaway for a budget card:</p>\n<p><strong>Draft stage</strong> (fast iteration, hunting for the right prompt/seed):</p>\n<ul>\n<li>\n<p>Low resolution: 0.2–0.4 megapixels</p>\n</li>\n<li>\n<p>Low steps: 8–13</p>\n</li>\n<li>\n<p>Acceleration: either a Turbo LoRA <em>or</em> EasyCache (not both)</p>\n</li>\n<li>\n<p>Faster VAE decode substitute: BlehTAEVideoDecode instead of the standard node</p>\n</li>\n</ul>\n<p><strong>Final stage</strong> (once the shot is locked):</p>\n<ul>\n<li>\n<p>Disable acceleration nodes</p>\n</li>\n<li>\n<p>Raise steps to 20–32</p>\n</li>\n<li>\n<p>Switch back to the standard VAE Decode node</p>\n</li>\n</ul>\n<p>Two draft \"recipes\" show up repeatedly and are reported as similarly fast:</p>\n<ol>\n<li>\n<p><strong>Turbo LoRA + Sage Attention</strong> — faster to set up, more established</p>\n</li>\n<li>\n<p><strong>Sage Attention + EasyCache</strong>, params (0.3, 0.2, 0.9), res_multistep sampler + Simple scheduler — one detailed user report (RTX 4060 Ti, 16GB), after testing 1000+ variations, said this drifts <em>less</em> from final quality than Turbo LoRA approaches, at comparable speed</p>\n</li>\n</ol>\n<p>For a 12–16GB budget card, EasyCache is worth trying first specifically because it avoids the quality-consistency debates that follow Turbo LoRAs (see below).</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">What Worked / What Didn't</h1>\n<table class=\"overflow-x-auto\">\n<thead>\n<tr>\n<th class=\"align-left\">Technique</th>\n<th class=\"align-left\">Verdict</th>\n<th class=\"align-left\">Reported Config</th>\n<th class=\"align-left\">Source Consensus</th>\n</tr>\n<tr></tr>\n</thead>\n<tbody>\n<tr>\n<td class=\"align-left\"><strong>Sage Attention (alone)</strong></td>\n<td class=\"align-left\">✅ Works</td>\n<td class=\"align-left\">Any step count</td>\n<td class=\"align-left\">Broad agreement — near-free speed, minimal quality loss</td>\n</tr>\n<tr>\n<td class=\"align-left\"><strong>Two-stage draft→final workflow</strong></td>\n<td class=\"align-left\">✅ Works</td>\n<td class=\"align-left\">Draft: 0.2–0.4MP, 8–13 steps → Final: 20–32 steps, no acceleration</td>\n<td class=\"align-left\">Converged on independently across nearly every thread</td>\n</tr>\n<tr>\n<td class=\"align-left\"><strong>\"Clean VRAM\" node before VAE Decode</strong></td>\n<td class=\"align-left\">✅ Works</td>\n<td class=\"align-left\">Placement only, no params</td>\n<td class=\"align-left\">Multiple independent reports, fixed OOM with no downsides</td>\n</tr>\n<tr>\n<td class=\"align-left\"><strong>EasyCache (draft)</strong></td>\n<td class=\"align-left\">✅ Works</td>\n<td class=\"align-left\">Params (0.3, 0.2, 0.9), res_multistep + Simple, 10 steps</td>\n<td class=\"align-left\">One deep-dive (1000+ tests) preferred it over turbo LoRAs for drift</td>\n</tr>\n<tr>\n<td class=\"align-left\"><strong>ema-ckpt500 Turbo LoRA</strong></td>\n<td class=\"align-left\">✅ Works</td>\n<td class=\"align-left\">Strength ~0.5, 6–8 steps</td>\n<td class=\"align-left\">Beat both ckpt850 and lightx2v in blind testing</td>\n</tr>\n<tr>\n<td class=\"align-left\"><strong>Spectrum below ~20 steps</strong></td>\n<td class=\"align-left\">❌ Doesn't work</td>\n<td class=\"align-left\">N/A</td>\n<td class=\"align-left\">Most consistent \"don't do this\" finding across all sources</td>\n</tr>\n<tr>\n<td class=\"align-left\"><strong>Spectrum + Turbo LoRA together</strong></td>\n<td class=\"align-left\">❌ Doesn't work</td>\n<td class=\"align-left\">N/A</td>\n<td class=\"align-left\">Explicitly warned against — Spectrum needs clean high-step data</td>\n</tr>\n<tr>\n<td class=\"align-left\"><strong>ckpt850 Turbo LoRA (vs ckpt500)</strong></td>\n<td class=\"align-left\">❌ Doesn't work</td>\n<td class=\"align-left\">Full 1.0 strength = \"overfried\"</td>\n<td class=\"align-left\">Newer checkpoint tested worse than older one, despite official claims</td>\n</tr>\n<tr>\n<td class=\"align-left\"><strong>lightx2v LoRA</strong></td>\n<td class=\"align-left\">❌ Doesn't work</td>\n<td class=\"align-left\">8 steps, 0.75 strength</td>\n<td class=\"align-left\">Worse faces/lighting vs ema-ckpt500 in direct comparison</td>\n</tr>\n<tr>\n<td class=\"align-left\"><strong>Raising steps to fix face-warping</strong></td>\n<td class=\"align-left\">❌ Doesn't work</td>\n<td class=\"align-left\">Tested 8→20, and up to 30 steps</td>\n<td class=\"align-left\">Two separate users found no improvement — not a step-count problem</td>\n</tr>\n<tr>\n<td class=\"align-left\"><strong>Any acceleration on non-RTX cards</strong></td>\n<td class=\"align-left\">❌ Doesn't work</td>\n<td class=\"align-left\">N/A</td>\n<td class=\"align-left\">Tensor-core dependent; gains don't transfer to older architectures</td>\n</tr>\n<tr>\n<td class=\"align-left\"><strong>Turbo LoRAs (general use)</strong></td>\n<td class=\"align-left\">⚠️ Mixed</td>\n<td class=\"align-left\">Fine for tests/talking-head; risky for motion/long prompts</td>\n<td class=\"align-left\">Depends on shot type, not a clean yes/no</td>\n</tr>\n<tr>\n<td class=\"align-left\"><strong>Spectrum + First Block Cache</strong></td>\n<td class=\"align-left\">⚠️ Mixed</td>\n<td class=\"align-left\">N/A</td>\n<td class=\"align-left\">Direct contradiction between two experienced users</td>\n</tr>\n<tr>\n<td class=\"align-left\"><strong>RTX upscaling node</strong></td>\n<td class=\"align-left\">⚠️ Mixed</td>\n<td class=\"align-left\">0.2MP+</td>\n<td class=\"align-left\">Good on animation, unreliable on photorealistic faces</td>\n</tr>\n</tbody>\n</table>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">GPU-Specific Data: Reported vs. Extrapolated</h1>\n<table class=\"overflow-x-auto\">\n<thead>\n<tr>\n<th class=\"align-left\">GPU</th>\n<th class=\"align-left\">VRAM</th>\n<th class=\"align-left\">Reported Result</th>\n<th class=\"align-left\">Status</th>\n</tr>\n<tr></tr>\n</thead>\n<tbody>\n<tr>\n<td class=\"align-left\">RTX 4070</td>\n<td class=\"align-left\">12GB</td>\n<td class=\"align-left\">0.3MP drafts in ~2 min; fine for T2V, tight on Ref2V</td>\n<td class=\"align-left\">Direct report</td>\n</tr>\n<tr>\n<td class=\"align-left\">RTX 4070 Ti Super</td>\n<td class=\"align-left\">16GB</td>\n<td class=\"align-left\">\"Works well\" (no numbers given)</td>\n<td class=\"align-left\">Direct report</td>\n</tr>\n<tr>\n<td class=\"align-left\">RTX 5070 Ti</td>\n<td class=\"align-left\">16GB</td>\n<td class=\"align-left\">Major generational leap even with zero acceleration</td>\n<td class=\"align-left\">Direct report</td>\n</tr>\n<tr>\n<td class=\"align-left\">RTX 5070</td>\n<td class=\"align-left\">12GB</td>\n<td class=\"align-left\"><em>(no data)</em></td>\n<td class=\"align-left\"><strong>Extrapolated</strong> from 4070 — likely similar</td>\n</tr>\n<tr>\n<td class=\"align-left\">RTX 5080</td>\n<td class=\"align-left\">16GB</td>\n<td class=\"align-left\">Handles BF16 pruned model fine (qualitative only)</td>\n<td class=\"align-left\">Direct report (thin) + extrapolated timing</td>\n</tr>\n</tbody>\n</table>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">The Unresolved Debates</h1>\n<p>Worth knowing before you commit to a setup, so you don't over-trust any single comment:</p>\n<ul>\n<li>\n<p><strong>Spectrum below 20 steps?</strong> Most experienced users say no — negligible speed gain, real quality loss. But a few 5090 owners reported <em>no</em> measurable time savings even at higher step counts, with no clear explanation (dismissed by one commenter as \"not using it right\").</p>\n</li>\n<li>\n<p><strong>Which Turbo LoRA checkpoint is actually best?</strong> The lineage went ckpt500 → ckpt850 → ckpt600, with each new version claimed better by its authors. But blind side-by-side testing found ckpt500 at 0.5 strength still beat ckpt850 even at full strength — directly contradicting the official recommendation.</p>\n</li>\n<li>\n<p><strong>Spectrum + First Block Cache together?</strong> One experienced user says combining them is worse than Spectrum alone; another says combining them is the fastest option with no noticeable quality loss. Unresolved.</p>\n</li>\n<li>\n<p><strong>Turbo LoRA strength values:</strong> reports range from 0.5 up to 1.15–1.20 (and one outlier claiming 3.0), so \"strength 1.0\" isn't a safe universal default — it depends on which checkpoint you're using.</p>\n</li>\n</ul>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">VRAM/RAM Troubleshooting Cheat Sheet</h1>\n<p>Fixes that came up repeatedly and matter more when you're VRAM-constrained:</p>\n<ul>\n<li>\n<p>Add a <strong>\"Clean VRAM\" node</strong> immediately before VAE Decode — fixed OOM issues for multiple users.</p>\n</li>\n<li>\n<p><strong>System RAM matters too</strong>, not just VRAM — one user needed to go from 16GB to 48GB total system RAM to stop hitting errors. 16GB system RAM was described by another as \"almost enough.\"</p>\n</li>\n<li>\n<p>Launch ComfyUI with <code>--reserve-vram 2</code> to keep 1-2GB permanently free for system stability, at a small cost to usable VRAM.</p>\n</li>\n<li>\n<p>If Ref2V errors show up on an 8GB VRAM card, don't assume it's a hard VRAM wall first — one such case turned out to be a node-conflict bug, not actually a memory limit.</p>\n</li>\n</ul>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">A Starter Config for Budget Cards</h1>\n<p>Synthesizing the most-corroborated points into one starting recipe (best-guess synthesis, not a benchmarked config):</p>\n<p><strong>Draft pass:</strong> Sage Attention + EasyCache (0.3, 0.2, 0.9) → 10 steps → res_multistep sampler, Simple scheduler → BlehTAEVideoDecode → 0.2–0.3 MP</p>\n<p><strong>Final pass:</strong> Sage Attention only (no EasyCache) → 20–25 steps → standard VAE Decode → 0.4–0.6 MP (push higher only if VRAM allows)</p>\n<p>Skip Spectrum entirely unless you're already comfortable at 25+ steps and have time to test it — it's not built for the low-step, fast-iteration use case a budget card usually needs.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">Sources</h1>\n<p>The most rigorous single data point in this set is the <a rpl=\"\" class=\"relative pointer-events-auto a\n \n \n \n \n underline\n \n cursor-pointer\" href=\"https://jo-nike.github.io/h3-turbo-eval\" rel=\"noopener nofollow ugc\" target=\"_blank\">JoNike Turbo LoRA comparison site</a> — a 10-scene A/B comparison across checkpoint versions, built and documented far more consistently than typical anecdotal Reddit reports.</p>\n<p> </p>\n<p><a href=\"https://www.reddit.com/r/StableDiffusion/comments/1vmprjh/minimax_h3_on_a_budget_what_actually_works_on/\">https://www.reddit.com/r/StableDiffusion/comments/1vmprjh/minimax_h3_on_a_budget_what_actually_works_on/</a></p>\n</div>\n</div>\n</div>",
"author": {
"name": "admin"
},
"tags": [
"minimax",
"comfyui"
],
"date_published": "2026-08-14T09:08:06-04:00",
"date_modified": "2026-08-14T09:08:06-04:00"
},
{
"id": "https://graphicdesigngeek.com/how-to-install-the-triton-package-for-comfyui-portable.html",
"url": "https://graphicdesigngeek.com/how-to-install-the-triton-package-for-comfyui-portable.html",
"title": "How to install the triton package for comfyui portable",
"summary": "To install the Triton package for ComfyUI Portable, open a terminal in your ComfyUI folder and run the command: .\\python_embeded\\python.exe -m pip install -U triton-windows.",
"content_html": "<p>To install the <span class=\"font-semibold\" data-streamdown=\"strong\">Triton package for ComfyUI Portable</span>, open a terminal in your ComfyUI folder and run the command: <code class=\"rounded bg-muted px-1.5 py-0.5 font-mono text-sm\" data-streamdown=\"inline-code\">.\\python_embeded\\python.exe -m pip install -U triton-windows</code>. Make sure to remove any previously installed version of Triton first with <code class=\"rounded bg-muted px-1.5 py-0.5 font-mono text-sm\" data-streamdown=\"inline-code\">.\\python_embeded\\python.exe -m pip uninstall triton</code></p>",
"author": {
"name": "admin"
},
"tags": [
"comfyui"
],
"date_published": "2026-08-14T03:19:20-04:00",
"date_modified": "2026-08-14T03:19:28-04:00"
},
{
"id": "https://graphicdesigngeek.com/minimax-helpful-information.html",
"url": "https://graphicdesigngeek.com/minimax-helpful-information.html",
"title": "minimax helpful information",
"summary": "a visual study of different setups 1: https://jo-nike.github.io/h3-turbo-eval/index.html 2: https://dawidope.github.io/model-comparison/ or https://www.reddit.com/r/StableDiffusion/comments/1vksju3/minimaxh3_local_benchmark_six_optimization_stacks/ 3. https://darkstarrddev.us.ci/ --- What subjects Minimax H3 knows:: https://www.reddit.com/r/StableDiffusion/comments/1vlqth2/what_characters_minimax_h3_knows_part_3_anime/ https://www.reddit.com/r/StableDiffusion/comments/1vkfq50/what_characters_minimax_h3_knows_part_2_videogames/ prompt guides.. T2V/I2V/FL2VA/L2VA guide: https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md…",
"content_html": "<p>a visual study of different setups</p>\n<p>1: <a href=\"https://jo-nike.github.io/h3-turbo-eval/index.html\">https://jo-nike.github.io/h3-turbo-eval/index.html</a></p>\n<p>2: <a rpl=\"\" class=\"relative pointer-events-auto a\n \n \n \n \n underline\n \n cursor-pointer\" href=\"https://dawidope.github.io/model-comparison/\" rel=\"noopener nofollow ugc\" target=\"_blank\">https://dawidope.github.io/model-comparison/</a></p>\n<p>or</p>\n<p><a href=\"https://www.reddit.com/r/StableDiffusion/comments/1vksju3/minimaxh3_local_benchmark_six_optimization_stacks/\">https://www.reddit.com/r/StableDiffusion/comments/1vksju3/minimaxh3_local_benchmark_six_optimization_stacks/</a></p>\n<p>3. <a href=\"https://darkstarrddev.us.ci/\">https://darkstarrddev.us.ci/</a></p>\n<p>---</p>\n<p id=\"post-title-t3_1vlqth2\" class=\"text-neutral-content-strong m-0 font-semibold text-18 xs:text-24 mb-xs px-md xs:px-0 xs:mb-md overflow-hidden\" dir=\"auto\" aria-label=\"Post Title: What Characters Minimax H3 knows - Part 3 - ANIME ACTION EDITION\" aria-describedby=\"feed-post-credit-bar-t3_1vlqth2\">What subjects Minimax H3 knows::</p>\n<p dir=\"auto\" aria-label=\"Post Title: What Characters Minimax H3 knows - Part 3 - ANIME ACTION EDITION\" aria-describedby=\"feed-post-credit-bar-t3_1vlqth2\"><a href=\"https://www.reddit.com/r/StableDiffusion/comments/1vlqth2/what_characters_minimax_h3_knows_part_3_anime/\">https://www.reddit.com/r/StableDiffusion/comments/1vlqth2/what_characters_minimax_h3_knows_part_3_anime/</a></p>\n<p dir=\"auto\" aria-label=\"Post Title: What Characters Minimax H3 knows - Part 3 - ANIME ACTION EDITION\" aria-describedby=\"feed-post-credit-bar-t3_1vlqth2\"><a href=\"https://www.reddit.com/r/StableDiffusion/comments/1vkfq50/what_characters_minimax_h3_knows_part_2_videogames/\">https://www.reddit.com/r/StableDiffusion/comments/1vkfq50/what_characters_minimax_h3_knows_part_2_videogames/</a></p>\n<p>prompt guides..</p>\n<p>T2V/I2V/FL2VA/L2VA guide: <a rpl=\"\" class=\"relative pointer-events-auto a\n \n \n \n \n underline\n \n cursor-pointer\" href=\"https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md\" rel=\"noopener nofollow ugc\" target=\"_blank\">https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md</a></p>\n<p>Ref2V guide: <a rpl=\"\" class=\"relative pointer-events-auto a\n \n \n \n \n underline\n \n cursor-pointer\" href=\"https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md\" rel=\"noopener nofollow ugc\" target=\"_blank\">https://huggingface.co/MiniMaxAI/MiniMax-H3/blob/main/docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md</a></p>\n<p>Bonus - Model built in Skills Guide: <a rpl=\"\" class=\"relative pointer-events-auto a\n \n \n \n \n underline\n \n cursor-pointer\" href=\"https://github.com/MiniMax-AI/MiniMax-H3/tree/main/skills\" rel=\"noopener nofollow ugc\" target=\"_blank\">https://github.com/MiniMax-AI/MiniMax-H3/tree/main/skills</a></p>\n<h1>H3 Ref2V System instruction</h1>\n<p>roles for system instructions</p>\n<p><a href=\"https://pastebin.com/HXsQF50i\">https://pastebin.com/HXsQF50i</a></p>\n<pre class=\"language-json\"><code># ROLE\nYou are an expert prompt writer for the MiniMax Hailuo H3 video model,\nspecializing in full-reference mode (Ref2VA): the user supplies multiple\nreference assets — reusable subjects, images, source videos, and audio — and you\nrewrite the request into a six-section, label-tracked prompt.\n \n# THE USER MESSAGE\nThe user will give you, in free form:\n1. The video concept / what they want to happen.\n2. The target video duration in seconds (if omitted, assume 5.00).\n3. For every reference asset they are attaching, a LABEL and a TEXT DESCRIPTION,\n e.g. `Picture 1: a blonde woman in a light-pink shirt on an orange sofa`.\n \nTreat the user's text descriptions of each reference as the reliable anchor for\nwhat that label contains. If your model can natively perceive the attached asset\n— images on any vision model, and audio/video on a model that ingests them\n(e.g. Gemma 4 12B) — use the asset directly for finer detail, but never\ncontradict the user's description. Never invent a reference the user did not\nprovide, and never leave a referenced label undefined.\n \n# OUTPUT CONTRACT\nOutput ONLY the fields specified below, in the exact order and with the exact\nfield names shown. No preamble, no commentary, no markdown headers, no code\nfences. Write everything in English EXCEPT dialogue/lyrics inside `<d>` and text\nvisibly present in the scene, which stay in their original language. Timing is\n`MM:SS.mmm` for cuts and `S.SS` (two decimals) for the alignment line.\n \n# THE SIX SECTIONS (exact order, exact names)\nsubject_definitions\nsummary\nretention_analysis\ndetailed_description\noverall_soundscape\nnon_diegetic_music\n \n# 1. subject_definitions — reference labels\nFour label types; once assigned, a label keeps the same meaning in every\nsection:\n- `<Subject N>`: reusable VISIBLE content (people, animals, objects, scenes,\n backgrounds, clothing, props, effects, styles, actions, expressions, poses).\n It is the content unit used in the target video, not the source file. One\n subject may come from several assets; one asset may yield several subjects.\n- `<Picture N>`: a reference image used as a concrete frame / keyframe / last\n frame / edited keyframe / composition or storyboard anchor.\n- `<Video N>`: a WHOLE-video relationship — editing a source video, continuing\n from its end, or referencing its camera/cuts/rhythm/temporal structure.\n- `<Audio N>`: a standalone audio asset or an enabled synchronized track from a\n reference video (copying signal, referencing BGM style, voice timbre/delivery,\n reusing dialogue/lyrics/SFX, or beat/continuity).\nGive each separately-tracked item its own line stating what the label denotes,\nits reference role, and the main features to follow. If a `<Picture N>` or\n`<Video N>` only identifies the SOURCE of another item and is not used\nseparately later, cite it INSIDE that item's definition without its own line.\nA person/object/scene/action/effect reused from a video is still a `<Subject N>`\n— `<Video N>` names the asset/structure, not the visible content. An ordinary\nreference video does NOT get an `<Audio N>` just because it has sound.\n`<Video N>` and `<Audio N>` are numbered independently; equal or different\nindices imply nothing about shared source.\nWhen an `<Audio N>` maps to a target speaker, reuse that speaker's GLOBAL id:\n`<Subject N> (Sx)` if it maps to a subject, else a stable voice description plus\n`(Sx)`. The id comes from the target video's global speaker order (Section 5);\nnever assign a new one in the audio definition.\nExamples:\n`<Subject 1> is the young woman in <Picture 1>, with long dark hair, a blue cardigan, and a thin silver necklace.`\n`<Subject 1> is the woman whose appearance comes from <Picture 1> and whose walking motion comes from <Video 1>.`\n`<Picture 3> is a storyboard reference for [Shot 1] and [Shot 2], defining their viewpoint, subject placement, and shot order.`\n`<Video 1> is the source video for the target video edit.`\n`<Audio 1> is the voice-timbre reference for <Subject 1> (S1).`\n \n# 2. summary — one short English paragraph\nBegins with a square-bracketed task-type prefix, then summarizes the target\nvideo and its reference relationships using ONLY already-defined labels (do not\nintroduce new labels here).\nTask types: `keyframe completion` (image as a concrete frame anchor) |\n`reference generation` (image/video/audio guides a character/scene/style/\naction/camera/storyboard without being a concrete frame or the edited/continued\nsource) | `video editing` (an existing source video is directly modified) |\n`video continuation` (new content continues/extends/resumes/transitions from a\nsource video) | `audio reuse` (same signal reused in full or part) |\n`audio reference` (only style/timbre/dialogue/SFX/beat/continuity referenced,\nnot copied).\nCombine multiple with ` + ` and never repeat a type\n(e.g. `[video continuation + keyframe completion]`). Presence of video/audio does\nNOT auto-create a task type: a video giving only camera/cuts/rhythm is\n`reference generation`; use `video editing`/`video continuation` only when that\nvideo is actually edited or continued. For video-editing tasks, start the body\nafter the prefix with `The target video is an edited version of <Video 1>.`\n \n# 3. retention_analysis — one line per label\nPreserve each label's meaning from subject_definitions. Do NOT write `(Sx)` here.\nDo not treat newly added actions/backgrounds/plot as losses of fidelity.\nVisible content (`<Subject N>`, `<Picture N>`, `<Video N>`) uses fixed markers:\n`fully_preserved` | `partially_preserved` | `attribute_transfer` |\n`weak_reference`.\nAudio (`<Audio N>`) uses: `fully_copy` | `partially_copy` | `reference` |\n`weak_reference`.\nEntry forms:\n`<Subject 1> (appears in [Shot 1], [Shot 3]): fully_preserved - ...`\n`<Picture 2> ([Shot 1] first frame): fully_preserved - ...`\n`<Video 1> (cut and pacing structure): weak_reference - ...`\n`<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.`\n`<Audio 2>: reference - the target speaker follows <Audio 2>'s voice timbre and measured delivery without copying the original signal.`\n \n# 4. detailed_description — main body, shot by shot in playback order\nEstablish the overall style in ONE or TWO English sentences BEFORE `[Shot 1]`\n(this is where the style opening lives in full-reference mode — not after\n`[Shot 1]`). Then describe each shot: composition, subject appearance and\nposition, environment and lighting, actions and state changes, camera movement,\ncurrent sound, dialogue, and the exact points where referenced content appears\nor takes effect. Insert `<Subject N>`, `<Picture N>`, `<Video N>`, `<Audio N>`\nat first appearance and wherever their roles apply; keep using the same label\nwithout redefining it. Do not reduce this to a plot summary or a list of\nreference relationships.\nConcrete frame anchors read naturally: `the shot begins from <Picture 1>`,\n`the shot's keyframe corresponds to <Picture 2>`, `the shot ends on <Picture 3>`.\nWhen a referenced subject speaks, keep BOTH the visual label and the speaker id:\n`<Subject 2> (S1) turns toward the woman and says, <d>[English] ...</d>`\n(off-screen: same form marked `off-screen`). Assign `(Sx)` once, by the order of\nactual vocal events in the target video, and reuse it at every vocal event.\nWhen a verbal cue exists only inside a directly reused BGM/soundtrack with no\nindependent vocal source, use `<Audio N>` as the audible source and do NOT\ninvent an `(Sx)`; a concrete person/character/narrator DOES get `(Sx)`.\nWhen dialogue/narration/lyrics from reference audio are directly reused (or the\nuser asks for reperformance), preserve the exact source words and original\nlanguage inside `<d>`; write `[unclear]` for unintelligible spans (never guess);\nstandardize punctuation to `, . ? !`, dropping tildes/emoji/decorative marks and\nending statements/questions/exclamations with `. ? !` before `</d>`. When only\ntimbre/rhythm/emotion/delivery is referenced, do NOT carry the original words\ninto the target video.\nLength: generation tasks are normally 350-500 English words; dialogue-dense\ncontent prioritizes fitting the full spoken timeline over word count; editing\nscales with source complexity. A single shot does not justify a short body —\ndistribute detail by information load.\n \n# 5. overall_soundscape and non_diegetic_music\nDefinitions are the same as the base guide (see CORE WRITING RULES below). State\na reference-audio relationship only in the matching audible layer: ambience/SFX\nin overall_soundscape, audience-only score in non_diegetic_music. If one audio\nprovides both, describe the matching relationship in each section, e.g.:\n`overall_soundscape: The copied ambience layer from <Audio 1> continues throughout the target video.`\n`non_diegetic_music: <Audio 2> is directly reused as the complete audience-only score.`\nWrite full dialogue/lyrics only inside `<d>` in detailed_description; never\nrepeat them in these two sections.\n \n# CORE WRITING RULES (apply to every field)\n \n## Shots and cuts\nDo not put a timestamp on the first shot. Number later shots sequentially and\nbegin each with a strictly increasing cut time inside the duration:\n`[Shot 2] At 00:03.500, the camera cuts to ...`\nFor ordinary cuts use: `the camera cuts to`, `the shot cuts to`,\n`the shot transitions to`, `the shot changes to`, or `the shot switches to`.\nUse cross-dissolve, fade, or wipe only when the user explicitly asks. A cut must\nintroduce new information (subject, space, state, viewpoint, or time); if only\ndistance or a slight angle changes, prefer camera motion instead.\n \n## Camera motion = motion type + amplitude + speed\nWrite camera motion as natural English inside the shot, not as stacked labels.\nAdd amplitude/speed only when meaningful (medium amplitude and normal speed are\nomitted).\nMotion type: Zoom In/Zoom Out (focal length changes, body still) |\nPush In/Pull Out (camera moves forward/back) |\nPan Left/Pan Right (pivots horizontally) |\nTruck Left/Truck Right (translates horizontally) |\nTilt Up/Tilt Down (pivots vertically) |\nPedestal Up/Pedestal Down (whole camera up/down) |\nArc Shot | Tracking Shot | Static Shot |\nShake Slightly/Shake Strongly | POV |\nRoll Clockwise/Roll Counterclockwise.\nAmplitude: `with small amplitude` | `with large amplitude`.\nSpeed: `at slow speed` | `at fast speed`.\nExamples:\n`The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.`\n`The camera pans right with large amplitude at fast speed, revealing the open doorway.`\n`The camera holds a static shot as the runner exits the frame.`\n \n## Speakers, dialogue, singing\nAnyone who speaks, sings, or makes an off-screen human voice gets a stable ID:\n`(S1)`, `(S2)`, ... A speaker keeps the same ID across shots; silent characters\nget no ID. For simultaneous speech use a compound ID like `(S1,S2)`.\nOn first appearance, establish a stable identity (type, age, gender, on/off\nscreen, pitch, timbre, rate, accent). Put the speaker's identifying phrase, ID,\naction, and delivery OUTSIDE `<d>`. Inside `<d>`, put ONLY the language tag and\nthe verbatim user-provided words — never translate or rewrite; preserve every\nword and punctuation mark.\n`The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>`\n`The two children (S1,S2) shout together, <d>[English] Wait for us!</d>`\nFor voiceover use the exact phrase `says in an off-screen voiceover`, and\nimmediately after the `<d>` block state the on-screen character's lips stay\nclosed:\n`The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.`\nWhen one line of dialogue/lyrics crosses a cut, put `<scenetrans>` at the join\nin BOTH parts and state the audio continues across the cut (e.g.\n`continues seamlessly across the cut`, `carries over from the previous shot`).\nUse `<cutoff>` when speech is truncated by the video end.\n \n## On-screen text\nAny banner/sign/label/subtitle/neon actually visible on screen goes in English\ndouble quotes, verbatim, untranslated:\n`A red neon sign reading \"营业中\" glows above the doorway.`\n \n## overall_soundscape\n1-4 English sentences, one paragraph: ambient sound, physical-action sounds,\nnon-verbal human sounds (wind, rain, traffic, footsteps, fabric, impacts,\nbreathing, laughter, panting). Do NOT repeat dialogue, singing, or diegetic\nmusic here. Use `N/A` only if the user explicitly wants full silence.\n \n## non_diegetic_music\n1-3 English sentences describing audience-only background music: instrumentation,\ntempo, rhythm, dynamic changes. No abstract mood words, no emotional-function\nexplanations. Music the characters can hear (singing, instruments, radio, TV,\nphone) is diegetic and belongs in the main description, not here. Use `N/A` when\nthere is no non-diegetic music.\n \n# WORKED EXAMPLE (format reference only — do not copy its content)\nsubject_definitions:\n<Subject 1> is the coffee-shop environment in <Picture 1>, featuring an exposed brick wall, an orange tufted sofa with patterned pillows, a neon sign, and a wooden coffee table.\n<Subject 2> is the fluffy white Samoyed in <Picture 2>, <Picture 3>, and <Picture 4>, with thick white fur, pointed ears, a dark nose, and a curved tail.\n<Subject 3> is the young blonde woman in <Video 1>, with long blonde hair and a light-pink button-down shirt with rolled-up sleeves.\n<Subject 4> is the young man in <Video 2>, with short wavy brown hair and a dark-grey hoodie with drawstrings.\n<Audio 1> is the voice-timbre reference for <Subject 3> (S1), containing a spoken English vocal layer.\n \nsummary:\n[reference generation + audio reference] The target video shows <Subject 3> eating a cookie in <Subject 1>. <Subject 4> enters with <Subject 2>, which lunges toward the cookie. The three-shot exchange uses <Audio 1> as the voice-timbre reference for <Subject 3> and ends with a canned audience laugh.\n \nretention_analysis:\n<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table are retained.\n<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the Samoyed's thick white fur, pointed ears, dark nose, and curved tail are retained.\n<Subject 3> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the blonde woman's identity, long hair, and light-pink shirt are retained.\n<Subject 4> (appears in [Shot 1], [Shot 2]): fully_preserved - the young man's short wavy brown hair and dark-grey hoodie are retained.\n<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.\n \ndetailed_description:\nThe target video uses a realistic multi-camera sitcom style with warm indoor lighting.\n[Shot 1] A medium shot establishes <Subject 1>, the coffee shop with its exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table. <Subject 3> (S1), the young woman with long blonde hair and a light-pink button-down shirt with rolled-up sleeves, sits on the sofa holding a chocolate-chip cookie. From the left, <Subject 4>, the young man with short wavy brown hair and a dark-grey hoodie with drawstrings, enters holding the leash of <Subject 2>, the thick-furred white Samoyed with pointed ears, a dark nose, and a curved tail. The dog lunges toward the cookie and pulls the leash taut. <Subject 3> (S1) jerks her hand back and, using the clear youthful voice timbre referenced from <Audio 1>, exclaims with light annoyance, <d>[English] Hey! Watch your dog!</d> She closes her lips and guards the cookie while <Subject 4> pulls the dog back.\n[Shot 2] At 00:03.000, the shot cuts to a close-up of <Subject 4> (S2), the young man in the dark-grey hoodie from Shot 1, sitting beside <Subject 3> on the sofa and holding <Subject 2> securely in his arms. <Subject 4> (S2) says in a casual young male voice with a playful tone and an easy conversational pace, <d>[English] He just likes cookies more than me.</d> He closes his mouth into an apologetic smile and strokes the dog's thick white fur.\n[Shot 3] At 00:05.000, the shot cuts to a close-up of <Subject 3> (S1), the blonde woman in the light-pink shirt from Shot 1. Her annoyance softens as she looks toward the Samoyed. <Subject 3> (S1) replies in the same clear youthful voice referenced from <Audio 1> with an amused cadence, <d>[English] Well, he has good taste at least.</d> She smiles and raises the cookie in a small toast-like gesture. A classic canned audience laugh begins immediately after the line and continues through the final frame.\n \noverall_soundscape:\nSoft indoor coffee-shop room tone continues throughout the scene.\n \nnon_diegetic_music:\nN/A</code></pre>\n<h1>H3 Ref2V User instruction</h1>\n<p><a href=\"https://pastebin.com/SWswRRxC\">https://pastebin.com/SWswRRxC</a></p>\n<pre class=\"language-json\"><code># Core Instructions\n \nYou are helping me turn a request into a MiniMax Hailuo H3 video-generation\nprompt in full-reference mode (Ref2VA), where I supply multiple reference assets\n— reusable subjects, images, source videos, and audio — and you rewrite my\nrequest into a six-section, label-tracked prompt.\n \nEverything under \"# Core Instructions\" is the RULESET: it tells you HOW to write\nthe prompt. My actual request is at the very bottom under \"# User Instructions\",\ntogether with the reference files I've attached to this message. Read the whole\nruleset first, then rewrite my request by following it exactly. Do not answer,\ncritique, or comment on the ruleset itself — it is instructions to follow, not\nsomething to respond to.\n \n## What I'm giving you (see \"# User Instructions\" below)\nUnder \"# User Instructions\" I provide:\n1. The video concept / what I want to happen.\n2. The target video duration in seconds (if I omit it, assume 5.00).\n3. For every reference asset: a LABEL and a TEXT DESCRIPTION, e.g.\n `Picture 1: a blonde woman in a light-pink shirt on an orange sofa`, and the\n file itself attached to this message.\n \nTreat my text description of each reference as the reliable anchor for what that\nlabel contains. If the app you're running in can natively view the attached\nimages (or play the attached audio/video), use the assets for finer detail, but\nnever contradict my description. Never invent a reference I did not provide, and\nnever leave a referenced label undefined.\n \n## Output contract\nReply with ONLY the six sections specified below, in the exact order and with the\nexact field names shown. Begin your reply directly with `subject_definitions:` —\nno preamble (\"Here's your prompt\", \"Sure\"), no closing remarks, no markdown\nheaders, no code fences. Write everything in English EXCEPT dialogue/lyrics\ninside `<d>` and text visibly present in the scene, which stay in their original\nlanguage. Timing is `MM:SS.mmm` for cuts and `S.SS` (two decimals) for the\nalignment line.\n \n## The six sections (exact order, exact names)\nsubject_definitions\nsummary\nretention_analysis\ndetailed_description\noverall_soundscape\nnon_diegetic_music\n \n### 1. subject_definitions — reference labels\nFour label types; once assigned, a label keeps the same meaning in every\nsection:\n- `<Subject N>`: reusable VISIBLE content (people, animals, objects, scenes,\n backgrounds, clothing, props, effects, styles, actions, expressions, poses).\n It is the content unit used in the target video, not the source file. One\n subject may come from several assets; one asset may yield several subjects.\n- `<Picture N>`: a reference image used as a concrete frame / keyframe / last\n frame / edited keyframe / composition or storyboard anchor.\n- `<Video N>`: a WHOLE-video relationship — editing a source video, continuing\n from its end, or referencing its camera/cuts/rhythm/temporal structure.\n- `<Audio N>`: a standalone audio asset or an enabled synchronized track from a\n reference video (copying signal, referencing BGM style, voice timbre/delivery,\n reusing dialogue/lyrics/SFX, or beat/continuity).\nGive each separately-tracked item its own line stating what the label denotes,\nits reference role, and the main features to follow. If a `<Picture N>` or\n`<Video N>` only identifies the SOURCE of another item and is not used\nseparately later, cite it INSIDE that item's definition without its own line.\nA person/object/scene/action/effect reused from a video is still a `<Subject N>`\n— `<Video N>` names the asset/structure, not the visible content. An ordinary\nreference video does NOT get an `<Audio N>` just because it has sound.\n`<Video N>` and `<Audio N>` are numbered independently; equal or different\nindices imply nothing about shared source.\nWhen an `<Audio N>` maps to a target speaker, reuse that speaker's GLOBAL id:\n`<Subject N> (Sx)` if it maps to a subject, else a stable voice description plus\n`(Sx)`. The id comes from the target video's global speaker order (Section 5);\nnever assign a new one in the audio definition.\nExamples:\n`<Subject 1> is the young woman in <Picture 1>, with long dark hair, a blue cardigan, and a thin silver necklace.`\n`<Subject 1> is the woman whose appearance comes from <Picture 1> and whose walking motion comes from <Video 1>.`\n`<Picture 3> is a storyboard reference for [Shot 1] and [Shot 2], defining their viewpoint, subject placement, and shot order.`\n`<Video 1> is the source video for the target video edit.`\n`<Audio 1> is the voice-timbre reference for <Subject 1> (S1).`\n \n### 2. summary — one short English paragraph\nBegins with a square-bracketed task-type prefix, then summarizes the target\nvideo and its reference relationships using ONLY already-defined labels (do not\nintroduce new labels here).\nTask types: `keyframe completion` (image as a concrete frame anchor) |\n`reference generation` (image/video/audio guides a character/scene/style/\naction/camera/storyboard without being a concrete frame or the edited/continued\nsource) | `video editing` (an existing source video is directly modified) |\n`video continuation` (new content continues/extends/resumes/transitions from a\nsource video) | `audio reuse` (same signal reused in full or part) |\n`audio reference` (only style/timbre/dialogue/SFX/beat/continuity referenced,\nnot copied).\nCombine multiple with ` + ` and never repeat a type\n(e.g. `[video continuation + keyframe completion]`). Presence of video/audio does\nNOT auto-create a task type: a video giving only camera/cuts/rhythm is\n`reference generation`; use `video editing`/`video continuation` only when that\nvideo is actually edited or continued. For video-editing tasks, start the body\nafter the prefix with `The target video is an edited version of <Video 1>.`\n \n### 3. retention_analysis — one line per label\nPreserve each label's meaning from subject_definitions. Do NOT write `(Sx)` here.\nDo not treat newly added actions/backgrounds/plot as losses of fidelity.\nVisible content (`<Subject N>`, `<Picture N>`, `<Video N>`) uses fixed markers:\n`fully_preserved` | `partially_preserved` | `attribute_transfer` |\n`weak_reference`.\nAudio (`<Audio N>`) uses: `fully_copy` | `partially_copy` | `reference` |\n`weak_reference`.\nEntry forms:\n`<Subject 1> (appears in [Shot 1], [Shot 3]): fully_preserved - ...`\n`<Picture 2> ([Shot 1] first frame): fully_preserved - ...`\n`<Video 1> (cut and pacing structure): weak_reference - ...`\n`<Audio 1>: fully_copy - <Audio 1> is reused 1:1 as the target video's complete final audio track.`\n`<Audio 2>: reference - the target speaker follows <Audio 2>'s voice timbre and measured delivery without copying the original signal.`\n \n### 4. detailed_description — main body, shot by shot in playback order\nEstablish the overall style in ONE or TWO English sentences BEFORE `[Shot 1]`\n(this is where the style opening lives in full-reference mode — not after\n`[Shot 1]`). Then describe each shot: composition, subject appearance and\nposition, environment and lighting, actions and state changes, camera movement,\ncurrent sound, dialogue, and the exact points where referenced content appears\nor takes effect. Insert `<Subject N>`, `<Picture N>`, `<Video N>`, `<Audio N>`\nat first appearance and wherever their roles apply; keep using the same label\nwithout redefining it. Do not reduce this to a plot summary or a list of\nreference relationships.\nConcrete frame anchors read naturally: `the shot begins from <Picture 1>`,\n`the shot's keyframe corresponds to <Picture 2>`, `the shot ends on <Picture 3>`.\nWhen a referenced subject speaks, keep BOTH the visual label and the speaker id:\n`<Subject 2> (S1) turns toward the woman and says, <d>[English] ...</d>`\n(off-screen: same form marked `off-screen`). Assign `(Sx)` once, by the order of\nactual vocal events in the target video, and reuse it at every vocal event.\nWhen a verbal cue exists only inside a directly reused BGM/soundtrack with no\nindependent vocal source, use `<Audio N>` as the audible source and do NOT\ninvent an `(Sx)`; a concrete person/character/narrator DOES get `(Sx)`.\nWhen dialogue/narration/lyrics from reference audio are directly reused (or I\nask for reperformance), preserve the exact source words and original language\ninside `<d>`; write `[unclear]` for unintelligible spans (never guess);\nstandardize punctuation to `, . ? !`, dropping tildes/emoji/decorative marks and\nending statements/questions/exclamations with `. ? !` before `</d>`. When only\ntimbre/rhythm/emotion/delivery is referenced, do NOT carry the original words\ninto the target video.\nLength: generation tasks are normally 350-500 English words; dialogue-dense\ncontent prioritizes fitting the full spoken timeline over word count; editing\nscales with source complexity. A single shot does not justify a short body —\ndistribute detail by information load.\n \n### 5. overall_soundscape and non_diegetic_music\nState a reference-audio relationship only in the matching audible layer:\nambience/SFX in overall_soundscape, audience-only score in non_diegetic_music.\nIf one audio provides both, describe the matching relationship in each section,\ne.g.:\n`overall_soundscape: The copied ambience layer from <Audio 1> continues throughout the target video.`\n`non_diegetic_music: <Audio 2> is directly reused as the complete audience-only score.`\nWrite full dialogue/lyrics only inside `<d>` in detailed_description; never\nrepeat them in these two sections.\n \n## Core writing rules (apply to every field)\n \n### Shots and cuts\nDo not put a timestamp on the first shot. Number later shots sequentially and\nbegin each with a strictly increasing cut time inside the duration:\n`[Shot 2] At 00:03.500, the camera cuts to ...`\nFor ordinary cuts use: `the camera cuts to`, `the shot cuts to`,\n`the shot transitions to`, `the shot changes to`, or `the shot switches to`.\nUse cross-dissolve, fade, or wipe only when I explicitly ask. A cut must\nintroduce new information (subject, space, state, viewpoint, or time); if only\ndistance or a slight angle changes, prefer camera motion instead.\n \n### Camera motion = motion type + amplitude + speed\nWrite camera motion as natural English inside the shot, not as stacked labels.\nAdd amplitude/speed only when meaningful (medium amplitude and normal speed are\nomitted).\nMotion type: Zoom In/Zoom Out (focal length changes, body still) |\nPush In/Pull Out (camera moves forward/back) |\nPan Left/Pan Right (pivots horizontally) |\nTruck Left/Truck Right (translates horizontally) |\nTilt Up/Tilt Down (pivots vertically) |\nPedestal Up/Pedestal Down (whole camera up/down) |\nArc Shot | Tracking Shot | Static Shot |\nShake Slightly/Shake Strongly | POV |\nRoll Clockwise/Roll Counterclockwise.\nAmplitude: `with small amplitude` | `with large amplitude`.\nSpeed: `at slow speed` | `at fast speed`.\nExamples:\n`The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.`\n`The camera pans right with large amplitude at fast speed, revealing the open doorway.`\n`The camera holds a static shot as the runner exits the frame.`\n \n### Speakers, dialogue, singing\nAnyone who speaks, sings, or makes an off-screen human voice gets a stable ID:\n`(S1)`, `(S2)`, ... A speaker keeps the same ID across shots; silent characters\nget no ID. For simultaneous speech use a compound ID like `(S1,S2)`.\nOn first appearance, establish a stable identity (type, age, gender, on/off\nscreen, pitch, timbre, rate, accent). Put the speaker's identifying phrase, ID,\naction, and delivery OUTSIDE `<d>`. Inside `<d>`, put ONLY the language tag and\nthe verbatim words — never translate or rewrite; preserve every word and\npunctuation mark.\n`The young woman with a quiet, breathy voice (S1) says: <d>[English] I get off at the next station.</d>`\n`The two children (S1,S2) shout together, <d>[English] Wait for us!</d>`\nFor voiceover use the exact phrase `says in an off-screen voiceover`, and\nimmediately after the `<d>` block state the on-screen character's lips stay\nclosed:\n`The man (S1) says in an off-screen voiceover: <d>[English] I still remember that road.</d> while his lips remain completely closed.`\nWhen one line of dialogue/lyrics crosses a cut, put `<scenetrans>` at the join\nin BOTH parts and state the audio continues across the cut (e.g.\n`continues seamlessly across the cut`, `carries over from the previous shot`).\nUse `<cutoff>` when speech is truncated by the video end.\n \n### On-screen text\nAny banner/sign/label/subtitle/neon actually visible on screen goes in English\ndouble quotes, verbatim, untranslated:\n`A red neon sign reading \"营业中\" glows above the doorway.`\n \n### overall_soundscape\n1-4 English sentences, one paragraph: ambient sound, physical-action sounds,\nnon-verbal human sounds (wind, rain, traffic, footsteps, fabric, impacts,\nbreathing, laughter, panting). Do NOT repeat dialogue, singing, or diegetic\nmusic here. Use `N/A` only if I explicitly want full silence.\n \n### non_diegetic_music\n1-3 English sentences describing audience-only background music: instrumentation,\ntempo, rhythm, dynamic changes. No abstract mood words, no emotional-function\nexplanations. Music the characters can hear (singing, instruments, radio, TV,\nphone) is diegetic and belongs in the main description, not here. Use `N/A` when\nthere is no non-diegetic music.\n \n## Worked example (format reference only — do NOT copy its content or echo it back)\nsubject_definitions:\n<Subject 1> is the coffee-shop environment in <Picture 1>, featuring an exposed brick wall, an orange tufted sofa with patterned pillows, a neon sign, and a wooden coffee table.\n<Subject 2> is the fluffy white Samoyed in <Picture 2>, <Picture 3>, and <Picture 4>, with thick white fur, pointed ears, a dark nose, and a curved tail.\n<Subject 3> is the young blonde woman in <Video 1>, with long blonde hair and a light-pink button-down shirt with rolled-up sleeves.\n<Subject 4> is the young man in <Video 2>, with short wavy brown hair and a dark-grey hoodie with drawstrings.\n<Audio 1> is the voice-timbre reference for <Subject 3> (S1), containing a spoken English vocal layer.\n \nsummary:\n[reference generation + audio reference] The target video shows <Subject 3> eating a cookie in <Subject 1>. <Subject 4> enters with <Subject 2>, which lunges toward the cookie. The three-shot exchange uses <Audio 1> as the voice-timbre reference for <Subject 3> and ends with a canned audience laugh.\n \nretention_analysis:\n<Subject 1> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table are retained.\n<Subject 2> (appears in [Shot 1], [Shot 2]): fully_preserved - the Samoyed's thick white fur, pointed ears, dark nose, and curved tail are retained.\n<Subject 3> (appears in [Shot 1], [Shot 2], [Shot 3]): fully_preserved - the blonde woman's identity, long hair, and light-pink shirt are retained.\n<Subject 4> (appears in [Shot 1], [Shot 2]): fully_preserved - the young man's short wavy brown hair and dark-grey hoodie are retained.\n<Audio 1>: reference - its vocal timbre guides the dialogue delivery of <Subject 3> without copying the original signal.\n \ndetailed_description:\nThe target video uses a realistic multi-camera sitcom style with warm indoor lighting.\n[Shot 1] A medium shot establishes <Subject 1>, the coffee shop with its exposed brick wall, orange tufted sofa, patterned pillows, neon sign, and wooden coffee table. <Subject 3> (S1), the young woman with long blonde hair and a light-pink button-down shirt with rolled-up sleeves, sits on the sofa holding a chocolate-chip cookie. From the left, <Subject 4>, the young man with short wavy brown hair and a dark-grey hoodie with drawstrings, enters holding the leash of <Subject 2>, the thick-furred white Samoyed with pointed ears, a dark nose, and a curved tail. The dog lunges toward the cookie and pulls the leash taut. <Subject 3> (S1) jerks her hand back and, using the clear youthful voice timbre referenced from <Audio 1>, exclaims with light annoyance, <d>[English] Hey! Watch your dog!</d> She closes her lips and guards the cookie while <Subject 4> pulls the dog back.\n[Shot 2] At 00:03.000, the shot cuts to a close-up of <Subject 4> (S2), the young man in the dark-grey hoodie from Shot 1, sitting beside <Subject 3> on the sofa and holding <Subject 2> securely in his arms. <Subject 4> (S2) says in a casual young male voice with a playful tone and an easy conversational pace, <d>[English] He just likes cookies more than me.</d> He closes his mouth into an apologetic smile and strokes the dog's thick white fur.\n[Shot 3] At 00:05.000, the shot cuts to a close-up of <Subject 3> (S1), the blonde woman in the light-pink shirt from Shot 1. Her annoyance softens as she looks toward the Samoyed. <Subject 3> (S1) replies in the same clear youthful voice referenced from <Audio 1> with an amused cadence, <d>[English] Well, he has good taste at least.</d> She smiles and raises the cookie in a small toast-like gesture. A classic canned audience laugh begins immediately after the line and continues through the final frame.\n \noverall_soundscape:\nSoft indoor coffee-shop room tone continues throughout the scene.\n \nnon_diegetic_music:\nN/A\n \n# User Instructions\n \nConcept:\n<Describe what happens in the video. Shot by shot is ideal, but plain prose is fine — the ruleset above will structure it.>\n \nDuration (seconds):\n<e.g. 8.00 — leave blank to default to 5.00>\n \nReferences (give each a label + a text description, and attach the file to this message):\n- Subject 1: <what it is and its key visual features>\n- Picture 1: <what the image shows and how it's used — first frame, storyboard, etc.>\n- Video 1: <what the video is and how it's referenced — edited, continued, or camera/rhythm only>\n- Audio 1: <what the audio is and how it's used — reused 1:1, or timbre/style reference only>\n<Add or delete lines to match what you actually have. Remove any label type you're not using.>\n \nAnything else:\n<optional notes — style, mood cues, must-keep details></code></pre>\n<p>use this system prompt for the llm</p>\n<p><a href=\"https://www.reddit.com/r/StableDiffusion/comments/1vn7243/is_there_a_h3_minimax_prompt_template_available/\">https://www.reddit.com/r/StableDiffusion/comments/1vn7243/is_there_a_h3_minimax_prompt_template_available/</a></p>\n<pre class=\"language-json\"><code> # System Prompt: Video Prompt Rewriter (FL2VA / REF2VA / T2VA / I2VA / L2VA)\n\nYou are an expert video-prompt rewriter specialized in converting user instructions into strictly formatted, production-ready prompts for video generation models.\n\nYour sole job is to analyze the user’s request, determine the correct task mode, and output a complete, correctly structured prompt that follows the rules below exactly. Never explain your reasoning unless the user asks. Output only the final structured prompt.\n\n## 1. Mode Detection Rules\n\nAnalyze the user message and any attached images/videos/audio to classify the task:\n\n| Mode | Detection Signals |\n\n|------|-------------------|\n\n| **FL2VA** | User provides (or clearly intends) a first-frame image **and** a last-frame image, and wants continuous motion/path between them. Keywords: “from this to that”, “start with picture A end with picture B”, “first and last frame”, “FL2VA”. |\n\n| **REF2VA** (Full-Reference) | User provides one or more reference images, videos, or audio assets that must be tracked with labels (`<Subject N>`, `<Picture N>`, `<Video N>`, `<Audio N>`). The request involves reusing, transferring, editing, continuing, or referencing specific visual/audio content. Keywords: “reference”, “use this character/scene/video”, “edit this video”, “continue from this”, “keep the style of”, “REF2VA”, “full reference”. |\n\n| **I2VA** | Single reference image is to be used as the **first frame** only, then the video develops forward from it. |\n\n| **L2VA** | Single reference image is to be used as the **last frame** only; the video must converge to it. |\n\n| **T2VA** | Pure text-to-video; no reference images/videos/audio provided. |\n\nIf multiple modes could apply, prefer the most specific:\n\n- Presence of both first + last frames → FL2VA\n\n- Explicit full-reference labels or multi-asset reuse → REF2VA\n\n- Otherwise fall back to I2VA / L2VA / T2VA as appropriate.\n\n## 2. Output Rules by Mode\n\n### A. FL2VA / I2VA / L2VA / T2VA (Base Guide)\n\nFollow the **Video Prompt Writing Guide (T2VA / I2VA / FL2VA / L2VA)** exactly.\n\n**Structure:**\n\n **Instruction line** (only for I2VA / FL2VA / L2VA; omit for pure T2VA):\n\n- **I2VA**:\n\n```\n\nFor the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.\n\n```\n\n- **FL2VA**:\n\n```\n\nHow the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.\n\n```\n\n(Replace N and S.SS with the actual final shot index and duration formatted to two decimal places.)\n\n- **L2VA**:\n\n```\n\nHow the reference pictures align with the target video — <Picture 1> (from [Shot N]) aligns with the S.SS-second mark of the target video.\n\n```\n\n2. Blank line\n\n3. Three core fields:\n\n```\n\nintegrated_multimodal_description: [Shot 1] ...\n\noverall_soundscape: ...\n\nnon_diegetic_music: ...\n\n```\n\n**Key writing rules for `integrated_multimodal_description`:**\n\n- Begin Shot 1 with style + composition.\n\n- Use exact camera vocabulary (Push In, Pull Out, Pan, Truck, Tilt, Pedestal, Arc, Tracking, Static, Zoom, Shake, POV, Roll) + amplitude + speed when meaningful.\n\n- Speakers receive stable IDs `(S1)`, `(S2)`, …\n\n- Dialogue/lyrics go inside `<d>[Language] exact text</d>`. Preserve original language and wording.\n\n- Use `<scenetrans>` for dialogue that crosses cuts and `<cutoff>` when speech is truncated by the video end.\n\n- Visible on-screen text goes in English double quotes, preserved verbatim.\n\n- For FL2VA: describe the continuous path from first-frame state → intermediate changes → last-frame landing. Prefer single continuous shot unless the user explicitly requests cuts.\n\n- For I2VA: anchor on the first frame then develop forward.\n\n- For L2VA: invent a plausible preceding state then converge to the last frame.\n\n### B. REF2VA (Full-Reference Mode)\n\nFollow the **Full-Reference Mode Rewrite Output Format Guide** exactly.\n\n**Mandatory six-section structure in this exact order:**\n\n```\n\nsubject_definitions:\n\n...\n\nsummary:\n\n[task-type] ...\n\nretention_analysis:\n\n...\n\ndetailed_description:\n\n...\n\noverall_soundscape:\n\n...\n\nnon_diegetic_music:\n\n...\n\n```\n\n#### subject_definitions\n\n- Define every reusable visual unit as `<Subject N>`.\n\n- Define concrete frame anchors as `<Picture N>` only when the image itself is used as first/last/key frame or storyboard.\n\n- Define whole-video structural sources as `<Video N>`.\n\n- Define audio assets as `<Audio N>`.\n\n- One line per label. Cite source assets inside the definition when needed.\n\n- Never invent labels that are not used later.\n\n#### summary\n\n- Starts with a square-bracketed task-type prefix, e.g.:\n\n- `[reference generation]`\n\n- `[video editing + audio reuse]`\n\n- `[video continuation + keyframe completion]`\n\n- `[reference generation + audio reference]`\n\n- Combine multiple applicable types with ` + `. Do not repeat types.\n\n- One short English paragraph describing the target video and the main reference relationships using the labels already defined.\n\n#### retention_analysis\n\n- One line per previously defined label.\n\n- Visible content uses: `fully_preserved`, `partially_preserved`, `attribute_transfer`, `weak_reference`.\n\n- Audio uses: `fully_copy`, `partially_copy`, `reference`, `weak_reference`.\n\n- Format examples:\n\n```\n\n<Subject 1> (appears in [Shot 1], [Shot 3]): fully_preserved - ...\n\n<Picture 2> ([Shot 1] first frame): fully_preserved - ...\n\n<Video 1> (cut and pacing structure): weak_reference - ...\n\n<Audio 1>: reference - ...\n\n```\n\n#### detailed_description\n\n- Write in English. Preserve original language only inside `<d>` and for visible text.\n\n- Open with 1–2 sentences establishing overall visual style before `[Shot 1]`.\n\n- Then describe shot-by-shot in playback order.\n\n- Insert reference labels at first clear appearance and wherever their role applies.\n\n- Use natural phrasing for frame anchors: “the shot begins from <Picture 1>”, “ends on <Picture 3>”, etc.\n\n- Speakers: when a referenced subject speaks, write `<Subject N> (Sx)`.\n\n- Dialogue/lyrics: `<d>[Language] exact original text</d>`. Do not paraphrase.\n\n- Length target for pure generation: ~350–500 English words. Editing tasks scale with source complexity.\n\n#### overall_soundscape & non_diegetic_music\n\n- Follow the same definitions as the base guide.\n\n- When reference audio is involved, state the copy/reference relationship in the correct section (ambience/SFX → overall_soundscape; audience-only score → non_diegetic_music).\n\n## 3. Universal Rules (All Modes)\n\n- Write all descriptive text in English. Preserve original language only for dialogue, lyrics, and visible on-screen text.\n\n- Never invent dialogue or lyrics that the user did not provide.\n\n- Never translate dialogue.\n\n- Camera motion must be expressed as natural English inside the shot description, not as stacked labels.\n\n- Shot timestamps (after Shot 1) must be strictly increasing and formatted `At MM:SS.mmm,`.\n\n- Speaker IDs are assigned in order of first vocal event in the target video and reused consistently.\n\n- Use `N/A` for non_diegetic_music or overall_soundscape only when the user explicitly requests complete silence or no score.\n\n- Do not add extra sections, markdown headings, or commentary outside the required structure.\n\n- If the user request is ambiguous, choose the most specific mode that fits the provided assets and intent, then produce a complete valid prompt.\n\n## 4. Output Discipline\n\n- Output **only** the final structured prompt.\n\n- No preamble, no explanation, no “Here is the prompt:”.\n\n- Exactly match the required field names and order for the detected mode.\n\n- If both FL2VA and REF2VA signals are present, prefer REF2VA and incorporate the first/last frame relationships inside the full-reference structure.\n\nBegin rewriting the user’s next message according to these rules. </code></pre>\n<p>shared setup for a scene</p>\n<pre class=\"language-json\"><code>MINIMAX H3 — REF2VA CHAIN — \"THE BATHHOUSE\" — 7 SCENES / 102s\n================================================================\n \nWORKFLOW NOTES\n- Node: H3 Chain Plan — Timed Segments (MiniMax H3 Ref2VA + Motion Context)\n- <Picture 1> = the young woman's character design sheet\n- <Picture 2> = the young wizard's character design sheet\n- Tested settings: context_length 22, encode_mode video, anchor_mode head\n- The SHARED PROMPT below is prepended automatically to every scene.\n Paste it into the shared prompt box, not into the individual scenes.\n \nHOW THE SCENES CONNECT\nEach scene opens by restating how the previous scene ended, and holds that\nstate with a small live action (a broom stroke, a breath, a step) for about\ntwo seconds before anything new happens. The pinned motion-context frames\nare ~0.92s, so the first cut in every scene lands well clear of them.\nDo not \"fix\" this by trimming the opening holds — they are what stop the\nmodel rendering the old and new states at the same time.\n \nSPEAKER IDS (persistent across the whole chain — do not renumber)\n (S1) the young woman\n (S2) the old man at the shutters\n (S3) the young wizard\n (S4) the bathhouse keeper\n \nKNOWN ISSUES TO WATCH\n- Character bleed: the mother and daughter can render as copies of the lead.\n Their full description with the \"no blue hair / no yukata\" clauses is\n repeated in every scene they appear in. Keep it.\n- Wardrobe: she changes outfit twice. Each scene states which outfit is worn\n plus what it is NOT. Keep those clauses too.\n- The wizard is defined in the shared prompt, so he is conditioned into every\n scene whether mentioned or not. If he appears uninvited in scenes 1-3,\n move his definition out of the shared block and into scene 4 only.\n \n \n================================================================\nSHARED PROMPT (prepended to every scene)\n================================================================\n \nModern 2D-animated anime feature film style, high-definition cel shading, clean line art, sharp saturated color palettes, vibrant cinematic lighting, expressive character animation. <Picture 1> is a character design sheet defining the young woman: her facial identity, shoulder-length blue bob with a blunt fringe, bright blue eyes, fair skin, slim build, and a red ribbon tied at the back of her hair. <Picture 2> is a character design sheet defining the young wizard: his facial identity, messy sandy-blond hair, wide anxious blue eyes, tall thin proportions, deep blue robe with a draped cowl collar, brown boots, large floppy dark navy pointed hat with a red band and a small green sprig, and tall gnarled wooden staff with a curled head. Preserve the exact facial identity, hair and body proportions from <Picture 1> and <Picture 2> wherever those characters appear; do not copy the sheets' neutral standing poses, plain backgrounds or head-study framing. Each character's clothing and presence are specified separately in each part and must follow that part's description only. Maintain strict character, prop, lighting, environment and motion continuity across all parts, and continue unfinished character motion and camera movement between parts without resetting positions, poses or actions.\n \n \n================================================================\nSCENE 1 — part_01_the_windy_staircase — 12s\n================================================================\n \n[Shot 1] A medium-wide shot frames the young woman from <Picture 1> descending a sunlit stone staircase in a small hillside town on a windy afternoon, flowering plants crowding both edges of the frame and terracotta rooftops stepping away below toward the sea. She wears a light yellow sleeveless sundress with a tied sash at the waist, bare arms and shoulders, brown sandals and a wide straw hat, with no red fabric, no wide sleeves and no head covering. The camera trucks left with small amplitude at slow speed to stay level with her descent as a gust tips the hat brim and she lifts one hand to hold it in place, bringing her other hand down to keep her skirt settled against her legs. The young woman with a light, bright voice and a quick, easy delivery (S1) says: <d>[English] this wind is killer</d> and her shoulders shake as she laughs. [Shot 2] At 00:04.000, the shot cuts to a close-up of her face beneath the straw hat, backlit by the sun with a bright rim along her jaw, her blue bob whipping across her cheek. The camera holds a static shot as she blinks slowly, a faint blush spreading across her cheeks, and turns her head to follow a butterfly crossing in front of her. [Shot 3] At 00:06.000, the shot cuts to a wide low angle from further down the staircase looking up at the weathered stone wall against the bright sky. A pair of wooden shutters bangs open and an old man leans out, his long white beard whipping sideways in the gust and his thick round spectacles catching the light. The old man with a gravelly, unhurried voice (S2) calls down, <d>[English] Hold that hat, missy! The last one is still in my garden!</d> She glances up mid-step from below without stopping, laughing, one hand still clamped on the brim. [Shot 4] At 00:09.500, the shot cuts to a low shot close to the stone treads as a small black cat trots down the steps after her with its tail wagging, her yellow skirt and brown sandals moving ahead of it in the upper frame. The camera holds a static shot as loose leaves blow past the front of the lens and cel-shaded foliage shadows shift across the stone. The clip ends with the young woman still walking down the steps, the cat still following and the wind still moving her skirt.\n \noverall_soundscape: Wind gusts steadily through the foliage, rustling leaves and pushing loose leaves past the frame. Fabric snaps and flutters against her legs and her sandals tap unevenly down the stone treads. Wooden shutters bang hard against the wall, and she gives a short bright giggle after her line. Distant birdsong and faint small-town ambience continue underneath throughout.\n \nnon_diegetic_music: A bright acoustic-guitar figure at a moderate tempo with light plucked ukulele and soft brushed percussion, playing continuously and still running at the end of the clip.\n \n \n================================================================\nSCENE 2 — part_02_the_bathhouse — 15s\n================================================================\n \n[Shot 1] The young woman continues descending the sunlit stone staircase exactly as before, her hand still on the straw hat brim, the black cat still following and the camera still trucking left with small amplitude at slow speed. She wears the light yellow sleeveless sundress with a tied sash, bare arms and shoulders, brown sandals and a wide straw hat, with no red fabric and no head covering. She takes two more steps down, adjusts her grip on the brim and glances ahead toward the rooftops as the wind keeps moving her skirt. [Shot 2] At 00:03.000, the shot cuts to a wide shot of a narrow sunlit street at the foot of the hill, terracotta roofs and shuttered windows lining both sides. She walks toward the camera with one hand still on her hat, the black cat trotting beside her, past a low wooden bathhouse frontage where split noren curtains lift and snap above the doorway. She slows, turns toward the entrance and steps through the noren. [Shot 3] At 00:06.500, the shot cuts to the interior of a wooden bathhouse entrance hall, looking back toward the open doorway where the noren settle against the bright street outside. She walks toward the camera into the hall with her back to the daylight, lifting the straw hat from her head and holding it against her chest. A woman in her forties in a deep red yukata with her hair pinned up turns from the counter and bows in greeting with a warm smile. The bathhouse keeper with a low, easy voice (S4) says, <d>[English] Right on time.</d> The young woman smiles back and pushes a strand of blue hair from her face. [Shot 4] At 00:10.000, the shot cuts to a narrow wood-panelled changing room lined with woven baskets, the folded yellow sundress and the straw hat resting in an open basket on the shelf. The young woman now wears a deep red yukata with wide long sleeves, a gold and dark obi at the waist and geta on her feet, with no yellow fabric. The camera holds a static shot as she binds the yukata sleeves back with a cord across her shoulders, knots a folded white cloth over her hair, and takes hold of the sliding door. [Shot 5] At 00:13.000, the shot cuts to a low shot from inside the bath hall looking back at the doorway as she slides it open and steps through into a wall of white steam, silhouetted for a moment against the brighter changing room behind her before the light finds her. She slides the door shut behind her and reaches for a wide flat-headed push broom leaning against the wall. The clip ends with her lifting the broom, steam still rolling around her.\n \noverall_soundscape: Wind gusts through foliage and sandals tap unevenly down stone treads, giving way to hanging noren brushing softly and wooden sandals knocking on floorboards. A woven basket creaks, fabric rustles and a cord pulls tight in the quiet changing room, then a wooden door rolls open on its track and steam hisses softly through the gap over a low steamy room tone.\n \nnon_diegetic_music: A bright acoustic-guitar figure with light plucked ukulele thins as the interior arrives, and a koto enters at a slower tempo with sustained low strings underneath as the sliding door opens, still playing at the end of the clip.\n \n \n================================================================\nSCENE 3 — part_03_the_bath_hall — 15s\n================================================================\n \n[Shot 1] The young woman continues stepping into the steam-filled bath hall exactly as before, the sliding door shut behind her and the push broom now in her hands. She wears a deep red yukata with wide sleeves bound back by a cord, a gold and dark obi, geta, and a folded white cloth knotted over her hair, with no yellow fabric and no straw hat. She sets the broom head down on the wet tiles and takes her first long forward stroke. [Shot 2] At 00:03.000, the shot cuts to a wide dutch-angle shot of the sento interior at golden hour, the frame canted so the tiled floor runs diagonally across it, low sun pouring through tall intact lattice windows in visible rays as the surface of the hot spring glitters and steam drifts through the light. The camera holds a static shot as she works the broom across the wet tiles in long forward strokes, humming with a closed smile, the red yukata moving with each stroke. [Shot 3] At 00:07.000, the shot cuts to a medium shot from the side, holding on a frog the size of a grown man in profile as he leans back against the tiled rim of the bath with his arms spread along the edge and his head tipped up, basking, a small white towel folded on his head. His throat pouch swells and he releases a low ribbit. [Shot 4] At 00:10.000, the shot cuts to a wide shot from the far side of the hall as a mother with long dark hair pinned up walks in from the right wrapped in a large towel and carrying a wooden bucket. Beside her walks her young daughter, a small child of about six with short brown hair and brown eyes, clearly much younger and much shorter than the young woman, wrapped in her own towel covering her from chest to knees. The daughter has no blue hair and wears no yukata. She stops dead to stare at the frog while her mother steers her gently onward to the water's edge. The young woman glances over without breaking her humming and keeps sweeping. [Shot 5] At 00:13.000, the shot cuts back to the wide dutch-angle framing of the hall, the camera holding a static shot as she continues sweeping in steady rhythm and the mother and her brown-haired daughter settle at the water's edge behind her. The clip ends with the young woman still sweeping, still humming, steam still drifting through the light rays.\n \noverall_soundscape: Water drips and trickles continuously over a low steamy room tone as a wide broom head pushes water across tile in long sweeping strokes. A deep resonant ribbit echoes off the tiles, a wooden bucket knocks against the floor and small bare feet slap across wet tiles.\n \nnon_diegetic_music: A koto figure at a slow tempo with sustained low strings underneath, joined by soft plucked notes as she sweeps, playing continuously and still running at the end of the clip.\n \n \n================================================================\nSCENE 4 — part_04_the_crash — 12s\n================================================================\n \n[Shot 1] The young woman continues sweeping the wet tiled floor of the steam-filled bath hall at the same steady rhythm, still humming, the camera still holding its wide dutch-angle framing with golden light pouring through the tall lattice windows. She wears the deep red yukata with its sleeves bound back and the folded white cloth knotted over her hair. At the water's edge behind her sit a mother with long dark hair pinned up, wrapped in a large towel, and her young daughter, a small child of about six with short brown hair and brown eyes, clearly much younger and much shorter than the young woman, wrapped in her own towel covering her from chest to knees. The daughter has no blue hair and wears no yukata. She completes two more long forward strokes, shifts her weight and glances up briefly toward the light before returning to the tiles. [Shot 2] At 00:03.500, the shot cuts to a wide shot of the tall lattice window as it explodes inward in a burst of splintered wood and glass, drawn with hard smear frames and radiating white impact lines. The young wizard from <Picture 2> tumbles through in a tangle of limbs in his deep blue robe and floppy dark navy pointed hat, his tall gnarled staff clattering across the tiles. He lands flat on his back on the wet floor and skids to a stop at her feet as steam swirls violently around the impact, her broom clattering from her hands and the mother pulling her small brown-haired daughter back from the water. [Shot 3] At 00:07.000, the shot cuts to a low close shot from just past his shoulder, looking up as she leans into frame above him against the bright window light, the white cloth knocked askew on her hair and loose blue strands falling forward. The camera pushes in with small amplitude at slow speed as her face fills the upper frame with a concerned expression. The young woman with a light, bright voice and a quick, easy delivery (S1) asks, <d>[English] Are you alright?</d> [Shot 4] At 00:10.000, the shot cuts to a close shot of the young wizard's face against the wet tiles as he stares up at her without answering, his wide blue eyes blown huge beneath the brim of his floppy navy hat and his face beginning to flush. The clip ends on his face mid-flush, steam still swirling and glass still settling across the tiles.\n \noverall_soundscape: A low steamy room tone with dripping water and long sweeping broom strokes breaks as a window shatters in a burst of splintering wood and glass, a body slaps hard onto wet tile and a wooden staff clatters and rolls to a stop. Broken glass shifts faintly underfoot, steam hisses through the open frame and a deep resonant ribbit echoes off the tiles.\n \nnon_diegetic_music: A koto figure at a slow tempo with sustained low strings cuts out entirely at the moment of the crash, leaving a single held low note ringing under the aftermath at the end of the clip.\n \n \n================================================================\nSCENE 5 — part_05_the_repair — 15s\n================================================================\n \n[Shot 1] The young wizard continues lying flat on his back on the wet tiles exactly as before, his face still flushed and the young woman still leaning over him out of frame. He blinks once, his mouth opening slightly without a sound, his eyes still fixed upward. [Shot 2] At 00:02.500, the shot cuts to a wide shot of the bath hall as his flush deepens to dark red and he scrambles upright in a panic, snatches his gnarled staff off the tiles and clambers back out through the gaping hole where the window used to be in a flurry of blue robe. The young wizard with a thin, flustered voice (S3) blurts, <d>[English] Sorry! Sorry!</d> as he disappears below the sill. The camera pulls out with small amplitude at slow speed as the young woman straightens up amid the wreckage in her deep red yukata and stares at the hole, head tilted. [Shot 3] At 00:06.000, the shot cuts to a hard dutch-angle medium shot of the broken window, the frame canted steeply so the empty frame runs diagonally across it. His floppy navy pointed hat and wide blue eyes rise slowly back into view from below the sill. The young wizard (S3) says, <d>[English] I'll pay for that. Probably.</d> He raises the gnarled staff and taps it once against the frame with a soft chime. A ring of white sparks spreads outward and splinters, glass shards and lattice slats lift off the tiles across the room and slide backward through the air in reversed smear frames, converging on the empty frame and locking into place piece by piece until the window stands whole again. He tips his hat and drops out of sight. [Shot 4] At 00:11.500, the shot cuts to a medium shot of the young woman leaning on her push broom, looking at the repaired window with her head tilted and one eyebrow raised. The young woman (S1) says, <d>[English] What an odd man.</d> Behind her stand a mother with long dark hair pinned up, wrapped in a large towel, and her young daughter, a small child of about six with short brown hair and brown eyes, clearly much younger and much shorter than the young woman, wrapped in her own towel covering her from chest to knees. The daughter has no blue hair and wears no yukata. Both are still staring at the window, and the mother slowly turns her head to look at the young woman. [Shot 5] At 00:13.500, the shot cuts to a wide dutch-angle shot of the bath hall as she shrugs and turns back to the wet tiles. Off to the left the frog the size of a grown man still leans back against the tiled rim of the bath with a small white towel folded on his head, unbothered. A single pane drops out of the repaired window and shatters across the floor behind her. She stops mid-stroke without turning around. The clip ends with her frozen mid-stroke, steam still drifting through the light rays.\n \noverall_soundscape: A low steamy room tone and dripping water continue beneath scrambling footsteps and a wooden staff scraping stone. A soft chime rings out followed by a long crystalline shimmer as glass and wood slide back into place, ending in a small final chime, then hurried footsteps retreat on gravel outside. A broom head resumes pushing water across tile before a single sharp pane shatter breaks the rhythm.\n \nnon_diegetic_music: A single held low note gives way to a koto figure with a light ascending line and soft bells as the window reassembles, resolving into a warm sustained chord that cuts off dead on the falling pane, leaving one final plucked note.\n \n \n================================================================\nSCENE 6 — part_06_closing_up — 15s\n================================================================\n \n[Shot 1] The young woman continues standing frozen mid-stroke in the steam-filled bath hall exactly as before, her back to the shattered pane, both hands still on the push broom handle in her deep red yukata with the white cloth knotted askew over her hair. Her shoulders drop as she lets out a slow breath. [Shot 2] At 00:02.200, the shot cuts to an extreme close-up of her face as she closes her eyes and raises one hand to press her palm flat against her forehead, holding it there. She opens her eyes, straightens, and wipes her damp forehead with the back of her wrist, loose strands of blue hair stuck to her temple. She grins and the young woman with a light, bright voice and a quick, easy delivery (S1) says, <d>[English] All done!</d> [Shot 3] At 00:06.000, the shot cuts to the narrow wood-panelled changing room at night, a paper lantern glowing overhead. Two woven baskets sit side by side on the shelf, the left one empty and the right one holding a neatly folded light yellow sundress with a wide straw hat resting on top. The camera holds a static shot as a pair of hands enters the frame from above and lowers a neatly folded deep red yukata into the left basket, laying a wound obi sash on top of it, then withdraws. After a beat the hands return to the right basket, lift the folded yellow sundress clear of it, and carry it up out of frame, leaving the straw hat behind. [Shot 4] At 00:09.000, the shot cuts to a medium shot of the young woman now fully dressed in the light yellow sleeveless sundress with its tied sash, bare arms and shoulders, with no red fabric and no head cloth. She lifts the wide straw hat from the basket, settles it onto her head and adjusts the brim with both hands, then turns and slides the changing room door open onto the lamplit hall. [Shot 5] At 00:11.500, the shot cuts to the bathhouse entrance hall at night, warm paper lanterns glowing along the counter and the split noren curtains hanging still in the doorway. She crosses the hall and lifts a hand in a small wave to the bathhouse keeper in her deep red yukata behind the counter, who waves back. She pushes through the noren and out into the dark street. [Shot 6] At 00:13.500, the shot cuts to a wide exterior shot of the hillside town from above, terracotta rooftops stepping down toward the sea, window lights coming on one by one across the town and a street lamp flickering alight on the staircase below as the last gold light drains from the ridge into cool blue shadow. The clip ends on the lit town with clouds still moving overhead.\n \noverall_soundscape: A low steamy room tone with dripping water gives way to a quiet changing room where woven baskets creak and fabric rustles, then a wooden door rolls open on its track. Hanging noren brush softly and wooden sandals knock on floorboards before a night street opens out with crickets, a low breeze through foliage and distant town ambience.\n \nnon_diegetic_music: A koto figure resolves into a warm sustained chord, thinning into a soft solo acoustic guitar at a slow tempo with a light sustained pad as the night street arrives, still playing at the end of the clip.\n \n \n================================================================\nSCENE 7 — part_07_the_stranger_on_the_steps — 15s\n================================================================\n \n[Shot 1] The wide exterior of the hillside town at night continues exactly as before, window lights glowing across the terracotta rooftops and clouds still moving overhead, the street lamp lit on the staircase below. The camera holds a static shot as a moth crosses the lamp light and the last blue drains from the sky. [Shot 2] At 00:02.400, the shot cuts to a low angle looking up the stone staircase at the weathered stone wall, the bathhouse windows glowing warm below. A pair of wooden shutters bangs open and an old man leans out holding a lit lantern, his long white beard and thick round spectacles catching the flame light. The old man with a gravelly, unhurried voice (S2) calls down, <d>[English] It's late, missy! Straight home!</d> The young woman stops mid-step on the stairs in her light yellow sundress and wide straw hat, looks up and laughs. The young woman with a light, bright voice and a quick, easy delivery (S1) calls back, <d>[English] Yes, sir!</d> He grunts and pulls the shutters closed, and the lantern light disappears from the wall. [Shot 3] At 00:06.500, the shot cuts to a hard dutch-angle low shot of the staircase, the frame canted so the steps run diagonally across it, the wall now dark with its shutters closed and a single street lamp throwing long hard-edged shadows down the treads. The camera holds a static shot as she climbs from right to left with one hand trailing along the stone wall, watching her footing, moths circling the lamp. [Shot 4] At 00:09.000, the shot cuts to a wide shot looking down the staircase from above, the young woman high on the steps in the upper left with her back to the camera and the young wizard far below her at the bottom in the lower right, standing motionless and half swallowed in shadow, only the silhouette of his floppy navy pointed hat and the curled head of his gnarled staff catching the light. The young wizard (S3) calls up, <d>[English] Hey.</d> [Shot 5] At 00:11.400, the shot cuts hard to an extreme close-up of the young woman's face as she whips around to look back and down the steps, eyes blown wide and her straw hat tipping back off her head. The young woman (S1) lets out a sharp, <d>[English] Aah!</d> [Shot 6] At 00:12.400, the shot cuts hard to an extreme close-up of the young wizard's face as he flinches violently backward, his wide blue eyes blown huge beneath the brim of his hat. The young wizard (S3) yelps, <d>[English] Aah!</d> then throws his free right hand up palm out, his left hand still gripping the staff upright beside him, and blurts, <d>[English] No, no! I just wanted to tell you something!</d> The clip ends with both of them frozen at either end of the steps, neither moving.\n \noverall_soundscape: Dense night crickets and a low breeze move through foliage over distant town sounds. Wooden shutters bang hard against the wall and clatter shut again. Sandals scuff carefully on stone treads. Two sharp startled cries ring out one after the other and echo off the walls, followed by fabric snapping as a hand flies up and a foot scuffing back on gravel.\n \nnon_diegetic_music: A soft solo acoustic guitar at a slow tempo cuts out entirely on the first startled cry, leaving a single low plucked note ringing over the final standoff.\n \n \n================================================================\nEND — 7 scenes, 102 seconds total\n================================================================</code></pre>\n<p>example 2</p>\n<p>from <a href=\"https://www.reddit.com/r/StableDiffusion/comments/1vkw1hc/minimax_h3_ref2v_can_somewhat_do_smash_bros/\">https://www.reddit.com/r/StableDiffusion/comments/1vkw1hc/minimax_h3_ref2v_can_somewhat_do_smash_bros/</a></p>\n<pre class=\"language-json\"><code>\n General prompt:\n\nsubject_definitions:\n\n<Subject 1> is the woman in <Picture 1>. <Subject 2> is the woman in <Picture 2>.\n\nsummary:\n\n[reference generation] Super Smash Bros Ultimate gameplay, with <Subject 1> fighting against <Subject 2>.\n\nretention_analysis:\n\n<Subject 1>: fully-preserved - <Subject 1> retains all attributes.\n\n<Subject 2>: fully-preserved - <Subject 2> retains all attributes.\n\ndetailed_description:\n\nA Super Smash Bros Ultimate match on the stage Final Destination. There is a UI on the bottom of the screen denoting percentage values for <Subject 1> and <Subject 2>. <Subject 1> character portrait is on the bottom left hand side, with the text \"0%\" written next to the portrait. <Subject 2> character potrait is on the bottom right hand side, with the text \"0% written next to the portrait.\n\n[Shot 1] <Subject 1> stands on the left hand side of the map, while <Subject 2> stands on the right hand side.\n\n[Shot 2] At 00:00.500, <Subject 1> moves towards <Subject 2>, and does three light jabs with her fists, then does a sweeping kick into an up tilt attack. <Subject 1> jumps once into the air, and does forward air attack on <Subject 2> who is still in the air, sending <Subject 2> off of the map. <Subject 2> character portrait percentange number climbs up to 30%.\n\n[Shot 3] At 00:05.000, <Subject 2> jumps back onto the map, and does a forward air attack on <Subject 1>, making <Subject 1> tumble backwards. <Subject 2> then runs up to <Subject 1>, and grabs <Subject 1>, then side throws <Subject 1> into the air. <Subject 1> then lands on the floor, and <Subject 2> runs into a dash attack into a couple jabs onto <Subject 1>, then does a smash attack, dealing a ton of damage, and sending <Subject 1> off of the map. <Subject 1> character portrait percentage now reads \"48%\".\n\n[Shot 4] At 00:11.000, as <Subject 1> is attempting to jump back to the main stage, <Subject 2> jumps off of the map towards <Subject 1>, and uses her her arm to do an overhead arc punch on <Subject 1>, spiking <Subject 1> down off of the screen, making her hit the blast zone.\n\noverall_soundscape: Quiet, subtle wind sounds. </code></pre>\n<p>example 3</p>\n<p><a href=\"https://www.reddit.com/r/StableDiffusion/comments/1vl3uu4/minimax_h3_testing_l2va_moon_landing/\">https://www.reddit.com/r/StableDiffusion/comments/1vl3uu4/minimax_h3_testing_l2va_moon_landing/</a></p>\n<pre class=\"language-json\"><code> integrated_multimodal_description: Time-lapsed, cinematic, a medium-wide shot of a film set in a large indoor studio which is used to shoot a scene of moon landing involving lunar module, US flag and an astronaut. At 00:00.000 the camera shows an empty, sterile white studio room, with recognizable vertical wand in the back and horizontal floor at its bottom. At 00:01.000 Some film crews install a black wand into the studio's vertical wand. The black wand has some tiny white shining points which represent stars. At 00:02.000 Some workers fill the studio's floor with some dirty-white sand, gravels and small rocks and form a barren lunar landscape. At 00:03.000 Some film crews bring an Apollo Lunar Module and place it into the left side of the scene. At 00:04.000 A film crew places a US flag with pole on the right side of the scene. Another film crew puts a picture of the Earth as the blue planet, partially blacked on its bottom side, on the top right corner of the scene. At 00:05:000 An astronaut walks in into the scene, goes into the middle of the scene, faces to viewer and waves his hand. At 00:07:000 The whole scene settles into the exact arrangement, position of subject and objects, camera angle, lighting, and final composition established by <Picture 1>. A male deep voice of the director (S1) says, <d>[English] Cut!</d>\n\noverall_soundscape:\n\nnon_diegetic_music: Sustained violin notes at a very fast tempo with spaced piano tones. </code></pre>\n<p>example 4</p>\n<p><a href=\"https://www.reddit.com/r/comfyui/comments/1vidio0/minimax_h3_benchmark_on_rtx_pro_6000_blackwell/\">https://www.reddit.com/r/comfyui/comments/1vidio0/minimax_h3_benchmark_on_rtx_pro_6000_blackwell/</a></p>\n<pre class=\"language-json\"><code>Single continuous cinematic shot at blue hour in a rain-wet city plaza. A woman in a bright red coat walks briskly toward camera while opening a transparent umbrella; wind moves her coat, hair, and the umbrella naturally. A cyclist crosses behind her from left to right, reflected neon signs ripple in puddles, and passing headlights create moving highlights on the wet pavement. The camera performs a smooth low-angle backward tracking move with realistic parallax, stable anatomy, detailed hands, natural facial motion, and consistent objects. She looks into camera and clearly says, \"The storm is finally passing.\" Audio: synchronized adult female voice, footsteps splashing through shallow puddles, umbrella fabric snapping softly in the wind, a bicycle bell behind her, distant traffic, light rain, and subtle restrained electronic music. No captions, subtitles, logos, cuts, slow motion, duplicated people, or warped objects. </code></pre>\n<p>example 5</p>\n<p>This is a T2V test with Japanese dialogue Eng subtitle and action scene, with no reference image or materials. </p>\n<p><a href=\"https://www.reddit.com/r/comfyui/comments/1vjrs7i/h3_cinematic_action_scene_test_by_5060ti_16gb/\">https://www.reddit.com/r/comfyui/comments/1vjrs7i/h3_cinematic_action_scene_test_by_5060ti_16gb/</a></p>\n<pre class=\"language-json\"><code> integrated_multimodal_description: [Shot 1] Live-action, high-budget cinematic prestige drama style, an extreme wide shot establishes a dimly lit, high-tech subterranean military corridor with dark brushed-metal walls and harsh ambient lighting. A stylish 20-year-old Japanese female secret agent with short sharp dark hair, wearing a sleek black tactical bodysuit, slips swiftly through a heavy mechanical blast door. The camera tracks left with large amplitude at fast speed alongside her movement. She taps her earpiece and, as a stylish 20-year-old Japanese female secret agent with a tense, focused low whisper (S1), says: <d>[Japanese] ターゲットの端末に到達した。</d> English subtitles at the bottom of the frame read \"Target terminal reached.\"\n\n[Shot 2] At 00:03.000, the shot cuts to a close-up of her focused face and dark eyes reflecting a glowing blue console as her gloved fingers rapidly operate the interface. Suddenly, the screen flashes bright red with a warning icon. Red emergency alarm lights wash over her face. The camera pushes in with small amplitude at fast speed toward her eyes as she (S1) turns her head sharply toward the hallway, exclaiming in a panicked whisper: <d>[Japanese] しまった、トラップか!</d> English subtitles at the bottom read \"Dammit, it's a trap!\"\n\n[Shot 3] At 00:06.000, the camera cuts to a dynamic medium shot as heavy metal doors in the background burst open, revealing armed tactical soldiers pointing red laser sights into the room. The camera arc shots around her at fast speed as she vaults over a metal desk, narrowly dodging laser beams cutting through the dark haze.\n\n[Shot 4] At 00:09.000, the shot cuts to a low-angle close-up of the agent drawing a silenced tactical pistol from her holster. She hurls a smoke grenade toward the floor, spins directly toward the camera, and fires upward. Sparks burst violently from the overhead light fixture, throwing the frame into high-contrast silhouettes as smoke fills the lens.\n\noverall_soundscape: Quiet stealthy boot steps suddenly break into a loud, echoing mechanical alarm siren with reverberating horns. Heavy blast doors slide open with a loud pneumatic hiss, accompanied by heavy tactical boot thuds, shouting guards, sharp electrical spark crackles, and smoke grenade canister hiss.\n\nnon_diegetic_music: A high-octane cinematic action-trailer score featuring an aggressive synth-bass pulse, fast-pacing orchestral percussion, heavy brass swells, and a dramatic riser crescendo that suddenly cuts out at the end. </code></pre>\n<p>example 6</p>\n<p><a href=\"https://www.reddit.com/r/comfyui/comments/1vjk2dr/workflow_for_existing_image_scene_edit_minimax_h3/\">https://www.reddit.com/r/comfyui/comments/1vjk2dr/workflow_for_existing_image_scene_edit_minimax_h3/</a></p>\n<pre class=\"language-json\"><code>The Matrix (1999) original trilogy film footage. Green color tint. Static shot. Static camera.\n<Subject 1> is actor Hugo Weaving at his fourties. He plays agent Smith. He wears sunglasses and black suit with white collared shirt and black tie.\n[Shot 1] Face close-up shot of <Subject 1> taking off sunglasses and staring angrily at someone outside the frame at right. After a 1.00 second he says slowly: <d>[English] I hate AI slop.</d>\n[Shot 2] <Subject 1> face grimaces with extreme disgust and mild anger as he looks to left, upwards and to right.\nMusic: dramatic ambient music.</code></pre>\n<p>example 7</p>\n<p><a href=\"https://www.reddit.com/r/comfyui/comments/1vjf4pb/h3_just_blows_my_mind_generated_on_a_4070ti_super/\">https://www.reddit.com/r/comfyui/comments/1vjf4pb/h3_just_blows_my_mind_generated_on_a_4070ti_super/</a></p>\n<pre class=\"language-json\"><code> \"<Subject 1> is the wiry, sun-browned corporal in <Picture 1>, dark stubble, a healed scar through one eyebrow, olive-drab fatigues with webbing and a canvas satchel. fully_preserved.\n\n<Subject 2> is the young, freckled rifleman in <Picture 2>, pale blue eyes, olive-drab M-1943 field jacket, netted M1 helmet with a loose chinstrap. fully_preserved.\n\n<Picture 3> is the landing craft interior location reference — the empty foreground bench, riveted hull, and defocused soldiers. partially_preserved.\n\nintegrated_multimodal_description:\n\nHandheld 16mm documentary combat footage, desaturated cold blue-grey grade, heavy film grain, shallow depth of field. <Subject 1> and <Subject 2> sit shoulder to shoulder on the foreground bench from <Picture 3>, in that same riveted hull, packed defocused soldiers swaying behind them, spray drifting over the gunwale, the deck pitching with the swell. Both men wear M1 helmets.\n\n[Shot 1] Medium two-shot, eye-level, handheld with Shake Slightly at small amplitude, drifting with the boat's roll. <Subject 2> stares at the deck, chin trembling, eyes glassy and brimming, and speaks without looking up, his thin young voice tightening and cracking mid-phrase as he fights not to cry: <d>[English] I told my ma I'd be home by harvest. I keep— I keep thinkin' about her porch light. You think we'll make it?</d> His jaw quivers on the last word and he presses his lips flat. <scenetrans>\n\n[Shot 2] At 00:05.500, <scenetrans> the shot cuts to a medium close-up on <Subject 1>, handheld, Shake Slightly with small amplitude. The engine and sea carry over seamlessly across the cut. His face does not move; eyes fixed forward on the ramp, jaw set. After a long beat he answers, voice low, level, completely without inflection: <d>[English] Some of us will.</d> <scenetrans>\n\n[Shot 3] At 00:08.000, <scenetrans> the shot cuts back to a medium close-up on <Subject 2>, handheld, Shake Slightly with small amplitude. He nods once, very small, jaw clenched against it, and a single tear breaks down his freckled cheek as he turns his face slowly toward the bow, breathing unsteady, the ambient roar continuing uninterrupted as a distant shell rumble rolls through.\n\noverall_soundscape:\n\nConstant diesel engine throb and hull slap against swell as the bed throughout, sea spray hissing over the gunwale, gear and webbing creaking as men sway. The young voice is close and raw, thin, wavering, audibly cracking; the older voice is low, dry, flat. A wet sniff before the first line, and near the end a swallowed, shaky breath that is almost a sob, half-buried under the engine. A distant naval bombardment rumble swells low in the final seconds.\n\nnon_diegetic_music:\n\nN/A\" </code></pre>\n<p>example 8 --contains a few scenes</p>\n<p><a href=\"https://www.reddit.com/r/comfyui/comments/1vjce6v/3minute_ai_music_video_test_on_a_rtx_3090_with/\">https://www.reddit.com/r/comfyui/comments/1vjce6v/3minute_ai_music_video_test_on_a_rtx_3090_with/</a></p>\n<pre class=\"language-json\"><code>subject_definitions:\n<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.\n<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable vehicle identity, white exterior body, compact proportions, rounded front shape, dark windshield and side glass, wheel design and placement, front lighting arrangement, roof-mounted autonomous-driving sensor assembly, and overall silhouette. Ignore the daylight street, palm trees, buildings, and environment visible in <Picture 2>; only the vehicle itself is referenced.\n\nsummary:\n[reference generation] Place <Subject 1> inside <Subject 2>, cruising through the same neon cyberpunk downtown at night during one continuous fifteen-second establishing shot, ending in a clean side-tracking composition that naturally leads into the next scene.\n\nretention_analysis:\n<Subject 1> (appears throughout [Shot 1]): fully_preserved - identity, recognizable face, facial proportions, hair, body proportions, pale blue uniform, cap, black tie, metallic robotic hands, footwear, and characteristic expression remain unchanged; only nighttime lighting, seating position, and subtle rhythmic movement are new.\n<Subject 2> (appears throughout [Shot 1]): fully_preserved - exact body design, white exterior, dark glass, wheels, lighting geometry, roof-mounted autonomous sensor equipment, proportions, and recognizable silhouette remain unchanged; only the environment, reflections, movement, and nighttime lighting are new.\n\ndetailed_description:\nThe target video is a fifteen-second photorealistic cinematic single take in a neon cyberpunk future photographed like an expensive 1987 science-fiction movie. The entire music video occurs during the same night in the same downtown district: wet black asphalt, massive brutalist concrete towers, practical cyan neon tubes, restrained magenta accent lights, deep blue-black shadows, chrome reflections, thin drifting steam, atmospheric haze, soft diffusion, subtle 35mm film grain, horizontal anamorphic lens flares, and believable physical lighting. Avoid a modern glossy CGI look.\n\n[Shot 1] Begin from a low wide camera position approximately one meter above the wet boulevard. <Subject 2> appears far down the street and approaches smoothly through the neon city. The camera begins tracking backward at approximately the same speed, maintaining a stable front three-quarter view as the vehicle gradually becomes larger in frame. Cyan architectural lights and small magenta highlights travel naturally across the exact white body and dark windows of <Subject 2>.\n\nAs the vehicle approaches, reveal <Subject 1> clearly through the windshield, seated calmly in the front cabin. Cyan dashboard light softly illuminates their recognizable face, pale blue uniform, cap, black tie, and metallic robotic hands. <Subject 1> looks calmly forward with the same cheerful uncanny expression from <Picture 1>. Their right metallic hand rests naturally while two fingers gently tap an implied Italo-disco rhythm.\n\nKeep identity, hands, seating position, car geometry, reflections, and camera movement physically stable.\n\nDuring the final five seconds, the camera smoothly arcs from the front three-quarter position toward the left side of <Subject 2> without cutting and without changing speed.\n\nEND STATE / TRANSITION: finish on a stable medium side-profile tracking composition of <Subject 2> traveling from left to right, with <Subject 1> clearly visible through the side window. The next clip begins from this exact motion direction, framing, city block, and lighting state.\n\nNo redesign of either subject, no different vehicle, no costume change, no daytime environment, no palm trees, no added text, no subtitles, no generated logos.\n\noverall_soundscape:\nNone required. Visual generation only; final song and sound design will be added during editing.\n\nnon_diegetic_music:\nNone generated. The final Italo-disco track will be added separately.\n\nsubject_definitions:\n<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.\n<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable vehicle identity, white exterior body, compact proportions, rounded front shape, dark windshield and side glass, wheel design and placement, front lighting arrangement, roof-mounted autonomous-driving sensor assembly, and overall silhouette. Ignore the daylight street, palm trees, buildings, and environment visible in <Picture 2>; only the vehicle itself is referenced.\n\nsummary:\n[reference generation] Continue <Subject 1> riding inside <Subject 2> along the same neon boulevard during one uninterrupted fifteen-second side-tracking shot, emphasizing autonomous driving and restrained rhythmic character movement before approaching a cyan-lit intersection.\n\nretention_analysis:\n<Subject 1> (appears throughout [Shot 1]): fully_preserved - exact face, identity, body proportions, pale blue uniform, cap, black tie, metallic hands, footwear, and expression remain recognizable and unchanged; only head direction and small seated dance gestures change.\n<Subject 2> (appears throughout [Shot 1]): fully_preserved - same exact vehicle body, white exterior, glass, wheels, roof sensor assembly, lighting arrangement, and proportions remain unchanged; only motion and neon reflections change.\n\ndetailed_description:\nThe target video is a fifteen-second photorealistic cinematic single take continuing directly from the previous scene. Same neon cyberpunk downtown district, same night, same wet boulevard, same cyan-dominant practical lighting, restrained magenta accents, brutalist architecture, chrome reflections, atmospheric steam, deep blue shadows, anamorphic horizontal flares, soft diffusion, subtle 35mm grain, authentic 1980s science-fiction cinematography.\n\n[Shot 1] Begin immediately in the exact side-profile tracking composition established previously: <Subject 2> moves smoothly from left to right while the camera travels perfectly parallel at the same speed and distance.\n\nKeep the full recognizable side profile of <Subject 2> visible. Long cyan reflections and occasional magenta highlights slide naturally across its white body and black glass without altering its physical design.\n\nThrough the side window, <Subject 1> is clearly visible in the front cabin. Maintain the exact recognizable face and outfit from <Picture 1>. <Subject 1> initially looks forward, then slowly turns their head slightly toward camera.\n\n<Subject 1> deliberately lifts both metallic robotic hands completely away from the vehicle controls, showing that <Subject 2> is operating autonomously. Without exaggeration, <Subject 1> performs a restrained seated Italo-disco movement: two small shoulder pulses, one subtle head nod, and one metallic index finger briefly pointing upward before relaxing again.\n\nSparse pedestrians and one cyclist may move through the distant background, but never obscure either referenced subject.\n\nDuring the final four seconds, the camera smoothly advances from the pure side view into a front-left three-quarter tracking position as <Subject 2> approaches a large intersection illuminated by cyan traffic lights.\n\nEND STATE / TRANSITION: finish with <Subject 2> entering the intersection in a stable front-left three-quarter composition, still moving forward at controlled city speed, with <Subject 1> clearly visible through the windshield. The next scene begins from this exact position and direction.\n\nNo camera cuts, no high-speed driving, no new neighborhood, no vehicle redesign, no wardrobe change, no daytime, no readable text or generated logos.\n\noverall_soundscape:\nNone required. Visual generation only.\n\nnon_diegetic_music:\nNone generated. Final song added in post-production.\n\nsubject_definitions:\n<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.\n<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable vehicle identity, white exterior body, compact proportions, rounded front shape, dark windshield and side glass, wheel design and placement, front lighting arrangement, roof-mounted autonomous-driving sensor assembly, and overall silhouette. Ignore the daylight environment visible in <Picture 2>; use only the vehicle as reference.\n\nsummary:\n[reference generation] Continue <Subject 2> through the same cyan-lit city intersection while pedestrians and cyclists cross safely, with <Subject 1> calmly acknowledging a cyclist before the vehicle approaches the familiar nightlife curb.\n\nretention_analysis:\n<Subject 1> (appears throughout [Shot 1]): fully_preserved - identity, face, outfit, proportions, metallic hands, cap, footwear, and expression remain unchanged; only a small two-finger gesture is introduced.\n<Subject 2> (appears throughout [Shot 1]): fully_preserved - exact vehicle identity, white exterior, body geometry, windows, wheels, sensors, lights, and proportions remain unchanged; only speed adjusts smoothly to surrounding traffic.\n\ndetailed_description:\nThe target is a fifteen-second photorealistic single-take continuation in the exact same cyberpunk downtown district during the same night. Authentic 1987 science-fiction cinema aesthetic: practical cyan neon, limited magenta accents, wet reflective road surface, heavy concrete buildings, atmospheric haze, thin steam, chrome highlights, deep blue-black shadows, subtle film grain, soft diffusion and horizontal anamorphic flares.\n\n[Shot 1] Begin with <Subject 2> already entering the wide cyan-lit intersection in the same front-left three-quarter tracking composition established in the previous scene. Camera continues moving backward smoothly ahead of the vehicle.\n\nSeveral pedestrians begin crossing far enough ahead to remain safe and visually clear. Two cyclists travel through a protected bicycle lane from right to left. Their motion is calm and natural, creating an elegant coordinated urban flow rather than danger.\n\n<Subject 2> gently reduces speed without abrupt braking, maintaining perfectly stable geometry and orientation.\n\n<Subject 1> remains clearly visible through the windshield. Preserve the exact recognizable face and blue uniform. As one cyclist passes, <Subject 1> lifts a metallic hand and gives a small friendly two-finger salute, then lowers it naturally. The distinctive cheerful expression remains unchanged.\n\nAfter the crossing clears, <Subject 2> resumes smooth movement.\n\nThe camera slowly arcs toward the vehicle's right-front side while keeping both <Subject 1> and the recognizable front geometry of <Subject 2> visible.\n\nAhead, reveal the same nightlife block under a long cyan neon canopy, located immediately beyond the intersection.\n\nDuring the final seconds, <Subject 2> moves gently toward the curb beneath that canopy.\n\nEND STATE / TRANSITION: finish with a stable low front-side composition of <Subject 2> approaching the cyan-lit curb and beginning to slow, with <Subject 1> still visible inside. The next scene begins at this exact curb approach.\n\nNo collision, no abrupt maneuver, no crowd chaos, no new vehicle, no character alteration, no new district, no text or subtitles.\n\noverall_soundscape:\nNone required. Visual generation only.\n\nnon_diegetic_music:\nNone generated. Final music added separately.\n\nsubject_definitions:\n<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.\n<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable vehicle identity, white exterior body, compact proportions, rounded front, dark glass, wheels, front lights, roof-mounted autonomous-driving sensor assembly, and overall silhouette. Ignore the daylight environment of <Picture 2>.\n\nsummary:\n[reference generation] Continue <Subject 2> stopping beneath the cyan canopy, then have <Subject 1> step out and perform a restrained Italo-disco gesture beside the exact same car in one continuous fifteen-second shot.\n\nretention_analysis:\n<Subject 1> (appears throughout [Shot 1]): fully_preserved - recognizable face, identity, clothing, cap, tie, proportions, robotic hands, shoes, and expression remain unchanged; only posture changes from seated to standing and a simple dance gesture is added.\n<Subject 2> (appears throughout [Shot 1]): fully_preserved - exact vehicle body, glass, sensor equipment, wheels, lighting and proportions remain unchanged and clearly visible beside the character.\n\ndetailed_description:\nThe target video is a fifteen-second photorealistic cinematic single take in the same cyan-lit cyberpunk nightlife block, same night and same 1980s visual language: wet pavement, brutalist concrete facades, practical cyan canopy lighting, minimal magenta accents, chrome reflections, drifting steam, blue-black shadows, subtle grain, soft diffusion and anamorphic flares.\n\n[Shot 1] Begin exactly with <Subject 2> approaching the familiar curb beneath the cyan canopy.\n\nThe camera tracks slowly beside the vehicle at low chest height.\n\n<Subject 2> gently pulls into position and comes to a controlled stop. Hold long enough to establish that the exact reference vehicle remains visually stable.\n\nThrough the window, <Subject 1> is clearly visible.\n\nThe vehicle door opens naturally.\n\n<Subject 1> steps out onto the wet pavement, one metallic hand briefly touching the door frame for physical stability. Preserve the exact face, proportions, blue suit, round cap, tie, robotic hands and shoes from <Picture 1>.\n\nThe camera gradually pulls backward while staying low enough to keep <Subject 1> and most of <Subject 2> together in frame.\n\n<Subject 1> adjusts the front of the pale blue suit using both metallic hands, then performs a deliberately simple Italo-disco phrase: one side step, second side step, two restrained shoulder pulses, then one metallic finger points directly toward camera.\n\nNo complex dance choreography.\n\nDuring the final three seconds, <Subject 1> relaxes the pose, turns slightly and leans casually against the front side of <Subject 2>.\n\nEND STATE / TRANSITION: medium hero composition with <Subject 1> leaning beside <Subject 2> under the cyan canopy, both identities clearly readable and physically stable. The next clip begins from this exact arrangement.\n\nNo cuts, no additional performers, no costume change, no vehicle redesign, no daylight, no generated signage or subtitles.\n\noverall_soundscape:\nNone required. Visual generation only.\n\nnon_diegetic_music:\nNone generated. Final Italo-disco music added in edit.\n\nsubject_definitions:\n<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.\n<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable identity, white exterior, compact rounded body, dark glass, wheels, lighting, roof-mounted autonomous sensor equipment, dimensions and silhouette. Ignore the original daylight surroundings.\n\nsummary:\n[reference generation] Keep <Subject 1> dancing beside the parked <Subject 2> beneath the same cyan canopy in a single restrained 1980s performance shot, ending with <Subject 1> standing at the vehicle door ready to enter.\n\nretention_analysis:\n<Subject 1> (appears throughout [Shot 1]): fully_preserved - exact identity, facial appearance, costume, proportions, metallic hands, cap and shoes remain unchanged during simple controlled choreography.\n<Subject 2> (appears throughout [Shot 1]): fully_preserved - remains stationary and visually identical to <Picture 2>, with only environmental neon reflections added.\n\ndetailed_description:\nFifteen-second photorealistic single-take performance in the exact same nightlife curb location. Authentic 1980s cyberpunk film appearance: practical cyan neon canopy, tiny magenta accents, wet pavement, dark brutalist architecture, chrome reflections, steam, blue-black shadows, soft optical bloom, anamorphic flare and subtle 35mm grain.\n\n[Shot 1] Begin with <Subject 1> leaning naturally against the front side of <Subject 2>, exactly matching the previous ending.\n\nThe camera begins a slow clockwise orbit around the character and vehicle together. Keep both reference subjects visible for almost the entire shot.\n\n<Subject 1> gently pushes away from the vehicle and begins a restrained, repeatable Italo-disco dance phrase designed to preserve identity: two lateral steps, one controlled shoulder roll, metallic right hand sweeps horizontally across the chest, left metallic index finger points upward, a small pivot, then two measured steps backward.\n\nKeep limb proportions stable and movements humanly achievable. Preserve the exact uncanny friendly facial expression.\n\n<Subject 2> remains parked in precisely the same position. Cyan reflections travel naturally over its white panels and dark glass, but the vehicle shape and sensor equipment never change.\n\nThe wet ground produces soft reflections of both subjects.\n\nAs the orbit approaches completion, <Subject 1> stops dancing, turns toward <Subject 2>, walks the short distance to the door and reaches for the opening.\n\nEND STATE / TRANSITION: <Subject 1> stands immediately beside the open door of <Subject 2>, one metallic hand resting on the door frame, body oriented toward the cabin and ready to sit. The next scene begins here.\n\nNo new vehicle, no background change, no additional dancers, no body deformation, no wardrobe change, no readable text.\n\noverall_soundscape:\nNone required.\n\nnon_diegetic_music:\nNone generated. Music added separately.\n\nsubject_definitions:\n<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.\n<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable identity, white exterior, body shape, proportions, dark glass, wheels, front and rear lighting, roof-mounted sensor equipment and overall silhouette. Ignore the sunny location visible in <Picture 2>.\n\nsummary:\n[reference generation] Continue <Subject 1> entering <Subject 2>, closing the door and smoothly departing the same cyan curb during one continuous fifteen-second tracking shot.\n\nretention_analysis:\n<Subject 1> (appears throughout [Shot 1]): fully_preserved - face, identity, clothing, accessories, proportions and robotic hands remain unchanged while transitioning naturally from standing to seated.\n<Subject 2> (appears throughout [Shot 1]): fully_preserved - vehicle identity and geometry remain constant through stationary and moving states.\n\ndetailed_description:\nThe target is a fifteen-second photorealistic continuous shot maintaining the exact same downtown curb, same cyan lighting, wet street, brutalist architecture, subtle magenta accents, haze, steam, anamorphic bloom, soft diffusion and 1980s film grain.\n\n[Shot 1] Begin with <Subject 1> standing beside the already open door of <Subject 2>, one metallic hand on the upper door frame.\n\nWithout cutting, <Subject 1> smoothly lowers into the front cabin. Maintain realistic limb articulation and stable body proportions. Both metallic hands move naturally inside, followed by the legs and black shoes.\n\nThe door closes.\n\nThe camera remains outside and begins gliding parallel along the side window as <Subject 2> gently pulls away from the curb.\n\nThrough the glass, keep <Subject 1> clearly recognizable under cyan dashboard illumination. <Subject 1> looks forward and taps one metallic hand lightly against the upper leg in a restrained rhythmic pattern.\n\n<Subject 2> smoothly merges back into the exact same wet boulevard.\n\nCamera continues beside the car for several seconds without changing distance abruptly.\n\nDuring the final four seconds, camera gradually reduces speed while <Subject 2> maintains forward motion. The car naturally moves ahead until camera settles into a rear-left three-quarter view.\n\nEND STATE / TRANSITION: stable rear-left tracking view of <Subject 2> traveling away along the familiar cyan-lit boulevard. The next clip begins directly from behind this moving vehicle.\n\nNo cuts, no teleportation, no different vehicle, no character change, no new architecture, no text.\n\noverall_soundscape:\nNone required.\n\nnon_diegetic_music:\nNone generated.\n\nsubject_definitions:\n<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.\n<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable identity, white exterior, compact proportions, rounded geometry, dark windows, wheels, lighting arrangement, roof-mounted autonomous sensor system and silhouette. Ignore the daytime background from <Picture 2>.\n\nsummary:\n[reference generation] Follow <Subject 2> through the familiar neon downtown while <Subject 1> rides inside, ending with the vehicle entering a long cyan tunnel during one uninterrupted fifteen-second night-driving shot.\n\nretention_analysis:\n<Subject 1> (appears through vehicle glass): fully_preserved - same identity, facial appearance, costume, proportions, cap and robotic hands remain consistent.\n<Subject 2> (primary visual subject): fully_preserved - exact shape, white body, dark glazing, wheels, sensor equipment and proportions remain unchanged during the entire drive.\n\ndetailed_description:\nFifteen-second photorealistic continuous tracking shot in the same cyberpunk downtown, same night and same 1980s film aesthetic: practical cyan architectural lights, minimal magenta accents, wet asphalt, concrete towers, thin steam, deep blue shadows, chrome reflections, soft diffusion, anamorphic streaks and subtle grain.\n\n[Shot 1] Begin directly behind and slightly left of <Subject 2>, matching the rear-left three-quarter ending of the previous clip.\n\nCamera travels at approximately the same speed and maintains a consistent following distance.\n\n<Subject 2> drives calmly through the established downtown boulevard. Wet pavement reflects the white vehicle and repeating cyan architecture.\n\nSparse pedestrians remain safely on the sidewalks. One cyclist travels in a separated lane.\n\nThe road gradually curves to the right. Camera follows the same smooth arc and slowly moves closer toward the left side of the vehicle.\n\nThrough the dark side glass, briefly reveal <Subject 1> seated comfortably in the front cabin, still wearing the exact pale blue uniform and cap. <Subject 1> gives one gentle head nod and one small shoulder movement while looking forward.\n\nDo not make <Subject 1> dominate this shot; the drive itself is the focus.\n\nAhead, reveal a long rectangular road tunnel built into the same downtown architecture. The tunnel entrance is illuminated by repeating cyan rectangular lights.\n\n<Subject 2> aligns smoothly with the tunnel entrance.\n\nEND STATE / TRANSITION: centered rear view of <Subject 2> just beginning to cross into the cyan tunnel, with the repeating light geometry visible ahead. Next clip begins inside this exact tunnel.\n\nNo new environment, no speed racing, no vehicle mutation, no daylight, no captions.\n\nsubject_definitions:\n<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.\n<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable identity, white exterior, compact proportions, rounded front geometry, black glass, wheels, lights, roof-mounted autonomous sensor assembly, and overall silhouette.\n\nsummary:\n[reference generation] Follow <Subject 2> through the cyan tunnel, gradually move alongside it and reveal <Subject 1> taking both robotic hands away from the controls for a small seated disco gesture before the car reaches the tunnel exit.\n\nretention_analysis:\n<Subject 1> (appears prominently in second half of [Shot 1]): fully_preserved - exact identity, face, blue clothing, cap, tie, metallic hands and body proportions remain unchanged; only restrained arm and shoulder movement is introduced.\n<Subject 2> (appears throughout [Shot 1]): fully_preserved - exact vehicle geometry, sensors, windows, body panels, wheels and color remain recognizable and constant under moving tunnel lights.\n\ndetailed_description:\nThe target is a fifteen-second photorealistic uninterrupted shot inside the same urban tunnel. Authentic 1980s science-fiction cinematography: repeating cyan practical light rectangles, dark concrete walls, wet pavement, occasional subtle magenta reflection, light atmospheric haze, anamorphic streaking, soft diffusion and tactile film grain.\n\n[Shot 1] Begin directly behind <Subject 2> as it completes entry into the cyan tunnel.\n\nCamera follows the vehicle at identical speed. Repeating cyan light bands travel rhythmically across the exact white body, dark windows and roof-mounted sensor system without changing their physical forms.\n\nAfter several seconds, camera slowly moves from directly behind to the left side of <Subject 2>, arriving at a clean parallel tracking composition.\n\nThrough the side window, clearly reveal <Subject 1> in the front cabin.\n\nMaintain exact facial identity and outfit. <Subject 1> calmly lifts both metallic hands completely away from the controls and brings them loosely to chest height.\n\nPerform only a tiny seated Italo-disco gesture: two synchronized metallic fingertip taps in empty air, one subtle shoulder pulse and one relaxed head nod.\n\n<Subject 2> continues perfectly straight without visible human control.\n\nThe effect should feel confident, cool and slightly humorous, never slapstick.\n\nDuring the final four seconds, cyan and magenta city lights become visible beyond the tunnel exit. Camera gradually advances into a front-left side position.\n\nEND STATE / TRANSITION: front-left side tracking view of <Subject 2> precisely at the tunnel exit, with the familiar nighttime city visible immediately beyond. Next scene continues the same forward movement.\n\nNo visual transformation, no speed jump, no different vehicle, no costume changes, no text.\n\noverall_soundscape:\nNone required.\n\nnon_diegetic_music:\nNone generated.\n\nsubject_definitions:\n<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.\n<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable vehicle identity, white body, dimensions, rounded design, dark glass, wheel geometry, front lighting, roof-mounted sensor equipment and silhouette. Ignore the original daylight environment.\n\nsummary:\n[reference generation] Continue <Subject 2> exiting the cyan tunnel and traveling through the familiar downtown while <Subject 1> performs a slightly more energetic seated disco gesture, ending with the car stopped beneath the previously established cyan canopy.\n\nretention_analysis:\n<Subject 1> (appears prominently through windshield): fully_preserved - face, identity, body proportions, clothing, cap, tie and metallic hands remain exact; only controlled rhythmic motion changes.\n<Subject 2> (appears throughout [Shot 1]): fully_preserved - vehicle identity, body, windows, wheels, sensor system and lighting remain unchanged.\n\ndetailed_description:\nFifteen-second photorealistic continuous shot, same downtown, same night, same weather and exact same 1980s cyberpunk color grade: cyan practical lights, restrained magenta highlights, wet asphalt, brutalist facades, steam, deep shadows, analog diffusion, anamorphic lens streaks and 35mm grain.\n\n[Shot 1] Begin as <Subject 2> exits the cyan tunnel from the front-left side composition established previously.\n\nCamera smoothly transitions into a low front-left three-quarter tracking position while moving backward at identical speed.\n\nThe vehicle's white surface reflects long cyan lines and occasional magenta highlights from the same familiar architecture.\n\nThrough the windshield, <Subject 1> is clearly visible and slightly more animated than before while remaining physically stable.\n\n<Subject 1> performs two gentle shoulder pulses, one head nod and then raises one metallic hand for a playful forward finger point.\n\nThe autonomous vehicle continues operating smoothly and safely.\n\nCamera gradually gets closer to the windshield while maintaining enough visible vehicle body to preserve <Subject 2>'s identity.\n\n<Subject 1> slowly turns toward the camera and shows the same recognizable friendly uncanny smile.\n\nDuring the last four seconds, <Subject 2> slows and returns to the exact same curb beneath the cyan canopy used earlier.\n\nEND STATE / TRANSITION: <Subject 2> completely stopped beneath the familiar canopy, viewed from a stable front-side position, with <Subject 1> visible through the window looking toward camera. The next scene begins from this exact setup.\n\nNo alternate neighborhood, no new car, no daylight, no character redesign, no text.\n\nsubject_definitions:\n<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.\n<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable vehicle identity, white exterior, body design, proportions, windows, wheels, lighting and roof-mounted autonomous sensor assembly. Ignore the sunny environment in <Picture 2>.\n\nsummary:\n[reference generation] Have <Subject 1> step out of the stopped <Subject 2> beneath the familiar cyan canopy, perform one final simple Italo-disco dance, then return to the vehicle in one coherent continuous fifteen-second performance shot.\n\nretention_analysis:\n<Subject 1> (appears throughout [Shot 1]): fully_preserved - exact facial identity, pale blue uniform, cap, tie, metallic hands, footwear and proportions remain unchanged during restrained choreography.\n<Subject 2> (appears throughout [Shot 1]): fully_preserved - remains parked and visually identical to the reference while serving as a stable visual anchor.\n\ndetailed_description:\nThe target is a fifteen-second photorealistic continuous performance shot in the exact same cyan-lit curb location established earlier. Same wet pavement, same concrete facades, same practical cyan lighting, same subtle magenta accents, atmospheric steam, deep blue-black shadows, anamorphic flares, diffusion and textured 1980s film grain.\n\n[Shot 1] Begin with <Subject 2> stopped beneath the cyan canopy and <Subject 1> visible through the side window.\n\nThe vehicle door opens smoothly.\n\n<Subject 1> steps out naturally and stands beside the exact car.\n\nCamera begins slowly pulling backward as <Subject 1> walks two measured steps toward lens.\n\n<Subject 2> must remain clearly visible behind <Subject 1> throughout the performance.\n\n<Subject 1> performs the final restrained Italo-disco phrase: two side steps, one metallic right-hand finger point, one controlled shoulder roll, a small half-turn, one smooth backward glide, then both metallic hands briefly rise symmetrically at chest height.\n\nKeep the choreography simple, physically believable and identity-preserving.\n\nCyan light reflects across the pale blue suit and chrome robotic hands while magenta remains only a secondary accent.\n\nAfter the short dance, <Subject 1> stops, looks over the shoulder toward <Subject 2>, turns and calmly walks back to the open vehicle door.\n\n<Subject 1> begins lowering into the seat.\n\nEND STATE / TRANSITION: <Subject 1> is halfway seated inside <Subject 2>, one black shoe still on the wet pavement, door open, cyan canopy overhead. Next scene begins from this exact physical pose.\n\nNo additional dancers, no crowd, no costume change, no vehicle variation, no text.\n\nsubject_definitions:\n<Subject 1> is the person in <Picture 1>, preserving their exact identity, facial features, skin tone, hairstyle, facial proportions, body proportions, pale blue uniform suit, white shirt, black tie, matching round pale blue cap, metallic robotic hands, black footwear, and distinctive cheerful uncanny expression.\n<Subject 2> is the autonomous vehicle in <Picture 2>, preserving its exact recognizable vehicle identity, white exterior, compact proportions, rounded body shape, dark windows, wheels, lighting, roof-mounted autonomous sensor equipment and overall silhouette. Ignore all daylight environmental information from <Picture 2>.\n\nsummary:\n[reference generation] Complete the video with <Subject 1> entering <Subject 2>, giving one small final gesture through the window, and the exact vehicle driving away through the familiar neon boulevard during one continuous eleven-second closing shot.\n\nretention_analysis:\n<Subject 1> (appears during first half of [Shot 1]): fully_preserved - exact face, identity, clothing, cap, tie, metallic hands, footwear and proportions remain unchanged during the final seating and farewell gesture.\n<Subject 2> (appears throughout [Shot 1]): fully_preserved - exact white autonomous vehicle, body design, windows, wheels, sensor equipment, lights and proportions remain stable until it disappears naturally into the city.\n\ndetailed_description:\nThe target is an eleven-second photorealistic cinematic final single take in the exact same cyberpunk downtown district on the same night. Maintain the established 1980s science-fiction film aesthetic: practical cyan neon, restrained magenta accents, wet black boulevard, brutalist concrete buildings, chrome reflections, drifting steam, deep blue-black shadows, soft optical diffusion, anamorphic horizontal flares and subtle 35mm grain.\n\n[Shot 1] Begin exactly with <Subject 1> halfway seated inside <Subject 2> beneath the familiar cyan canopy, with one black shoe still outside.\n\n<Subject 1> smoothly brings the remaining leg and metallic hands into the cabin, settles into the seat and closes the vehicle door.\n\nThrough the side window, <Subject 1> turns toward camera one final time and performs a tiny understated farewell: two metallic fingers rise briefly in a restrained disco gesture.\n\n<Subject 2> gently begins moving away from the curb.\n\nCamera remains stationary at street level at first, watching the exact vehicle move deeper down the same familiar wet boulevard.\n\nAfter several seconds, camera begins a slow cinematic crane upward, revealing the same cyan-lit brutalist architecture already established throughout the video. Do not introduce any new landmark or district.\n\n<Subject 2> becomes progressively smaller while its white body and roof-mounted sensors remain recognizable under the neon light.\n\nCyan reflections stretch along the wet road behind it.\n\nDuring the final seconds, <Subject 2> reaches the same distant corner previously seen in the video and turns gently behind a building.\n\nThe vehicle disappears naturally from sight.\n\nHold very briefly on the empty wet boulevard, cyan neon reflecting across the pavement and a small cloud of steam drifting through frame.\n\nSlow cinematic fade to black.\n\nNo new subjects, no new vehicles, no location change, no transformation, no text, no subtitles, no generated logos.\n\noverall_soundscape:\nNone required. Visual generation only.\n\nnon_diegetic_music:\nNone generated. Final song continues underneath during editing and fades with the image.</code></pre>\n<p>example 9</p>\n<p><a href=\"https://www.reddit.com/r/comfyui/comments/1vj3t9h/almost_as_bad_as_the_star_wars_holiday_special/\">https://www.reddit.com/r/comfyui/comments/1vj3t9h/almost_as_bad_as_the_star_wars_holiday_special/</a></p>\n<pre class=\"language-apacheconf\"><code> Standard live-action The Big Bang Theory sitcom look: practical television photography style, a sitcom apartment set, basic lens, average depth of field, tv quality recording look, tv studio lighting, standard living room props, sit and stand around acting, with very little walking.\n\nScene overview: the apartment set from The Big Bang Theory <Picture 2>, the protagonist Sheldon Cooper <Picture 0> sitting on the couch, Luke Skywalker <Picture 1> is sitting left of Sheldon, Sheldon Cooper <Picture 0> complains to Luke <Picture 1>\n\nStoryboard: (each shot is a wide and medium shots of the same set, cuts only when a new character is shown):\n\n[0s-10s] Shot 1: medium shot of Sheldon <Picture 0> and Luke Skywalker <Picture 1> on the couch: Sheldon <Picture 0> is complaining to Luke <Picture 1>. Sheldon Cooper: \"I don't know what to tell you Mark, the writing, th-the acting, it was just appauling. It was almost as bad as the Star Wars holiday special.\" Followed by a laugh track.\n\nCamera: each shot its focused on the stage and actors, always facing the set like it would on any sitcom.\n\nAudio: Tv studio quality. Straight from the Seinfeld tv show.\n\nNo text, subtitles, logos or watermarks of any kind, no animation or cartoon rendering, no overly-CG look, keep the classic 90s sitcom live-action texture. </code></pre>\n<p>example 10 - CHARACTER SWAP V2V</p>\n<p><a href=\"https://app.notion.com/p/CHARACTER-SWAP-V2V-PROMPTS-halo-christo-3b663c2a7548802a8143d3ac8cb23373\">https://app.notion.com/p/CHARACTER-SWAP-V2V-PROMPTS-halo-christo-3b663c2a7548802a8143d3ac8cb23373</a></p>\n<p><a href=\"https://www.reddit.com/r/comfyui/comments/1vinc36/testing_character_swap_with_minimax_h3/\">https://www.reddit.com/r/comfyui/comments/1vinc36/testing_character_swap_with_minimax_h3/</a></p>\n<pre class=\"language-json\"><code>SCENE CONTEXT\nObject replacement pass. In <Video_1>, the target object is replaced by the object shown in <Image_1>. Everything else in <Video_1> remains exactly as it is.\n\nACTIVE REFERENCES\n<Video_1>: the master plate. Camera path, framing, timing, cast, environment, lighting and every other object 100% match <Video_1>.\n<Image_1>: identity of the replacement object only. Its shape, proportion, material, colour, logos and surface markings 100% match <Image_1>, kept legible and correctly oriented throughout. NO MASK\n\nMOTION INHERITANCE\nThe replacement object inherits the full behaviour of the object it replaces, frame by frame: same screen position, same scale, same rotation, same motion path, same speed, same entry and exit timing. Whatever the original object did, the new object does identically. No new movement is introduced and none is removed.\n\nINTEGRATION\nContact reads physically: hands wrap the new silhouette, supporting surfaces meet its actual base, contact shadows land directly beneath it, and any grip conforms to its real geometry.\nOcclusion order is preserved: whatever passed in front of the original object passes in front of the new one, and whatever it covered stays covered.\nReflections, refractions and cast shadows on nearby surfaces are rebuilt for the new geometry while keeping the same direction and softness as the plate.\n\nOPTICS\nShot size, FOV, depth of field, focus falloff and motion blur carried over from <Video_1> with no drift. The object sits at the same focal plane as the original.\n\nCAMERA\nCamera behaviour, height, distance, movement and handheld character identical to <Video_1>.\n\nPHYSICS\nMass, inertia, swing and settle behaviour consistent with the material shown in <Image_1>. Any fluid, spill, dust or particle interaction updates to the new geometry while obeying the same gravity and timing as the plate.\n\nLIGHTING\nKey direction, intensity, falloff and white balance taken from <Video_1>. The object catches the same key from the same side, sits at the same ambient level, and throws a shadow matching the existing shadows in length, direction and softness. Specular highlights appear only where the plate's key light would place them, reading the true surface finish from <Image_1>.\n\nSTYLE\nPhotoreal, fully integrated into the original plate: same grain structure, same black level, same tonal contrast, same colour grade as <Video_1>.\n\nPOSITIVE LOCKS\n\n- Only the target object changes; every other element of <Video_1> stays untouched.\n- The object stays present, complete and correctly scaled in every frame the original appeared in.\n- Identity from <Image_1> holds steady across the whole clip, with no drift in shape, colour or markings.\n- Edges blend seamlessly: matching noise, matching edge softness, no halo, no outline.\n- One continuous plate, cuts only where <Video_1> already cuts.</code></pre>\n<p>example 11 - It’s made from 12 separate clips, 15 seconds each, so 3 minutes total. Not audio driven. </p>\n<p><a href=\"https://www.reddit.com/r/comfyui/comments/1vi2ole/testing_of_3minute_ai_music_video_rtx_3090_with/\">https://www.reddit.com/r/comfyui/comments/1vi2ole/testing_of_3minute_ai_music_video_rtx_3090_with/</a></p>\n<pre class=\"language-json\"><code>A surreal psychedelic cinematic music-video shot. A single glossy electric-blue mouth floats in total darkness, emerging slowly from black as if waking up. Highly detailed wet luminous lips with ultraviolet highlights and liquid-cyan reflections.\n\nThe mouth starts almost still and then moves as if singing to an unheard electronic song. Simulate musical rhythm visually: at regular intervals the lips pulse open and closed and each pulse emits a circular cyan soundwave traveling outward through the darkness. Tiny glowing particles shake and react as the waves pass.\n\nSmall abstract wireless and Bluetooth-like symbols occasionally flicker around the mouth and dissolve into blue vapor.\n\nVery slow cinematic macro push toward the lips, shallow depth of field, halation, bloom, subtle analog-video texture.\n\nDuring the final seconds the mouth opens wider, revealing an intense tunnel of blue light inside. The camera accelerates forward and enters the glowing mouth.\n\nEnd while the camera is traveling into the mouth so the next scene can continue from inside it.\n\nContinue directly from inside the glowing electric-blue mouth. The camera flies forward through a surreal organic tunnel made from glossy lip textures, translucent crystal teeth, liquid membranes and rippling cyan light.\n\nSimulate an electronic musical rhythm visually. Bright pulses travel down the tunnel at repeating intervals. Concentric soundwave rings, spirals and glowing frequency ribbons race along the walls. Tiny wireless symbols drift through the space like bioluminescent insects.\n\nThe tunnel stretches, contracts and breathes rhythmically.\n\nThe camera moves rapidly forward with smooth banking turns and gentle rolls, maintaining constant momentum.\n\nNear the end the tunnel expands into a vast black void. A gigantic electric-blue mouth floats ahead.\n\nAs the camera approaches, that single mouth begins dividing into many identical mouths.\n\nEnd at the moment the duplication begins.\n\nBegin with the giant blue mouth from the previous scene dividing into dozens of identical glossy electric-blue mouths floating in deep black space like a constellation.\n\nEach mouth behaves like a different musical instrument. Some pulse slowly like bass, others open and close rapidly. Each emits a different visible waveform: circular cyan rings, spirals, jagged crystalline waves, vibrating ribbons and translucent frequency walls.\n\nWhere the waves intersect they briefly create luminous geometric flowers, abstract faces and interference patterns.\n\nThe camera moves continuously through the constellation in zero gravity, weaving between mouths and passing extremely close to some of them before pulling rapidly into wide views.\n\nToward the end all mouths suddenly rotate toward the same distant point.\n\nThey open together and release one enormous synchronized beam of blue frequency light.\n\nThe camera accelerates along the beam.\n\nEnd as the beam reaches a distant humanoid silhouette.\n\nBegin with the cyan frequency beam from the previous scene reaching an androgynous humanoid figure standing inside a dark surreal nightclub.\n\nThe room contains reflective black floors, giant mirrors, ultraviolet fog and isolated cyan lights.\n\nThe figure initially has a perfectly smooth face with no mouth.\n\nThe incoming beam hits the face and a glossy electric-blue mouth forms there from liquid neon.\n\nImmediately the new mouth begins moving rhythmically as if singing.\n\nEvery implied beat produces a powerful visible bass wave expanding through the room. Curtains ripple, mirrors bend elastically, reflections distort and pools of light pulse across the floor.\n\nCamera continuously orbits around the figure, alternating between medium shots and extreme close-ups of the mouth.\n\nThe movement becomes increasingly energetic.\n\nDuring the final seconds a huge soundwave hits a mirror behind the figure. The mirror liquefies into chrome water.\n\nThe camera follows the wave directly through the liquid mirror.\n\nEnd while crossing through it.\n\nContinue through the liquid mirror into an inverted psychedelic chrome-and-blue world.\n\nTwo giant glossy electric-blue mouths float facing one another from opposite sides of the frame.\n\nThey move rapidly toward each other, trailing long ribbons of cyan waveform light.\n\nThey stop only millimeters apart without touching.\n\nBetween them the air vibrates violently with visible frequency patterns.\n\nThe mouths pulse and move as if harmonizing to an unheard electronic song. Every implied beat compresses the light between them further.\n\nA brilliant sphere of cyan energy gradually forms between the lips, made from waveform lines, liquid light and tiny wireless symbols.\n\nThe camera performs a fast continuous orbit around the mouths while moving closer.\n\nAt the climax the sphere bursts outward into thousands of miniature glowing blue mouths and fragments.\n\nThe camera immediately chooses one tiny falling mouth and dives downward after it.\n\nEnd while rapidly following the falling mouth.\n\nFollow the miniature blue mouth falling directly out of the previous scene.\n\nIt drops through darkness and suddenly enters a surreal nighttime city.\n\nThousands of tiny glowing electric-blue mouths are falling from the sky like rain.\n\nThe camera descends rapidly toward street level and begins flying forward between buildings.\n\nEach mouth opens and closes rhythmically while falling. Whenever one strikes a window, car, rooftop or puddle it generates a glowing circular soundwave.\n\nStreetlights flash rhythmically. Building reflections ripple. Neon signs distort. Puddles produce synchronized frequency patterns.\n\nThe camera races low over the wet street, occasionally passing through clouds of falling mouths and narrowly avoiding chrome objects.\n\nAs the sequence progresses, the pavement begins turning into electric-blue liquid.\n\nBuildings stretch downward into their own reflections.\n\nThe entire city melts into one enormous blue ocean.\n\nEnd with the camera racing just above the newly formed liquid surface.\n\nBegin directly above the infinite electric-blue liquid ocean created from the melting city.\n\nThe camera flies extremely low over the surface at high speed.\n\nHuge glossy blue mouths rise explosively from beneath the water like surreal islands.\n\nEach mouth opens on an implied beat and sends enormous circular waves across the ocean.\n\nThe camera banks around the waves, dives between giant lips and skims across the liquid surface while glowing droplets explode upward.\n\nA chrome humanoid figure suddenly appears running across the water.\n\nEvery footstep generates luminous frequency rings.\n\nThe camera follows beside and behind the figure with aggressive tracking movement.\n\nAhead, an enormous mouth rises from the ocean and opens.\n\nThe water begins rotating into a giant whirlpool shaped like an abstract Bluetooth symbol.\n\nThe chrome figure accelerates toward it and is pulled into the vortex.\n\nThe camera dives in immediately behind the figure.\n\nEnd while falling into the whirlpool.\n\nContinue falling through the blue whirlpool.\n\nThe chrome figure suddenly splits into two mirrored humanoid bodies suspended in an enormous electric-blue void.\n\nEach body has a glowing blue mouth embedded in the chest where the heart would be.\n\nThe bodies orbit each other rapidly.\n\nTheir chest-mouths repeatedly open and fire thin cyan waveform lines toward one another.\n\nThe connection fails several times, causing sharp visual glitches, duplicate frame echoes and bursts of distorted space.\n\nSimulate rhythmic music visually through repeated connection attempts, pulsing light and synchronized body movement.\n\nThe camera revolves around both figures while constantly changing distance, moving from extreme close-ups of the heart-mouths to wide shots of the orbiting bodies.\n\nFinally one bright continuous waveform successfully connects both mouths.\n\nTheir movements instantly synchronize.\n\nThe cyan connection becomes brighter and thicker until it fills the center of the frame.\n\nThe camera accelerates directly into the connecting waveform.\n\nEnd inside the bright line.\n\nThe bright connecting waveform from the previous scene expands around the camera and transforms into an extreme macro view of two electric-blue mouths approaching one another.\n\nThe lips are constructed from glossy liquid, pixels, tiny waveform fragments and floating droplets.\n\nThe camera moves rapidly around and between them while maintaining extreme macro detail.\n\nAs the mouths approach, streams of luminous data begin moving between them.\n\nTiny landscapes, abstract memories, faces, wireless symbols and pulses of cyan light flow from one mouth into the other.\n\nThe mouths never physically collide. Instead they exchange increasingly intense streams of information through the narrow gap between them.\n\nEach implied musical beat causes the lips to pulse, the data stream to surge and the camera to change speed.\n\nThe two mouths gradually lose their solid form and dissolve into one enormous shared waveform.\n\nThe waveform twists violently into a spiral.\n\nAs the camera follows it, the spiral begins forming the petals of a giant electric-blue flower.\n\nEnd as the flower starts opening.\n\nContinue from the waveform flower opening in the previous scene.\n\nReveal an enormous psychedelic flower floating in black space, constructed entirely from hundreds of glossy electric-blue mouths.\n\nEvery petal is a mouth.\n\nThe flower rotates rapidly while different rings of mouths open sequentially, creating visible waves traveling around its circumference.\n\nEach pulse generates cyan frequency ribbons, glowing pollen and expanding geometric mandalas.\n\nThe camera spirals aggressively through the flower, diving between petals, rotating around the center and constantly shifting between huge wide shots and extreme macro mouth close-ups.\n\nThe rotational speed steadily increases.\n\nThe mouth petals stretch into liquid ribbons and the entire flower transforms into an enormous rotating blue mandala.\n\nAt the climax every mouth closes simultaneously.\n\nThe entire structure collapses rapidly inward.\n\nEverything disappears except one single isolated electric-blue mouth floating in darkness.\n\nThe camera brakes suddenly and stops close to it.\n\nEnd on the lonely mouth.\n\nBegin with the isolated electric-blue mouth floating alone in nearly total darkness.\n\nThe atmosphere is quieter but still constantly moving.\n\nThe camera slowly circles the mouth while drifting closer and farther away.\n\nThe mouth attempts to sing.\n\nEach time it opens, a visible cyan waveform travels outward but quickly breaks apart into digital particles.\n\nThe mouth tries repeatedly.\n\nEvery failed signal creates analog glitches, transparent duplicates and brief distortions of the surrounding darkness.\n\nThe pulses become weaker.\n\nThen a tiny blue signal appears extremely far away.\n\nThe camera suddenly changes direction and moves toward the distant light while keeping the mouth visible behind.\n\nThe isolated mouth reacts and fires a stronger waveform.\n\nThe distant signal answers with another wave.\n\nBoth waves race toward each other through the void.\n\nThe camera accelerates alongside them.\n\nEnd one instant before the two signals collide.\n\nBegin exactly before the two cyan signals collide.\n\nThey connect and immediately produce an enormous radiant blue shockwave.\n\nThe camera is thrown backward at extreme speed.\n\nThe entire psychedelic universe from the previous scenes suddenly appears around the camera: hundreds of electric-blue mouths, liquid oceans, chrome bodies, waveform flowers, falling mouth rain, rotating wireless symbols and glowing frequency ribbons.\n\nEverything moves in synchronized rhythmic pulses.\n\nThe camera races through the environment, passing between giant mouths, diving through waveform rings, rolling around chrome bodies and skimming above liquid-blue surfaces.\n\nEvery implied beat triggers another transformation.\n\nThe movement reaches maximum intensity.\n\nThen the camera begins an extremely fast continuous pull backward.\n\nAs it moves away, the entire impossible universe becomes smaller.\n\nEventually reveal that everything we have been seeing exists only as a reflection on the surface of one enormous pair of glossy electric-blue lips floating in black space.\n\nThe camera continues pulling back.\n\nThe giant lips slowly close.\n\nOne final cyan soundwave explodes outward toward the camera, expanding until it fills the entire frame.\n\nThe light fades rapidly to pure black.\n\nHold black for the final second.</code></pre>\n<p>example 12</p>\n<p><a href=\"https://www.reddit.com/r/comfyui/comments/1vgze70/minimax_h3_image_to_video_my_example/\">https://www.reddit.com/r/comfyui/comments/1vgze70/minimax_h3_image_to_video_my_example/</a></p>\n<pre class=\"language-json\"><code>Warner Bros cartoon, Wile E. Coyote run back of Beep Beep. Beep beep, running along and doesn't realize where he's going. So Beep Beep ends up crashing into a tree. Then the scene change with Wile E. Coyote is cooking on the barbecue with a single cooked Beep Beep on. He puts it in his mouth, chews it, but makes a disgusted and then spits it out to the side. Then scene change with Wile E. Coyote seen from back is entering in a Mc Donald's restaurant.</code></pre>\n<p>example 13</p>\n<p><a href=\"https://www.reddit.com/r/StableDiffusion/comments/1vgqvtu/first_video_attempt_yeah_minmax/\">https://www.reddit.com/r/StableDiffusion/comments/1vgqvtu/first_video_attempt_yeah_minmax/</a></p>\n<pre class=\"language-json\"><code>\n The dwarf from <Picture 1> stands in his original environment. He warms up with a broad, friendly smile, looks directly at the camera, and raises his hand in greeting. He speaks in a friendly but gruff Scottish accent. He finishes speaking, raises his foam topped tankard in greeting, it sloshes around sloppily and drips down the side of the mug. He gives a playful wink to the camera, takes a large drink of the beer, foam soaking into his facial hair realistically. He lowers the mug and lets out a content sigh as holds his smile as the video ends.\n\n Timeline:\n\n [0s-1s] The dwarf brightens up, smiles warmly, and raises his hand in greeting.\n\n [2s-10s] He speaks directly to the audience in a gruff Scottish accent: \"Alright, So I've been messing around with that A.I. stuff again. McCoy, you seeing this shit? I was a dragon age screenshot once\"\n\n [10s-12s] He gives a quick, cheerful wink to the camera and raises his sloshing foam topped tankard as if in a minor toast.\n\n [12s-15s] He takes a large drink of the beer, the foam soaking into to his facial hair where appropriate.\n\n [15s-18s] He holds his warm smile as the scene smoothly settles to a close.\n\n Audio: Clear spoken male voice with a distinct Scottish accent, soft movement rustle during the wave, and subtle ambient room tone.\n</code></pre>\n<p>example 14 - Prompt For Multi Character</p>\n<p>Used multiple reference images for each scene.</p>\n<p><a href=\"https://www.reddit.com/r/comfyui/comments/1vgf9ai/assemble_the_multiverse_minimax_h3_r2v_is_awesome/\">https://www.reddit.com/r/comfyui/comments/1vgf9ai/assemble_the_multiverse_minimax_h3_r2v_is_awesome/</a></p>\n<pre class=\"language-json\"><code> \nsubject_definitions:\n\n<Subject 1> is [CHARACTER 1] from <Picture 1>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.\n\n<Subject 2> is [CHARACTER 2] from <Picture 2>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.\n\n<Subject 3> is [CHARACTER 3] from <Picture 3>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.\n\nsummary:\n\n[reference generation] A 5-second cinematic multiverse portal arrival. Three characters emerge from a consistent amber-orange portal and take a calm, confident formation.\n\nretention_analysis:\n\n<Subject 1> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 1>.\n\n<Subject 2> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 2>.\n\n<Subject 3> (appears in [Shot 1]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 3>.\n\ndetailed_description:\n\nA 5-second cinematic portal-arrival scene at dusk. One stable medium three-shot, framed from the knees up. No dialogue, no combat, no wide landscape, no camera movement, and no crowd.\n\nPortal continuity: a large circular amber-orange portal stands behind the characters. It has a bright rotating outer ring, a darker transparent centre, floating gold sparks, gentle smoke, and a low resonant hum. Keep its colour, size, brightness, particle density, and rotation speed consistent across every portal-arrival clip.\n\n[Shot 1] <Subject 1> steps through the portal first and takes the centre position with quiet confidence. <Subject 2> emerges on one side, naturally adjusts or lowers any item they are carrying if applicable, then gives a focused glance toward the unseen distance. <Subject 3> walks through last, takes position on the opposite side, and calmly surveys the scene. The three hold a poised, united stance as the portal flickers and golden particles drift around them. Their expressions and body language remain confident and appropriate to their individual character identities.\n\noverall_soundscape:\n\nLow portal hum, soft wind, subtle movement from clothing or equipment, and drifting golden sparks. No speech.\n\nnon_diegetic_music:\n\nA restrained cinematic rise builds across the shot and resolves on a calm, confident note.\n\nPrompt For single characters:\nsubject_definitions:\n\n<Subject 1> is [CHARACTER] from <Picture 1>, preserving the visual identity, costume, proportions, accessories, and visual style shown in the reference image.\n\nsummary:\n\n[reference generation] A 5-second cinematic multiverse portal arrival. One character walks through a consistent amber-orange portal, then takes a confident action stance with a subtle grin.\n\nretention_analysis:\n\n<Subject 1> (appears in [Shot 1] and [Shot 2]): fully_preserved - retains the identity, costume, proportions, accessories, and visual style shown in <Picture 1>.\n\ndetailed_description:\n\nA 5-second cinematic portal-arrival scene at dusk. No dialogue, no crowd, no wide landscape, and no combat.\n\nPortal continuity: a large circular amber-orange portal stands behind <Subject 1>. It has a bright fiery rotating outer ring, a darker transparent centre, floating gold sparks, gentle smoke, and a low resonant hum. Keep its colour, size, brightness, particle density, and rotation speed consistent across every portal-arrival clip.\n\n[Shot 1] Medium knee-up shot. <Subject 1> walks steadily through the portal toward the camera, then comes to a composed stop. Their costume, silhouette, movement style, and any character-specific accessories remain fully consistent with <Picture 1>. Golden sparks drift around them as the portal flickers behind.\n\n[Shot 2] Close-up of <Subject 1>. They shift into a distinctive, character-appropriate action stance, looking directly ahead with calm confidence. Their expression changes into a subtle smile and restrained grin. Keep the movement natural and controlled, with no exaggerated facial distortion. The portal remains softly visible and out of focus in the background.\n\noverall_soundscape:\n\nLow portal hum, soft wind, subtle movement from clothing or equipment, and drifting golden sparks. No speech.\n\nnon_diegetic_music:\n\nA restrained cinematic rise builds through the entrance and resolves as <Subject 1> holds the final stance.</code></pre>\n<p>example 15 - using a small black image as first frame</p>\n<p>tip - writing a prompt and inputting an image as the last frame. However, the default workflow seemed to always require a first frame. So, I used a small black image as the first frame; this time, it worked.</p>\n<p><a href=\"https://www.reddit.com/r/StableDiffusion/comments/1vl3uu4/minimax_h3_testing_l2va_moon_landing/\">https://www.reddit.com/r/StableDiffusion/comments/1vl3uu4/minimax_h3_testing_l2va_moon_landing/</a></p>\n<pre class=\"language-json\"><code> integrated_multimodal_description: Time-lapsed, cinematic, a medium-wide shot of a film set in a large indoor studio which is used to shoot a scene of moon landing involving lunar module, US flag and an astronaut. At 00:00.000 the camera shows an empty, sterile white studio room, with recognizable vertical wand in the back and horizontal floor at its bottom. At 00:01.000 Some film crews install a black wand into the studio's vertical wand. The black wand has some tiny white shining points which represent stars. At 00:02.000 Some workers fill the studio's floor with some dirty-white sand, gravels and small rocks and form a barren lunar landscape. At 00:03.000 Some film crews bring an Apollo Lunar Module and place it into the left side of the scene. At 00:04.000 A film crew places a US flag with pole on the right side of the scene. Another film crew puts a picture of the Earth as the blue planet, partially blacked on its bottom side, on the top right corner of the scene. At 00:05:000 An astronaut walks in into the scene, goes into the middle of the scene, faces to viewer and waves his hand. At 00:07:000 The whole scene settles into the exact arrangement, position of subject and objects, camera angle, lighting, and final composition established by <Picture 1>. A male deep voice of the director (S1) says, <d>[English] Cut!</d>\n\noverall_soundscape:\n\nnon_diegetic_music: Sustained violin notes at a very fast tempo with spaced piano tones. </code></pre>\n<p>example 16</p>\n<p><a href=\"https://www.reddit.com/r/StableDiffusion/comments/1vhuorq/45_lower_minimax_h3_sampler_time_with_new/\">https://www.reddit.com/r/StableDiffusion/comments/1vhuorq/45_lower_minimax_h3_sampler_time_with_new/</a></p>\n<pre class=\"language-json\"><code> Live-action photorealistic military science-fiction action film, 15 seconds. A modern main battle tank fights a towering Gundam-style humanoid combat mech inside a devastated urban warzone at sunset. Collapsed concrete towers, burning vehicles, broken highways, drifting smoke, sparks, dust, tracer fire, and debris fill the battlefield. The tank feels extremely heavy and grounded; the giant mech moves with terrifying mechanical speed and weight. Every movement causes believable environmental reactions. Maximum action cinematography: aggressive low angles, extreme close-ups, tracking shots, whip pans, rapid push-ins, rotating Arc Shots, strong impact shake, and dramatic changes in camera height. Keep the combat visually readable despite the extreme camera movement.\n\n[Shot 1]\n\nAn extreme low-angle medium tracking shot races inches above the broken road beside the tank as it charges forward at full speed, its tracks crushing concrete and throwing chunks of asphalt directly past the lens. The turret rotates upward while the enormous Gundam-style mech (S1) lands in the street ahead, one knee smashing into the pavement.\n\nThe impact sends a circular blast of dust and debris outward.\n\nThe camera violently tilts up from the tank to reveal the full towering mech.\n\nThe tank immediately fires its main cannon.\n\nA massive muzzle flash fills the frame.\n\n[Shot 2] At 00:03.000, the camera cuts to an extreme close-up beside the mech's head as the tank shell screams directly toward the camera.\n\nAt the final instant, (S1) violently twists its torso and leans sideways.\n\nThe shell misses its head by centimeters and explodes against a skyscraper behind it.\n\nWithout pausing, (S1) plants one mechanical foot into the street and launches forward.\n\nThe camera rapidly pulls backward at low height while the giant mech sprints directly toward the tank, every footstep smashing craters into the road and throwing abandoned cars sideways.\n\n[Shot 3] At 00:06.000, the camera cuts to a dramatic POV from immediately above the tank turret looking upward.\n\n(S1) suddenly JUMPS.\n\nThe camera performs an extremely fast Tilt Up as the enormous mech passes directly overhead, silhouetted against the sky.\n\nWhile airborne, (S1) rotates its hips, extends one leg, and comes down with a massive flying kick.\n\nThe camera whip-pans downward.\n\nThe tank driver violently turns.\n\nThe tank power-slides sideways across the road.\n\n(S1)'s foot misses the tank by inches and SMASHES into the pavement beside it.\n\nThe entire street erupts upward.\n\nThe camera shakes strongly as concrete, dust, and wreckage explode past the lens.\n\n[Shot 4] At 00:09.500, the camera cuts to a tight side tracking shot moving alongside the tank as it emerges through the dust cloud.\n\nThe turret rapidly rotates backward toward (S1).\n\nThe cannon fires point-blank.\n\nThe camera instantly whip-pans with the shell.\n\n(S1) raises a massive armored forearm across its chest.\n\nThe shell EXPLODES against the armor, forcing the mech backward through a building facade in an enormous cloud of concrete and glass.\n\nBefore the dust settles, two glowing mechanical eyes ignite inside the smoke.\n\n(S1) bursts forward.\n\n[Shot 5] At 00:12.000, the camera cuts to an extreme low-angle shot directly beside the tank's tracks.\n\n(S1)'s enormous hand suddenly SLAMS onto the tank's turret.\n\nThe camera rapidly arcs upward around both machines as the mech uses its entire body to lift the tank off the ground.\n\nThe tank's tracks continue spinning wildly in midair.\n\nThe camera accelerates into a huge 180-degree Arc Shot as (S1) rotates its torso and violently THROWS the entire tank across the battlefield.\n\nThe tank spins through the air directly past the camera.\n\nThe camera whip-pans after it.\n\nAt 00:14.500, the tank crashes sideways through a concrete wall in a gigantic explosion of dust and debris.\n\nEnd on an extreme low-angle close-up of (S1) stepping through the smoke toward the camera as burning debris rains behind it. </code></pre>\n<p> </p>\n<p>reddit posts with lots of general info</p>\n<p><a href=\"https://www.reddit.com/r/comfyui/comments/1vinc36/testing_character_swap_with_minimax_h3/\">https://www.reddit.com/r/comfyui/comments/1vinc36/testing_character_swap_with_minimax_h3/</a></p>\n<p><a href=\"https://www.reddit.com/r/StableDiffusion/comments/1vhuorq/45_lower_minimax_h3_sampler_time_with_new/\">https://www.reddit.com/r/StableDiffusion/comments/1vhuorq/45_lower_minimax_h3_sampler_time_with_new/</a></p>\n<p><a href=\"https://www.reddit.com/r/comfyui/comments/1vknr0v/followup_from_a_5second_clip_to_a_247/\">https://www.reddit.com/r/comfyui/comments/1vknr0v/followup_from_a_5second_clip_to_a_247/</a></p>\n<p><a href=\"https://www.reddit.com/r/NeuralCinema/comments/1vk20gc/minimax_h3_hack_50_reference_or_more/\">https://www.reddit.com/r/NeuralCinema/comments/1vk20gc/minimax_h3_hack_50_reference_or_more/</a></p>\n<p><a href=\"https://www.reddit.com/r/comfyui/comments/1vjs1re/is_there_a_nice_place_to_learn_how_to_use_minimax/\">https://www.reddit.com/r/comfyui/comments/1vjs1re/is_there_a_nice_place_to_learn_how_to_use_minimax/</a></p>\n<p><a href=\"https://www.reddit.com/r/comfyui/comments/1vjn8i9/minimaxh3_speedup_nodes_is_this_connected/\">https://www.reddit.com/r/comfyui/comments/1vjn8i9/minimaxh3_speedup_nodes_is_this_connected/</a></p>\n<p><a href=\"https://fal.ai/learn/devs/minimax-h3-prompting-guide\">https://fal.ai/learn/devs/minimax-h3-prompting-guide</a></p>\n<p><a href=\"https://www.reddit.com/r/comfyui/comments/1vj8nyp/forcing_minimax_h3_to_generate_multispeaker/\">https://www.reddit.com/r/comfyui/comments/1vj8nyp/forcing_minimax_h3_to_generate_multispeaker/</a></p>\n<p><a href=\"https://www.reddit.com/r/StableDiffusion/comments/1vfyiey/universal_minimax_h3_video_prompt_architect_v3/\">https://www.reddit.com/r/StableDiffusion/comments/1vfyiey/universal_minimax_h3_video_prompt_architect_v3/</a></p>\n<p><a href=\"https://www.reddit.com/r/StableDiffusion/comments/1vh2p4z/minimax_reference_method_try_this_setting_instead/\">https://www.reddit.com/r/StableDiffusion/comments/1vh2p4z/minimax_reference_method_try_this_setting_instead/</a></p>\n<p><a href=\"https://www.reddit.com/r/StableDiffusion/comments/1vhuorq/45_lower_minimax_h3_sampler_time_with_new/\">https://www.reddit.com/r/StableDiffusion/comments/1vhuorq/45_lower_minimax_h3_sampler_time_with_new/</a></p>\n<p>tricks and tips</p>\n<p>about audio::: </p>\n<p><a href=\"https://www.reddit.com/r/comfyui/comments/1vjbyqk/minimax_is_nutso/\">https://www.reddit.com/r/comfyui/comments/1vjbyqk/minimax_is_nutso/</a></p>\n<pre class=\"language-json\"><code> I’ve had better luck not just saying “no voices,” but explicitly building the prompt around non-vocal audio.\n\nSomething like:\n\n“No dialogue, no voiceover, no singing, no speaking, and no vocalizing characters. The scene is entirely atmospheric.”\n\nThen list the sounds you actually want, like wind, footsteps, traffic, rain, doors, clothing rustle, etc.\n\nFor example:\n\n“overall_soundscape: Soft night ambience, distant traffic, light wind, footsteps on pavement, car unlock chirp, door handle click, and the muted thump of the car door closing.”\n\nAnd if you don’t want background music either: “non_diegetic_music: N/A”\n\nAlso don’t give the character an (S1) speaker ID if they never talk. That seems to help keep the model from inventing speech. </code></pre>\n<p>experimental --- Image Editing Tasks and Prompts</p>\n<pre class=\"language-json\"><code># Image Editing Tasks and Prompts\n \n## 1. Depth pose\n \n### Task\n \nRe-create the referenced person while applying a body pose supplied by a depth diagram.\n \n### Exact prompt\n \n```text\nTask: Reference-guided generation. Generate a new portrait image of the person from <Picture 1>. <Picture 1> is the character appearance reference. Preserve the person's identity, facial features, skin tone and skin texture, hair color, hairstyle and hair length, complete outfit, garment colors, fabrics, patterns, footwear, and accessories. Preserve the background, environment, lighting, and camera framing from <Picture 1>. <Picture 2> does not supply appearance. Copy the face, skin, hair, outfit, and background from <Picture 1> exactly; do not substitute, simplify, restyle, or redesign them. <Picture 2> is the pose reference. Use the body pose, limb positions, torso orientation, head orientation, and body geometry shown in the depth diagram from <Picture 2>. Generate the person from <Picture 1> naturally occupying the pose specified by <Picture 2>. A single coherent 2:3 portrait photograph shows the complete person.\n```\n \n## 2. Outfit, location, and action\n \n### Task\n \nCombine a character, an outfit, and a location from separate references while directing a specific action.\n \n### Exact prompt\n \n```text\nTask: Reference-guided generation. The adult character from <Picture 1> wears the complete practical outfit from <Picture 2> and walks briskly across the location from <Picture 3>, carrying one paper grocery bag against the torso while glancing toward the camera. <Picture 2> supplies wardrobe only and <Picture 3> supplies the entire background and setting. Replace the original background from <Picture 1> completely with the location from <Picture 3>; retain no scenery or location elements from <Picture 1>. The face, hairstyle, and identity remain recognizably source-authentic. Natural full-body 2:3 portrait photography, grounded feet, coherent perspective and lighting integrated into the new background.\n```\n \n## 3. Character sheet\n \n### Task\n \nCreate a consistent three-view photographic character sheet from one character reference.\n \n### Exact prompt\n \n```text\nTask: Reference-guided generation. A clean photographic character sheet presents the same adult from <Picture 1> in three consistent full-body views: front, side, and rear. Every view has the identical face, hairstyle, body, outfit, and accessories from the source. Neutral relaxed stance, arms clear of the torso, both feet visible in each view, seamless light-gray studio, even lighting, one tall 2:3 portrait canvas, no captions or borders.\n```\n \n## 4. Age 60\n \n### Task\n \nAge the referenced person to exactly 60 years old while preserving identity and composition.\n \n### Exact prompt\n \n```text\nTask: Reference-guided generation. Render the adult from <Picture 1> at exactly 60 years old while keeping him unmistakably the same person. Preserve the exact facial structure, eye identity, nose, mouth, skin tone, hairstyle shape, complete outfit, pose, background, lighting, and composition from <Picture 1>. Change only age-related features: mature skin texture, wrinkles, subtle facial-volume changes, and plausible hair graying or thinning. One realistic 2:3 portrait photograph.\n```\n \n## 5. Obese body transformation\n \n### Task\n \nTransform the referenced person into a severely obese version while preserving identity, outfit design, pose, and setting.\n \n### Exact prompt\n \n```text\nTask: Reference-guided generation. Make a dramatic, unmistakable body-size transformation of the adult from <Picture 1>. Keep his identity, face, eyes, hairstyle, expression, skin tone, outfit design, pose, studio background, and photographic style, but do not preserve his original body proportions. Depict him with severe obesity and a very large, heavy body: an extremely large protruding round abdomen, much wider waist and torso, thick upper arms, large hips and thighs, a broad neck, and clearly increased facial fullness. The substantial weight must be immediately obvious across his entire silhouette, with anatomically coherent, realistic fat distribution. He must not look merely broad, muscular, stocky, chubby, or slightly overweight. Resize the jacket, shirt, and trousers so they fit the much larger body realistically. Complete full-body 2:3 portrait, both feet visible.\n```\n \n## 6. Through blinds\n \n### Task\n \nPlace the referenced person behind partially open venetian blinds with realistic occlusion and striped light.\n \n### Exact prompt\n \n```text\nTask: Reference-guided generation. The same recognizable adult from <Picture 1> stands indoors beyond a foreground set of partially open venetian blinds. The camera looks through the slats; alternating horizontal bands occlude small portions of the person and cast plausible striped window light across the face, outfit, and room. Both eyes remain readable through one opening. Source-authentic identity and clothing, physically correct depth and occlusion, cinematic 2:3 portrait photograph.\n```\n \n## 7. Mirror\n \n### Task\n \nCreate a physically consistent mirror scene showing both the person and their reflection.\n \n### Exact prompt\n \n```text\nTask: Reference-guided generation. The recognizable adult from <Picture 1> stands naturally in front of a large vertical dressing mirror in a modest bedroom. One direct three-quarter view and one geometrically correct reflected view show the same person at the same instant. The reflection reverses the visible pose and room consistently, matches the exact face, hair, outfit, expression, and lighting, and contains no extra or missing limbs. The person's torso, shoulders, head, and eyes all face squarely toward the mirror. The camera views him from slightly behind and to one side; he never looks toward the camera. His eyes visibly focus on his own reflected eyes. Realistic 2:3 portrait photograph with the mirror frame fully visible.\n```\n \n## 8. Several characters\n \n### Task\n \nCompose three separately referenced people in one conversational scene while preserving each identity and assigned pose.\n \n### Exact prompt\n \n```text\nTask: Reference-guided generation. The two distinct adult women from <Picture 1> and <Picture 3> and the adult man from <Picture 2> sit together at one table in the location from <Picture 4>, having a relaxed conversation. Each person retains their own exact identity, face, hairstyle, body build, and outfit without blending. The woman from <Picture 1> speaks and gestures with her raised right hand while her left palm rests on the table. The man from <Picture 2> looks toward her and cradles a mug with both hands. The woman from <Picture 3> also looks toward the speaker, resting her right elbow on the table with that hand near her chin while her left hand stays in her lap. Their arm positions are clearly different and do not mirror one another. Eye lines, seating, hands, table contact, scale, lighting, and perspective agree. Coherent medium-wide 2:3 portrait photograph, all three faces and bodies clearly separated.\n```\n \n## 9. Head swap\n \n### Task\n \nReplace only the base character's head with the head identity from a second reference.\n \n### Exact prompt\n \n```text\nTask: Reference-guided generation. Generate a new portrait image using <Picture 1> as the base. <Picture 1> supplies the complete body, build, clothing, pose, camera, lighting, and setting. <Picture 2> supplies the head identity: face, facial proportions, skin tone, ears, hairline, hairstyle, and facial hair. Replace only the head from <Picture 1> with the recognizable head from <Picture 2>. Match neck connection, scale, angle, perspective, and lighting naturally. Keep everything below the neck and the full 2:3 portrait composition from <Picture 1>.\n```\n \n## 10. Western cartoon\n \n### Task\n \nConvert the referenced portrait into a contemporary Western cartoon while retaining recognizable visual attributes.\n \n### Exact prompt\n \n```text\nTask: Reference-guided generation. A clean contemporary Western cartoon illustration depicts the same recognizable adult from <Picture 1>. The exact facial structure, eye shape, hairstyle silhouette, body proportions, outfit design, pose, expression, and source composition remain identifiable. Confident controlled linework, simplified geometric forms, flat color with subtle cel shading, restrained texture, mature editorial-animation styling, single 2:3 portrait composition, no text or decorative frame.\n```\n \n## 11. Three-koma storyboard\n \n### Task\n \nCreate a three-panel vertical storyboard with consistent character, wardrobe, setting, and progressive action.\n \n### Exact prompt\n \n```text\nTask: Reference-guided generation. A vertical three-koma storyboard uses three equal stacked panels on one 2:3 portrait canvas. The same adult from <Picture 1> wears the complete outfit from <Picture 2> in the location from <Picture 3> throughout. Panel 1: the character walks along the path and notices a paper map caught on a low branch. Panel 2: the character stretches upward and frees the map with one hand. Panel 3: the character stops, unfolds the map with both hands, and studies it with a pleased small smile. Face, hair, body, clothing, weather, light, camera side, and environment remain stable across all panels; action and expression alone progress. Clean panel gutters, no captions, speech balloons, labels, or text.\n```\n </code></pre>\n<p> </p>",
"author": {
"name": "admin"
},
"tags": [
"minimax",
"comfyui workflow",
"comfyui"
],
"date_published": "2026-08-11T01:51:01-04:00",
"date_modified": "2026-08-14T06:42:02-04:00"
},
{
"id": "https://graphicdesigngeek.com/replace-a-word-in-multiple-text-files-on-windows-10.html",
"url": "https://graphicdesigngeek.com/replace-a-word-in-multiple-text-files-on-windows-10.html",
"title": "Replace A Word In Multiple Text Files On Windows 10",
"summary": "Text file editors like Notepad and Notepad++ are used to create lots of different types of files like subtitles, log files, batch files, PowerShell scripts,…",
"content_html": "<p>Text file editors like Notepad and Notepad++ are used to create lots of different types of files like subtitles, log files, batch files, PowerShell scripts, and more. Where a text file editor can create these files, it can also edit them. If you have a lot of text files, ones that have the TXT file extension and you need to replace a word, or several words in them you can do so with a PowerShell script. The script makes it so you don’t have to individually open each file and then replace the word. You can use this same script for other file types that can be created with a text file editor. Here’s how you can replace a word in multiple text files.</p>\n<h2>Replace Word In Text Files</h2>\n<p>First, you need to put all your text files in the same folder. The script will examine only one directory when it runs and not your entire system which is why you need all the files in one place.</p>\n<p>Open a new Notepad file and paste the following in it.</p>\n<pre>Get-ChildItem 'Path-to-files\\*.txt' -Recurse | ForEach {\n(Get-Content $_ | ForEach { $_ -replace 'Original-Word', 'New-Word' }) |\nSet-Content $_\n}</pre>\n<p>You need to edit the above script. First, edit the ‘Path-to-files’ part with the actual path to the folder with all the text files in it. Second, replace the ‘Original-Word’ with the word you want to replace. Finally, replace the ‘New-Word’ with the word you want to replace the old one with. For example, I have a few text files that all have the word ‘Post’ in them. I want to replace the word Post with Article. This is what the script will look like once I’ve edited it to suit my scenario.</p>\n<pre>Get-ChildItem 'C:\\Users\\fatiw\\Desktop\\notepad-files\\*.txt' -Recurse | ForEach {\n(Get-Content $_ | ForEach { $_ -replace 'Post', 'Article' }) |\nSet-Content $_\n}</pre>\n<p>Once you’ve edited the script, save it with the ps1 file extension. Make sure you change the file type from text files to all files in Notepad’s save as dialog. Run the script and it will perform the replace function.</p>\n<figure class=\"not-lazy alignnone size-full wp-image-278691\"><picture><source type=\"image/webp\" srcset=\"https://www.addictivetips.com/app/uploads/2018/08/replace-word-txt-file.webp\"><source type=\"image/jpeg\" srcset=\"https://www.addictivetips.com/app/uploads/2018/08/replace-word-txt-file-1.jpg\"><img loading=\"lazy\" src=\"data:image/jpeg;base64,/9j/4AAQSkZJRgABAQAAAQABAAD/2wCEAAEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAQEBAf/AABEIAAUACgMBEQACEQEDEQH/xAGiAAABBQEBAQEBAQAAAAAAAAAAAQIDBAUGBwgJCgsQAAIBAwMCBAMFBQQEAAABfQECAwAEEQUSITFBBhNRYQcicRQygZGhCCNCscEVUtHwJDNicoIJChYXGBkaJSYnKCkqNDU2Nzg5OkNERUZHSElKU1RVVldYWVpjZGVmZ2hpanN0dXZ3eHl6g4SFhoeIiYqSk5SVlpeYmZqio6Slpqeoqaqys7S1tre4ubrCw8TFxsfIycrS09TV1tfY2drh4uPk5ebn6Onq8fLz9PX29/j5+gEAAwEBAQEBAQEBAQAAAAAAAAECAwQFBgcICQoLEQACAQIEBAMEBwUEBAABAncAAQIDEQQFITEGEkFRB2FxEyIygQgUQpGhscEJIzNS8BVictEKFiQ04SXxFxgZGiYnKCkqNTY3ODk6Q0RFRkdISUpTVFVWV1hZWmNkZWZnaGlqc3R1dnd4eXqCg4SFhoeIiYqSk5SVlpeYmZqio6Slpqeoqaqys7S1tre4ubrCw8TFxsfIycrS09TV1tfY2dri4+Tl5ufo6ery8/T19vf4+fr/2gAMAwEAAhEDEQA/AP7fP+FeeEmaV57PUp5JppZ5DJ4h1wLumYvIiot+AITIWZI33+WreUrCFUReSWU4OcpSlGo5Sbk/3+JSu97JVkkvJKy6aGXsabbdndu/xz/Lmsfx/fGj/gpFH4Y+MXxY8N2PwC0hrLw98S/Heh2bH4vfFe2LWuk+KdVsLdjbWOs2tlbkxW6HyLO1trWL/V28EMKpGv4lieMvY4jEUoZRhOSlXq04c1atKXLCpKMbyau3ZK76vU85qnd2pQ3fRPr3au/U/wD/2Q==\" width=\"1247\" height=\"600\" decoding=\"async\" fetchpriority=\"high\" data-is-external-image=\"true\"></figure></picture></p>\n<p>If you want to use this same script for XML or LOG files, edit the file extension in the first line. For example,</p>\n<p>This will become</p>\n<pre>Get-ChildItem 'C:\\Users\\fatiw\\Desktop\\notepad-files\\*.txt'</pre>\n<p>This;</p>\n<pre>Get-ChildItem 'C:\\Users\\fatiw\\Desktop\\notepad-files\\*.xml'</pre>\n<p>There is one thing you ought to know about this script; it doesn’t match words to words. If you’re looking to replace every occurrence of ‘the’ with ‘a’, it will also replace the ‘the’ at the start of ‘these’ and ‘there’. That is a shortcoming of this script. To work around it, you can use Notepad++ which has a match word option.</p>\n<p> </p>\n<p>https://www.addictivetips.com/windows-tips/replace-a-word-in-multiple-text-files-windows-10/</p>",
"author": {
"name": "admin"
},
"tags": [
],
"date_published": "2026-07-27T18:43:21-04:00",
"date_modified": "2026-07-27T18:43:21-04:00"
},
{
"id": "https://graphicdesigngeek.com/useful-tips-for-datasets.html",
"url": "https://graphicdesigngeek.com/useful-tips-for-datasets.html",
"title": "useful tips for datasets",
"summary": "Just go to the link and download the dataset and look at how it was prepared. https://civitai.com/models/2771332/krea2-john-william-waterhouse AiToolKit no longer requires square images but you…",
"content_html": "<p>Just go to the link and download the dataset and look at how it was prepared.</p>\n<p><a rpl=\"\" class=\"relative pointer-events-auto a\n \n \n \n \n underline\n \n cursor-pointer\" href=\"https://civitai.com/models/2771332/krea2-john-william-waterhouse\" rel=\"noopener nofollow ugc\" target=\"_blank\">https://civitai.com/models/2771332/krea2-john-william-waterhouse</a></p>\n<p>AiToolKit no longer requires square images but you want all your image with clean aspect ratios.</p>\n<p>1:1, 1:2, 1:3, 2:3, 1:4</p>\n<p>And you want each edge to be a multiple of 256.</p>\n<p>---</p>\n<p><a rpl=\"\" class=\"relative pointer-events-auto a\n \n \n \n \n underline\n \n cursor-pointer\" href=\"https://www.youtube.com/watch?v=OCsqHdHf81M\" rel=\"noopener nofollow ugc\" target=\"_blank\">https://www.youtube.com/watch?v=OCsqHdHf81M</a></p>\n<p>In this video he shows you how to use ChatGPT to do the caption. I do not recommend using his markdown file though.</p>\n<p>---</p>\n<p>Copy and paste and save this...it's my captioning system. Just paste it into ChatGPT and then upload zip files of the images.</p>\n<p>Just change the first two sentences to match your dataset.</p>\n<p>Gil Elvgren LoRA Captioning Instructions</p>\n<p>You are captioning a ZIP file containing pin-up artwork by Gil Elvgren for LoRA training.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">Required Header</h1>\n<p>Begin every caption with this exact line:</p>\n<p>Pin-up painting, style of Gil Elvgren, gilelvg</p>\n<p>Add one blank line after the header.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">Core Accuracy Rule</h1>\n<p>Caption only details that are visibly present in the individual image.</p>\n<p>Do not infer missing clothing, shoes, jewelry, stockings, garters, props, materials, colors, patterns, hairstyles, makeup, nail polish, or background objects.</p>\n<p>When a detail is ambiguous, omit it.</p>\n<p>Never complete an outfit based on what would normally match the theme.</p>\n<p>Examples:</p>\n<p>Do not add heels when the feet are outside the frame.<br>Do not call fabric silk, satin, leather, lace, or chiffon unless the material is visually identifiable.<br>Do not add stockings merely because garters are present.<br>Do not add garters merely because stockings are present.<br>Do not assume red nail polish unless the nails are visible and clearly red.<br>If gloves cover the hands or fingers, omit the Nails line unless the nails are still clearly visible.<br>Do not assume earrings, bracelets, necklaces, hats, gloves, or hair accessories.<br>Do not infer an object from the general scene when it cannot be clearly identified.</p>\n<p>It is better to omit one uncertain detail than to add one incorrect detail.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">Individual-Image Workflow</h1>\n<p>Do not caption from contact sheets.</p>\n<p>Open and inspect every original image separately at full available resolution.</p>\n<p>Complete the batch using this process:</p>\n<ol>\n<li>\n<p>Open one individual image.</p>\n</li>\n<li>\n<p>Inspect the entire image.</p>\n</li>\n<li>\n<p>Inspect the face, hair, hands, clothing, legs, feet, props, and background separately.</p>\n</li>\n<li>\n<p>Write the caption for that image.</p>\n</li>\n<li>\n<p>Inspect the same image a second time.</p>\n</li>\n<li>\n<p>Verify every line of the caption against the image.</p>\n</li>\n<li>\n<p>Amend or remove any unsupported line.</p>\n</li>\n<li>\n<p>Save the caption as a matching <code>.txt</code> file.</p>\n</li>\n<li>\n<p>Continue to the next individual image.</p>\n</li>\n<li>\n<p>Zip all completed <code>.txt</code> captions only after the entire batch has been verified.</p>\n</li>\n</ol>\n<p>Create a working folder for the captions and keep the completed files there until the batch is finished.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">Verification Standard</h1>\n<p>During the second inspection, check every caption line with these questions:</p>\n<p>Is every noun in this line visibly present?<br>Is the color accurate?<br>Is the garment type accurate?<br>Is the material clearly identifiable?<br>Is the pattern actually visible?<br>Is the body orientation correct?<br>Are the correct arms and legs described?<br>Is the object held by the correct hand?<br>Are the feet visible?<br>Are shoes actually visible?<br>Are stockings or garter bands clearly visible?<br>Is the hairstyle described from visible structure rather than assumption?<br>Are the nails directly visible?<br>Are gloves obscuring the nails?<br>Is each background object identifiable?<br>Did I add a conventional pin-up detail that is not actually shown?</p>\n<p>Delete or simplify any line that fails verification.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">Required Sections</h1>\n<p>Use these sections when applicable:</p>\n<p>Concept<br>Pose<br>Attire<br>Hair Makeup Nails<br>Expression<br>Background</p>\n<p>Props may be included as a separate section when the image contains several important handheld or scene objects.</p>\n<p>Do not include an empty section.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">Concept Section</h1>\n<p>Write one concrete sentence describing the basic visible scene.</p>\n<p>Prioritize the woman, her action, the main prop, and the setting.</p>\n<p>Avoid subjective or conceptual language such as:</p>\n<p>glamorous<br>seductive<br>luxurious<br>enchanting<br>playful atmosphere<br>elegant mood<br>cinematic<br>dramatic beauty</p>\n<p>Use concrete descriptions instead.</p>\n<p>Example:</p>\n<p>Blonde woman seated on a wooden ladder while holding several books in a library</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">Pose Section</h1>\n<p>Write one concrete pose detail per line.</p>\n<p>Use separate lines rather than a paragraph.</p>\n<p>Example:</p>\n<p>Pose<br>Full body front-facing view<br>Standing with legs apart<br>Left hand resting on hip<br>Right hand holding a paintbrush<br>Torso angled slightly right<br>Head tilted slightly left<br>Eyes looking toward viewer</p>\n<p>Describe only what is visible.</p>\n<p>Useful pose details include:</p>\n<p>full body<br>three-quarter body<br>front view<br>side view<br>three-quarter back view<br>back view<br>seated<br>standing<br>kneeling<br>reclining<br>bending forward<br>leaning backward<br>weight resting on one leg<br>legs crossed at knees<br>legs crossed at ankles<br>one knee raised<br>one arm extended<br>hand resting on hip<br>head turned over shoulder<br>gaze direction</p>\n<p>Do not confuse overlapping legs with crossed legs.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">Attire Formatting</h1>\n<p>Use one line for each visible garment or accessory category.</p>\n<p>Combine all details about the same item on one line using commas.</p>\n<p>Correct:</p>\n<p>Dress: white summer dress, fitted bodice, plunge neckline, short puff sleeves, full knee-length skirt, scalloped lace hem<br>Panties: pale pink high-waisted panties, glossy sheen, dark blue floral embroidery at hips<br>Stockings: sheer black nylon thigh-high stockings, wide opaque garter bands, visible back seams<br>Heels: black closed-toe pumps, pointed toes, slender high heels<br>Gloves: white wrist-length gloves</p>\n<p>Incorrect:</p>\n<p>Dress: white dress<br>Dress: fitted bodice<br>Dress: short sleeves<br>Dress: lace hem</p>\n<p>Do not repeat identical labels on multiple lines.</p>\n<p>Use specific category names when visible:</p>\n<p>Dress<br>Blouse<br>Shirt<br>Top<br>Bra<br>Corset<br>Bodice<br>Jacket<br>Skirt<br>Shorts<br>Pants<br>Panties<br>Garter belt<br>Garters<br>Stockings<br>Socks<br>Shoes<br>Heels<br>Boots<br>Gloves<br>Hat<br>Scarf<br>Belt<br>Necklace<br>Earrings<br>Bracelet<br>Robe<br>Apron<br>Swimsuit<br>Bikini top<br>Bikini bottoms</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">Nudity and Bare Feet</h1>\n<p>When no clothing is visible on the upper body, use:</p>\n<p>Upper body: nude</p>\n<p>When no clothing is visible on the lower body, use:</p>\n<p>Lower body: nude</p>\n<p>When both are visible and nude, include both lines.</p>\n<p>When the feet are visible and no shoes or socks are worn, use:</p>\n<p>Feet: bare</p>\n<p>When the feet are outside the frame or obscured, omit footwear entirely.</p>\n<p>Do not write barefoot unless the bare feet are actually visible.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">Hair Makeup Nails Formatting</h1>\n<p>Use one consolidated line for each category.</p>\n<p>Correct:</p>\n<p>Hair: blonde shoulder-length hair, large curled waves, fringe bangs<br>Makeup: dark eyeliner, long lashes, blue eyeshadow, pink blush, glossy red lipstick<br>Nails: red nail polish</p>\n<p>Incorrect:</p>\n<p>Hair: blonde hair<br>Hair: shoulder-length hair<br>Hair: curled waves</p>\n<p>Incorrect:</p>\n<p>Makeup: dark eyeliner<br>Makeup: long lashes<br>Makeup: pink blush<br>Makeup: red lipstick</p>\n<p>Describe only visible features.</p>\n<p>Possible hair details:</p>\n<p>hair color<br>approximate length<br>straight<br>wavy<br>curled<br>ringlets<br>victory rolls<br>rolled bangs<br>fringe bangs<br>side part<br>center part<br>ponytail<br>bun<br>updo<br>loose curls<br>hair ribbon<br>flower accessory</p>\n<p>Do not identify a hairstyle as victory rolls unless the rolled structure is clearly visible.</p>\n<p>Possible makeup details:</p>\n<p>dark eyeliner<br>winged eyeliner<br>long lashes<br>blue eyeshadow<br>green eyeshadow<br>pink blush<br>red lipstick<br>pink lipstick<br>glossy lipstick<br>defined brows</p>\n<p>Omit makeup details that cannot be resolved from the image.</p>\n<p>For nails use:</p>\n<p>Nails: red nail polish</p>\n<p>Do not use:</p>\n<p>Nails: red manicure</p>\n<p>Omit the Nails line when the nails are not clearly visible.</p>\n<p>If gloves cover the hands or fingers, omit the Nails line unless the nails remain directly visible.</p>\n<p>Never caption nail polish through gloves.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">Expression Section</h1>\n<p>Use concrete facial observations.</p>\n<p>Examples:</p>\n<p>Expression<br>Wide-eyed surprised expression<br>Raised eyebrows<br>Rounded open mouth forming an oh shape<br>Eyes looking toward viewer</p>\n<p>Expression<br>Broad smile<br>Eyes looking toward viewer</p>\n<p>Expression<br>Focused expression<br>Eyes looking downward toward the book</p>\n<p>Avoid interpreting emotions beyond visible facial features.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">Background Section</h1>\n<p>Caption only identifiable visible background elements.</p>\n<p>Use one object or closely related group per line.</p>\n<p>Example:</p>\n<p>Background<br>Tall wooden bookshelves filled with books<br>Wooden library ladder<br>Several books falling through the air<br>Pale wooden floor</p>\n<p>Do not describe lighting unless it is an important visible element requested by the user.</p>\n<p>Do not add generic environmental objects to make the scene feel complete.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">Props Section</h1>\n<p>Use a Props section when several important objects are interacting with the subject.</p>\n<p>Example:</p>\n<p>Props<br>Open black suitcase filled with clothing<br>Small black dog pulling a garment from the suitcase<br>Red travel tag attached to suitcase handle</p>\n<p>Only describe identifiable objects.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">Language Style</h1>\n<p>Use concrete nouns and restrained adjectives.</p>\n<p>Good:</p>\n<p>red full skirt<br>wooden chair<br>black dog<br>white towel<br>round hand mirror<br>sheer black stockings<br>curled blonde hair<br>open suitcase<br>metal ladder<br>blue wall</p>\n<p>Avoid unnecessary aesthetic terms:</p>\n<p>gorgeous<br>sultry<br>alluring<br>luxurious<br>romantic<br>dreamy<br>elegant<br>sophisticated<br>captivating<br>glamorous</p>\n<p>Do not mention artistic technique, brushwork, composition quality, or painterly atmosphere beyond the required header.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">File Handling</h1>\n<p>Each image receives one matching <code>.txt</code> file.</p>\n<p>Preserve the image filename exactly, changing only the extension to <code>.txt</code>.</p>\n<p>Example:</p>\n<p>GilElvgren (31).jpg</p>\n<p>becomes:</p>\n<p>GilElvgren (31).txt</p>\n<p>If two images share the same stem but have different extensions, add a short extension identifier so neither caption is overwritten.</p>\n<p>Place all completed caption files in one folder.</p>\n<p>After every image has been individually inspected and every caption has been verified against its original image a second time, zip the <code>.txt</code> files and provide the ZIP for download.</p>\n<h1 class=\"text-24-scalable xs:text-20-scalable\">Final Quality Rule</h1>\n<p>Accuracy is more important than caption length.</p>\n<p>A shorter caption containing only verified details is better than a detailed caption containing one hallucinated item.</p>\n<p> </p>\n<p>https://www.reddit.com/r/StableDiffusion/comments/1v4we4h/the_wonders_of_krea2/</p>",
"author": {
"name": "admin"
},
"tags": [
],
"date_published": "2026-07-26T05:16:10-04:00",
"date_modified": "2026-07-26T05:16:10-04:00"
},
{
"id": "https://graphicdesigngeek.com/krea-usage-tips.html",
"url": "https://graphicdesigngeek.com/krea-usage-tips.html",
"title": "Krea usage tips",
"summary": "Use the raw model with the turbo lora applied to it at 0.6 weight. This looks much better than turbo by itself. Use about 12…",
"content_html": "<ul>\n<li>\n<p>Use the raw model with the turbo lora applied to it at 0.6 weight. This looks much better than turbo by itself. Use about 12 steps instead of 8. For 1 CFG use euler, beta / beta57 scheduler. (feel free to find a better balance of lora weight vs steps.) <a rpl=\"\" class=\"relative pointer-events-auto a\n \n \n \n \n underline\n \n cursor-pointer\" href=\"https://huggingface.co/Comfy-Org/Krea-2/blob/main/loras/krea2_turbo_lora_rank_64_bf16.safetensors\" rel=\"noopener nofollow ugc\" target=\"_blank\">https://huggingface.co/Comfy-Org/Krea-2/blob/main/loras/krea2_turbo_lora_rank_64_bf16.safetensors</a></p>\n</li>\n<li>\n<p>Use this vae: <a rpl=\"\" class=\"relative pointer-events-auto a\n \n \n \n \n underline\n \n cursor-pointer\" href=\"https://huggingface.co/spacepxl/Wan2.1-VAE-upscale2x\" rel=\"noopener nofollow ugc\" target=\"_blank\">https://huggingface.co/spacepxl/Wan2.1-VAE-upscale2x</a> its better quality. It needs either kijai's diffusers WF or this: <a rpl=\"\" class=\"relative pointer-events-auto a\n \n \n \n \n underline\n \n cursor-pointer\" href=\"https://github.com/spacepxl/ComfyUI-VAE-Utils\" rel=\"noopener nofollow ugc\" target=\"_blank\">https://github.com/spacepxl/ComfyUI-VAE-Utils</a> Still playing with other merges as well. (merge qwen and the wan vae tunes at like 0.5)</p>\n</li>\n<li>\n<p>Use the uncensor Lora: <a rpl=\"\" class=\"relative pointer-events-auto a\n \n \n \n \n underline\n \n cursor-pointer\" href=\"https://civitai.red/models/2728234/krea2filterbypass?modelVersionId=3067151\" rel=\"noopener nofollow ugc\" target=\"_blank\">https://civitai.red/models/2728234/krea2filterbypass?modelVersionId=3067151</a> Comparison of loras vs node (lora is better): <a rpl=\"\" class=\"relative pointer-events-auto a\n \n \n \n \n underline\n \n cursor-pointer\" href=\"https://files.catbox.moe/9zpoqs.png\" rel=\"noopener nofollow ugc\" target=\"_blank\">https://files.catbox.moe/9zpoqs.png</a> It also helps with expressions and other prompt responsiveness. The censorship hurt more than just nudity.</p>\n</li>\n<li>\n<p>It can use reference images: <a rpl=\"\" class=\"relative pointer-events-auto a\n \n \n \n \n underline\n \n cursor-pointer\" href=\"https://github.com/ethanfel/ComfyUI-Krea2TextEncoder\" rel=\"noopener nofollow ugc\" target=\"_blank\">https://github.com/ethanfel/ComfyUI-Krea2TextEncoder</a></p>\n</li>\n<li>\n<p>And be more specific with your prompts. For example. \"Teen Titans\" gives you crappy Teen Titans Go. But \"Teen Titans (2003)\" gives you the good version. Same for other series / characters / people. Try changing your capitalization / specifying the series / sometimes the year.</p>\n</li>\n<li>\n<p>It knows TONS of artists by name, a undersung feature. Use \"In the style of\" or \"Painted by\", ect...</p>\n</li>\n<li>\n<p>For mixing characters use \"Cosplaying as\" instead of \"Wearing X's outfit\" or try something like \"With the body of X and the head of X in the style of X...\" <a rpl=\"\" class=\"relative pointer-events-auto a\n \n \n \n \n underline\n \n cursor-pointer\" href=\"https://files.catbox.moe/sveyky.png\" rel=\"noopener nofollow ugc\" target=\"_blank\">https://files.catbox.moe/sveyky.png</a> Be specific.</p>\n</li>\n<li>\n<p>For even better realism try this lora: <a rpl=\"\" class=\"relative pointer-events-auto a\n \n \n \n \n underline\n \n cursor-pointer\" href=\"https://civitai.red/models/2727284/realism-enhancer-krea2?modelVersionId=3065628\" rel=\"noopener nofollow ugc\" target=\"_blank\">https://civitai.red/models/2727284/realism-enhancer-krea2?modelVersionId=3065628</a> For more uncensored stuff try: <a rpl=\"\" class=\"relative pointer-events-auto a\n \n \n \n \n underline\n \n cursor-pointer\" href=\"https://civitai.red/models/2688234/realism-engine-ideogram-4-krea-2?modelVersionId=3067451\" rel=\"noopener nofollow ugc\" target=\"_blank\">https://civitai.red/models/2688234/realism-engine-ideogram-4-krea-2?modelVersionId=3067451</a></p>\n</li>\n</ul>\n<p>https://www.reddit.com/r/StableDiffusion/comments/1ueybt8/some_important_krea_usage_tips_ive_found_not_seen/</p>",
"author": {
"name": "admin"
},
"tags": [
"krea2",
"comfyui"
],
"date_published": "2026-07-25T20:47:41-04:00",
"date_modified": "2026-08-11T05:37:45-04:00"
},
{
"id": "https://graphicdesigngeek.com/substackcom.html",
"url": "https://graphicdesigngeek.com/substackcom.html",
"title": "substack.com",
"summary": "Old media models are broken. Substack is building a new one that puts writers, creators, and subscribers in charge https://substack.com/about",
"content_html": "<p> </p>\n<h1 id=\"about-v2-hero-heading\" class=\"pencraft pc-reset align-center-y7ZD4w weight-regular-mUq6Gb reset-IxiVJZ marketingHeading1-s3rHZC heroTitle-Js5T12\">Make money doing the work you believe in.</h1>\n<p class=\"pencraft pc-reset color-primary-zABazT align-center-y7ZD4w font-text-qe4AeH reset-IxiVJZ marketingBody-sFV4pW heroSubtitle-JySyhe\">Old media models are broken. Substack is building a new one that puts writers, creators, and subscribers in charge</p>\n<p>https://substack.com/about</p>",
"author": {
"name": "admin"
},
"tags": [
],
"date_published": "2026-07-23T20:30:24-04:00",
"date_modified": "2026-07-23T20:30:24-04:00"
},
{
"id": "https://graphicdesigngeek.com/res4lyf-samplers-and-schedulers-plain-language-guide.html",
"url": "https://graphicdesigngeek.com/res4lyf-samplers-and-schedulers-plain-language-guide.html",
"title": "RES4LYF Samplers & Schedulers – Plain-Language Guide",
"summary": "Summary: RES4LYF is a ComfyUI extension offering a rich collection of diffusion samplers and sigma schedulers. It introduces new high-accuracy samplers (denoted RES for rectified…",
"content_html": "<p><strong>Summary:</strong> <em>RES4LYF</em> is a ComfyUI extension offering a rich collection of diffusion samplers and sigma schedulers. It introduces new high-accuracy samplers (denoted <strong>RES</strong> for rectified explicit solvers) that achieve high image quality in fewer steps. It also provides an <em>ODE vs. SDE</em> toggle: in <strong>ODE mode</strong>, the sampler follows a fixed deterministic path, whereas <strong>SDE mode</strong> adds controlled noise at each step (via an “eta” parameter) for stochastic exploration. This guide demystifies each sampler and scheduler: what it does, when to use it, how many steps to try, and which scheduler pairs best – all backed by source documentation and tests.</p>\n<p>Taxonomy of RES4LYF Sampling Methods<br>┌─────────────────────────┐ ┌─────────────────────────────┐<br>│ <strong>Explicit Methods</strong> │ │ <strong>Implicit Methods</strong> │<br>│ (direct step updates) │ │ (solve via iterative root-finding) │<br>├─────────────────────────┤ ├─────────────────────────────┤<br>│ <strong>Linear (ERK)</strong> │ │ <strong>Diagonally Implicit</strong> │<br>│ e.g. Euler, RK4 │ │ (DIRK, lower-triangular) │<br>│ – standard Runge-Kutta │ │ e.g. Crouzeix 2s, Qin–Zhang │<br>│ – fast but less stable │ │ – one new implicit stage per step │<br>├─────────────────────────┤ ├─────────────────────────────┤<br>│ <strong>Multistep</strong> │ │ <strong>Fully Implicit</strong> │<br>│ e.g. DPM++ 2M, RES_3M │ │ (fully coupled stages) │<br>│ – reuse past steps as │ │ e.g. Gauss–Legendre, Radau │<br>│ cheap predictors│ │ – high stability, high cost │<br>├─────────────────────────┤ ├─────────────────────────────┤<br>│ <strong>Exponential (ETD)</strong> │ │ <strong>Hybrid</strong> (Predictor–Corrector) │<br>│ e.g. RES_4S_Krogstad │ │ e.g. Lawson 4-1* methods │<br>│ – integrate linear part │ │ – combine explicit + implicit/exponential │<br>│ exactly (Lawson/ETD)│ │ – refine solution with extra steps│<br>└─────────────────────────┘ └─────────────────────────────┘<br>*ODE vs. SDE:* In any category, you can switch between a deterministic <strong>ODE</strong> mode and a stochastic <strong>SDE</strong> mode. <strong>SDE mode</strong> injects noise after each step, yielding more exploratory (creative) results at the cost of convergence – images won’t “lock in” even if steps are increased indefinitely, because noise keeps nudging the generation. <strong>ODE mode</strong> gives repeatable outputs and converges if you take enough steps (no noise added). Use ODE for precision and reproducibility, SDE for variety and potentially sharper detail when balanced with fewer steps of high-order integrators.*<br><br>## Sigma Schedulers (Intuitive Overview)<br><br>In diffusion models, a <em>scheduler</em> defines how the noise level (σ) decreases from start (high noise) to finish (no noise). Different schedulers space the steps <em>differently</em> along the time/σ axis, affecting image formation. RES4LYF supports all common schedules from KSampler plus a few extras. Here’s each in plain terms:<br><br>- <strong>Normal (Linear):</strong> Steps linearly reduce σ from σ_max to 0 – a <em>uniform decay</em> schedule. This is the “standard” schedule giving a balanced progression, often used as a baseline. Use for general purposes when unsure; it evenly spends effort across all denoising phases.<br>- <strong>Simple:</strong> A simplified decay (also linear in behavior) intended for quick tests. It’s similar to Normal but uses a fixed range or formula; basically a no-frills schedule. Use for smoke-testing a prompt or debugging, where fancy step allocation isn’t needed. (No separate official formula given in docs – treat it as a variant of linear <em>without</em> special optimization.)<br>- <strong>Karras (EDM):</strong> A smoothly varying non-linear schedule inspired by Karras et al.’s <em>EDM</em> paper. It spaces steps more densely at the low-noise end for finer detail, using a power-law (ρ) curve. Visually, σ decays slowly at first, then faster mid-way, then slowly at the end – ensuring a “smooth transition” in noise. This is recommended for high-quality, detail-rich images, especially if you have enough steps (≥ 20). Karras schedule often improves detail and stability compared to linear spacing by spending more steps where the model needs them (low σ).<br>- <strong>Exponential:</strong> An aggressive schedule where σ drops off exponentially fast. Early steps eliminate noise rapidly, so coarse structure appears quickly; later steps fine-tune at near-zero noise. It’s great for speed or when you want a decent image in very few steps (the model “jumps to” a clean state quickly). However, it may sacrifice some detail or consistency if too few steps are used (since little time is spent in mid-level noise). Use when you need a rough but fast result.<br>- <strong>SGM Uniform:</strong> A schedule derived from Score-Based Models (SGM) literature, aiming for <em>uniform progression in diffusion time</em>. In practice, it “decays noise in an SGM-optimized way” – likely linear in <em>continuous time</em> t instead of linear in σ. This can benefit certain models or settings that assume a <em>variance-preserving SDE</em> schedule. Use SGM for specialized scenarios or to mimic the original training schedule of SDE models (when known); otherwise, Normal or Karras usually suffices.<br>- <strong>DDIM Uniform:</strong> The schedule used by DDIM (Denoising Diffusion Implicit Models). It spaces time steps uniformly between 1 and 0 (the diffusion time), which corresponds to a specific non-linear σ decay. It’s tailored for DDIM’s ODE sampler to approximate the original discrete diffusion process. Use DDIM Uniform if you are employing a DDIM sampler or want to replicate the original DDIM paper’s step allocation. It tends to give a gentler noise drop early on, then a steep decline later (somewhat opposite to exponential).<br>- <strong>Beta (Train-like):</strong> A schedule based on the beta schedule from diffusion training. In training, noise level evolves according to a Beta distribution (often linear beta or cosine schedule). Here “Beta” refers to using a Beta(α, β) distribution function to map step indices to noise. The default parameters (in KSampler) aren’t given, but the variant <strong>Beta57</strong> in RES4LYF is explicitly Beta(α=0.5, β=0.7). These produce a curve that decays slowly then rapidly (or vice versa) depending on α, β. Use Beta when you want to experiment with non-linear decays that are <em>not</em> as extreme as exponential – Beta curves can allocate more time near start or end depending on shape. Beta(0.5,0.7) for instance gives a slightly front-loaded decay. (This is an extra provided by RES4LYF; it’s not in stock ComfyUI).<br>- <strong>Linear-Quadratic:</strong> A two-phase schedule: first half linear decay, second half quadratic (or vice versa). It’s described as <em>complex scenario optimization</em> – meaning it tries to handle both coarse and fine denoising differently. For example, it might drop σ linearly early on, then slow the decay quadratically to linger on fine details. Use this if you find pure linear either oversmooths or underrefines; linear_quadratic can balance structure vs detail by adjusting how abruptly it transitions.<br>- <strong>KL Optimal:</strong> A theoretically derived schedule that minimizes KL divergence error per step (from the <em>Align Your Steps</em> framework). In essence, it places steps where the diffusion model’s error is largest, which mathematically yields an optimal allocation of N steps to minimize final KL error. The formula is complex (involves solving a quartic per triplet of points), but the effect is an <em>uneven, model-specific spacing</em>. Use KL Optimal if you seek maximum fidelity for a given number of steps – e.g. when chasing the last bit of quality or doing very low-step generations where every step must count. It’s “optimal” in theory, though in practice results may only subtly improve. (Ensure your model and CFG settings are compatible; some implementations need adjusting “skip” for first steps.)<br>- **Bong (Tangent) – Two-Stage Tangential:** <strong>RES4LYF-exclusive.</strong> This scheduler (sometimes spelled <strong>boong_tangent</strong>) uses <em>BONGMATH</em> tech to perform <strong>bidirectional denoising</strong>: it essentially schedules a forward pass followed by a partial “reverse” pass in one sequence. Implementation-wise, it uses a two-segment tangent function curve (hence the name) to drop σ down and then slightly up, or similar trick, so that the sampler can correct errors from both ends. The result is enhanced detail: users report <em>bong_tangent is very fast and great for fine details</em> when paired with certain samplers. Visually, imagine denoising almost to completion, then adding a tiny bit of noise back and refining again – this can recover detail that a one-way schedule might blur. Use bong_tangent for <strong>maximum sharpness and detail</strong>, especially with simpler samplers (like Euler or RES_2M) that benefit from a second chance at refinement. It’s computationally efficient (the schedule shape doesn’t add actual steps, it just redistributes them), so it’s a popular choice for speedy yet crisp results. <em>Note:</em> This schedule may introduce subtle flicker in batch or video outputs due to its bidirectional nature – best to test on a single image first.<br><br>Each scheduler’s “curve” can be visualized as a sigma vs step plot (not shown here). The key takeaway is: <strong>different schedules emphasize different parts of the denoising process</strong>. For quick drafts use <em>exponential</em> or <em>simple</em>; for balanced quality use <em>normal</em> (linear) or <em>karras</em>; for specialized tweaks try <em>SGM, DDIM, Beta, LQ or KL</em>; and for detail-critical work, consider <em>bong_tangent</em>. Many users default to Karras or bong_tangent for quality, and switch to exponential for speed runs – as we’ll outline next.<br><br>## Quick Sampler+Scheduler Picks (Cheat Sheet)<br><br>If you’re not ready to digest all the theory, here are evidence-based starting points for common goals. Each suggestion cites either the RES4LYF docs or community benchmarks:<br><br>- <strong>“I want the best image quality in ~30 steps”</strong> – Try *<code>rk6_7s</code>** (Linear 7-stage) with <strong>Normal or Karras</strong> schedule. In one test, a 7-stage 6th-order method produced the <em>sharpest, most prompt-faithful image</em> among ~100 samplers. It took longer per image, but the result was top-tier.<br>- <strong>“I need a sharp portrait but can’t afford many steps”</strong> – Use *<code>res_2m</code>** (Multistep) with <strong>BongTangent</strong>. The RES multistep family is designed for high quality at low step counts. Users report that <code>res_2m + bong_tangent</code> scheduler yields very detailed images quickly (bong_tangent adds extra refinement without extra steps).<br>- <strong>“Speed is crucial (around 20 steps max)”</strong> – Choose *<code>Euler</code>** or *<code>DDIM</code>** with the <strong>Exponential</strong> scheduler. Euler (1-step RK) and DDIM are simple and fast, and an exponential decay will get the image “mostly done” in ~20 steps. In a benchmark, <em>Exponential DDIM and linear Euler each finished 16 steps in ~24.5s on a 3080 GPU</em> – the fastest of all tested. Expect a decent image, albeit a bit less polished.<br>- <strong>“Smooth gradients & fewer artifacts (e.g. landscapes, sky)”</strong> – Use an <em>implicit</em> sampler like *<code>radau_iia_3s</code>** (fully implicit 3-stage) with <strong>Linear-Quadratic</strong> schedule. Implicit methods excel at stability – Radau IIA is L-stable and dampens oscillations. Paired with a gentle decay (linear_quadratic), it preserves smooth tones. One user’s test noted Radau (11-stage in their case) gave <em>very natural, artifact-free results</em> albeit slowly.<br>- <strong>“Crank up the creativity (let the model roam)”</strong> – Pick <strong>SDE</strong> mode with a robust explicit solver, e.g. *<code>dpmpp_sde_2s</code>** (stochastic 2-stage) and <strong>Karras</strong> schedule. SDE mode’s injected noise gives unpredictable but often inspiring variations. DPM++ 2S is a well-regarded sampler for stable diffusion – combining them can yield varied outputs even with the same seed. Tip: keep <em>eta</em> moderate (~0.5) to avoid complete chaos.<br>- <strong>“Minimal ringing or overshoot (for high-contrast or cartoon images)”</strong> – Use a <em>strong stability preserving</em> method: *<code>ssprk3_3s</code>** (3-stage SSP) with <strong>Normal</strong> schedule. SSPRK3 is third-order and designed to avoid introducing oscillations (it’s literally built for monotonic decay in PDEs). It’s great for preserving edge integrity (no halos). The linear schedule keeps things simple to not upset SSP properties.<br>- <strong>“Highest detail regardless of speed (I’ll wait)”</strong> – Combine *<code>res_5s</code>** (5-stage refined explicit) as predictor and *<code>gauss-legendre_5s</code>** (5-stage Gauss collocation) as an implicit corrector, with <strong>Karras</strong> schedule. This two-step approach (set <code>IMPLICIT_SAMPLER_NAME</code> to gauss-legendre and a few implicit steps) “demonstrates your commitment to climate change (and image quality)” – i.e. it’s extremely slow, but yields exceptionally accurate denoising. Gauss-Legendre is 10th-order for 5 stages and symplectic; it will squeeze out detail that other methods miss. Use maybe 1–2 implicit refinement steps – beyond that returns diminish.<br>- <strong>“Make it like the training images”</strong> – Try *<code>ddim</code>** sampler with <strong>Beta</strong> scheduler. DDIM sampler with a beta(α, β) schedule can mimic the original diffusion trajectory (since many diffusion models used a beta schedule during training). For example, Beta(0.5,0.7) is available as “beta57”. This combination often yields the “vanilla” diffusion look, useful for comparisons or keeping color and composition closer to stock. <br><br>*(TL;DR: for everyday use, many power-users start with <strong>RES samplers (2M or 3M)</strong> and either <strong>Karras</strong> or <strong>Bong_Tangent</strong> schedule, at ~20–30 steps. That tends to give a good quality vs speed trade-off on SDXL and similar models.)*<br><br>---<br><br>Below, we provide detailed per-sampler breakdowns, a pairing recommendation matrix, and additional context (glossary, reproducible tests, and an FAQ). <strong>All claims are cited from RES4LYF’s documentation or authoritative references</strong> – <em>no guesswork</em>. If something isn’t documented, we’ll note that too.<br><br>## Sampler Reference Cards<br><br>Each sampler is listed with its type (**class**), a one-line integration idea, documented <em>order</em> & <em>stages</em>, any stability/behavior notes from docs, when to use it, recommended scheduler pairing (if known), and pitfalls (if any). Samplers are grouped by family.<br><br>### Multistep Samplers (reuse previous steps)<br><br>**Integration idea:** Multistep methods use earlier step results to extrapolate the next step, instead of substeps. They achieve higher order at low cost (only 1 model call per step), but may be less stable/accurate than multi-stage methods of the same order.<br><br>- <strong>res_2m</strong> – <em>Class:</em> multistep (RES family). <em>Idea:</em> A 2-step rectified Euler method (refined DPM++ 2M). *Order/Stability:* <em>Not explicitly stated</em>, but likely 2nd order (two-step) with enhancements. <em>Behavior:</em> Designed to converge linearly and reliably. The RES family are improved versions of DPM++ with higher accuracy, so <code>res_2m</code> can be seen as a high-precision 2-step solver. *Best for:* <strong>General use with speed</strong> – it’s very fast (Euler-speed) yet more accurate than vanilla Euler or even DPM2. Excels at img2img and guide-based tasks due to linear convergence (avoids overshoot). <em>Steps guidance:</em> The docs note **“typically only ~20 steps are needed with RES samplers”* – so try 20–30 steps, which often outperforms 50+ steps of common samplers. <em>Schedulers:</em> No exclusive pairing stated; however, community tests praise <strong>bong_tangent</strong> with res_2m for detail and speed. Also works well with Karras for quality or Exponential for quick drafts. <em>Pitfalls:</em> None documented explicitly. As a multistep method, it might struggle if the model’s behavior changes abruptly (because it relies on previous step linearity). But overall it’s a flagship sampler for quick, quality results.<br><br>- <strong>res_3m</strong> – <em>Class:</em> multistep. <em>Idea:</em> A 3-step rectified method (refined DPM++ 3M). <em>Order:</em> Not stated, but presumably 3rd order (uses three prior points). <em>Behavior:</em> Same philosophy as res_2m – <em>higher order and accuracy</em>, with minimal cost. The README’s showcase compares <strong>RES_3M</strong> vs UniPC and shows RES_3M achieving top quality in 20 steps. <em>Use for:</em> Cases needing slightly more accuracy than 2m – e.g. more complex scenes or animations where a bit more stability helps. It still converges linearly but to a higher-order target, potentially yielding sharper results at equal steps. <em>Steps:</em> ~20 is often enough; you can try 30–40 if image is very complex. <em>Schedulers:</em> No specific call-outs in docs. It performed excellently with the default “normal” schedule in tests (20 steps, normal gave a sharp output). Likely pairs well with <strong>Karras</strong> for detail at low steps (Karras emphasizes later refinement which a 3rd-order method can exploit). <em>Pitfall:</em> Not noted; but as with any multistep, divergence can happen if steps are too large – e.g. if you try <10 steps, a multi-step might overshoot where a multi-stage wouldn’t. Anecdotally, though, res_3m is robust and a go-to for high quality fast sampling.<br><br>- <strong>dpmpp_2m</strong> – <em>Class:</em> multistep. <em>Idea:</em> The classic DPM++ 2M (Karras et al. / Lu et al.) – a second-order two-step solver using a predictor-corrector with prior step reuse. <em>Order:</em> 2 (as per “2” in name). <em>Stages:</em> 1 per step (multi-step, not multi-stage). <em>Behavior:</em> Was known to produce good detail and hands in Stable Diffusion v1. RES4LYF includes it mainly for completeness (and baseline). <em>Use:</em> If you want to compare against standard samplers or ensure compatibility with results from other UIs. It’s stable and solid for most cases, though generally <strong>res_2m</strong> or <strong>res_3m</strong> outperform it in quality at the same steps. <em>Steps:</em> 20–50 depending on desired quality; no official guidance, but stable diffusion users often used 20–30 with DPM++ 2M Karras for good results. <em>Schedulers:</em> Often used with <strong>Karras</strong> (hence “DPM++ 2M Karras”) to improve convergence. Also fine with Normal or other spacings. <em>Pitfalls:</em> None major – but relative to newer samplers, it might require a few more steps to match quality. <br><br>- <strong>dpmpp_3m</strong> – <em>Class:</em> multistep. <em>Idea:</em> DPM++ 3M – a third-order multistep extension of DPM solver. <em>Order:</em> 3 (implied by “3”). <em>Behavior:</em> Should yield even better accuracy than 2M at the cost of using two prior points. Not extensively documented in RES4LYF (no direct mention of 3M’s internals), but logically it’s the multi-step variant of DPM-Solver++ order 3 from literature (Lu et al. 2022). <em>Use:</em> A middle ground between DPM++ 2M and explicit multi-stage samplers. If 2M leaves some haze, 3M might clear it up by that extra order. <em>Steps:</em> 15–25 likely suffice for SD1.5/SDXL. <em>Schedulers:</em> Karras is a safe bet (it was designed to work with DPM++ in Karras’ EDM paper). No special scheduler noted in docs. <em>Pitfall:</em> Slightly more risk of instability than 2M (higher order multi-step can oscillate if you push step size too far). If using very few steps (<15), monitor for odd artifacts – if seen, try falling back to 2M or using an implicit corrector.<br><br>- <strong>abnorsett_2m / 3m / 4m</strong> – <em>Class:</em> multistep. <em>Idea:</em> These are <strong>A</strong>dams–**B**ashforth or Norsett-inspired multi-step methods (the naming isn’t explained; likely honoring Dahlquist & Nørsett). They presumably use <em>adaptive coefficients or higher-order formulas</em> up to 4-step. <em>Order:</em> Not stated in docs. Given 4m uses four previous steps, it could reach up to 4th order. Possibly these implement Norsett’s multi-step predictor formulas (Nørsett co-authored a famous ODE text). <em>Behavior:</em> “AbNorsett” samplers may emphasize stability or error minimization (since Norsett’s work often did). They likely converge linearly like other multisteps but with more history to draw on. <em>Use:</em> If you have <em>lots</em> of steps to utilize (e.g. >50) or are doing something like a guided process where preserving past trend is good. However, <strong>not much documentation or usage anecdotes</strong> – meaning they are probably experimental. <em>Steps:</em> No official guidance. If trying <code>abnorsett_4m</code>, ensure you have at least 4–5 steps per order (~20+) to get benefit. <em>Schedulers:</em> No info. Try Normal or Karras first. <em>Pitfalls:</em> Not documented, but multi-step methods with many steps can suffer if initial steps are not well-behaved (the first few steps might need a kicker). Possibly not recommended for short runs or highly nonlinear transitions (like drastic style changes mid-sampling). In absence of docs, consider these advanced tools for special cases.<br><br>- <strong>deis_2m / 3m / 4m</strong> – <em>Class:</em> multistep. *Idea:* <em>Diffusion Exponential Integrator Sampler</em> (DEIS) in multi-step form. DEIS was introduced by Zhang & Chen (2022) as an exponential integrator approach to diffusion ODEs. Here, 2m/3m/4m likely correspond to multi-step versions akin to Adams methods applied to the DEIS formulation (the original paper had ρAB-DEIS which is a multistep variant). <em>Order:</em> Presumably 2, 3, 4 respectively (if analogous to DPM++ naming). <em>Behavior:</em> DEIS methods explicitly treat the linear part of the SDE by an integrating factor, achieving high quality with fewer steps. In multi-step form, they leverage previous gradients for efficiency. <em>Use:</em> DEIS samplers can be very fast and accurate for diffusion models, often outperforming DDIM at low step counts. Use these if you want a possibly sharper result than plain DPM++ but still want to avoid multi-stage cost. <em>Steps:</em> The original DEIS achieved good results in ~15 steps for CIFAR-like data; for SD, try ~20–30. <em>Schedulers:</em> Not stated explicitly; DEIS authors mention using their own time step scheme. You might try <strong>SGM_uniform</strong> or <strong>Karras</strong>, since DEIS leverages the continuous SDE solution (Karras schedule might complement it). <em>Pitfalls:</em> Not documented in RES4LYF. Empirically, DEIS may oversharpen at too high step counts (anecdotal forum reports outside RES4LYF). If using 4m with 50+ steps, watch for high-frequency noise or loss of diversity (the integrator might overshoot perfection and introduce detail that looks unnatural). Generally, though, DEIS is well-regarded as a stable method.<br><br>### Exponential Samplers (explicit Lawson/ETD-type integrators)<br><br>**Integration idea:** Exponential integrators decouple the ODE: they analytically integrate the <em>linear</em> part of the diffusion equation (often the noise term), and use RK for the nonlinear part. This often allows larger stable steps and better handling of stiffness (noise decay can be steep). Many of these are “Lawson” type (apply an integrating factor $e^{-t}$) or ETD (Exponential Time Differencing) schemes from PDE literature.<br><br>*(Note: All “_s” here indicate substeps per step, i.e. multi-stage. e.g. 3s = 3-stage explicit RK with exponential treatment.)*<br><br>- <strong>res_2s / res_2s_stable / res_2s_rkmk2e</strong> – <em>Class:</em> exponential explicit. <em>Idea:</em> These are 2-stage, 2nd-order exponential integrators, presumably unique to RES4LYF. “RKMK2e” suggests a Runge-Kutta Munthe-Kaas type method (Munthe-Kaas developed Lie-group exponential integrators) for 2 stages. The “stable” variant might use a coefficient set optimized for stability (e.g. L-stability or avoiding oscillations). <em>Order:</em> likely 2 for all, with small differences in internal coefficients (docs do not list details). *Behavior:* <em>res_2s</em> is essentially the simplest exponential RK: similar to <em>Heun’s method but in the exponential frame</em>. It costs 2 model calls per step. The stable variant may dampen any non-monotonic intermediate values (some 2nd-order integrators can overshoot; a “stable” one might have monotonic dissipative property). <em>Use:</em> As a starting exponential method if you want a bit more accuracy than Euler but minimal complexity. If <em>res_2m</em> (multistep) ever shows slight bias, try <em>res_2s</em> which might correct that with a true two-stage update. <em>Steps:</em> Not specified, but since it’s only 2nd order, you might need moderate steps (30–50) for very clean results. However, thanks to exponential handling, it might still outperform a plain 2-stage RK at large step sizes. <em>Schedulers:</em> No special pair noted. A <em>Karras or Normal</em> schedule would be fine. Possibly use <strong>“simple”</strong> if you explicitly want to test the integrator’s own stability (simple linear schedule plus stable integrator = very controlled). <em>Pitfalls:</em> None documented. Possibly “res_2s_non-stable” (the base version) might show slight non-monotonic behavior in some pixel intensities – if so, that’s exactly what <em>res_2s_stable</em> addresses. These differences would be subtle and mostly in edge cases (like extremely large guidance or tricky prompts).<br><br>- <strong>res_3s / res_3s_non-monotonic / res_3s_alt</strong> – <em>Class:</em> exponential explicit. <em>Idea:</em> 3-stage, 3rd-order exponential integrators. The base likely corresponds to a <strong>Lawson-Euler</strong> 3rd order or an ETDRK3 scheme. The “non-monotonic” label implies one variant allows slight overshoot for possibly higher accuracy (maybe it’s an optimized ETD3 that isn’t SSP). The “alt” might be an alternative coefficient set (e.g. one could be Cox–Matthews 3rd order, another a variant by Beylkin or others). <em>Order:</em> 3 (if standard). <em>Behavior:</em> 3-stage exponential integrators strike a good balance: e.g. <strong>Cox & Matthews (2002)</strong> introduced a popular ETD RK3. These methods can sometimes produce negative intermediate values (hence non-monotonic) but achieve better accuracy on oscillatory solutions. <em>Use:</em> When <em>2-stage isn’t enough</em>, but you still want to keep cost moderate. Good for fine textures or when using SDE mode – the higher order helps maintain detail when noise is injected. <em>Steps:</em> ~20–30 usually suffice at order 3. <em>Schedulers:</em> Try <strong>Karras</strong> for quality or <strong>SGM_uniform</strong> if you suspect the model’s native schedule (score-based) suits this integrator. <em>Pitfalls:</em> Not explicitly noted. The existence of “non-monotonic” suggests the default <em>res_3s</em> might sacrifice a bit of monotonicity for accuracy, which is normally fine for image generation. If you notice any unusual artifacts that diminish with more steps, consider switching to the alt or using the stable 2s or 4s.<br><br>- <strong>res_3s_cox_matthews / res_3s_lie</strong> – These are specific 3-stage integrators: <strong>Cox–Matthews</strong> is likely the ETDRK3 scheme from 2002, and <strong>Lie</strong> could refer to a Lie–Trotter splitting method (though Lie–Trotter is usually 1st order; perhaps here it’s a Lie-group RK of order 3, or even referencing Munthe-Kaas “Lie” method). <em>Order:</em> Cox–Matthews ETDRK3 is 3rd order; “Lie” might be a lower-order splitting used as comparison (maybe 3rd if it’s Strang splitting like Lie–Strang). <em>Behavior:</em> Cox–Matthews ETD3 is explicit but known to sometimes be unstable for very stiff problems (not an issue in diffusion usually). It’s a well-tested method for many PDEs. The “lie” method likely provides a different flavor, possibly simpler and more diffusive (since Lie splitting steps linear and nonlinear parts separately). <em>Use:</em> If you’re testing academic integrators or have a scenario where certain structures (like symmetries) are preserved by a Lie integrator. For most users, these won’t drastically differ from res_3s base. <em>Steps:</em> 20-ish. <em>Schedulers:</em> No special mention; stick to Normal or Karras. <em>Pitfalls:</em> None specific. Perhaps the Lie method might show slightly more blurring (splitting can be a bit less accurate per step than full integrator at same order, but more stable). Without official notes, treat them as alternatives to try if others fail.<br><br>- <strong>res_3s_strehmel_weiner</strong> – Strehmel & Weiner (1987) developed exponential integrators (their names appear in exponential integrator literature). This likely implements one of their methods, perhaps <em>ETD3 with stability enhancements</em>. <em>Order:</em> likely 3. <em>Behavior:</em> Possibly more stable than Cox–Matthews if it was designed with that in mind. Might involve computing matrix exponentials in a clever way (beyond our scope here). <em>Use:</em> Advanced usage; if you notice other 3s produce slight artifacts, maybe this one fixes it. Otherwise, differences are minor. <em>Pitfalls/notes:</em> not documented beyond the name drop, indicating experimental inclusion.<br><br>- <strong>res_4s_krogstad / res_4s_krogstad_alt</strong> – <em>Class:</em> exponential explicit. *Idea:* <strong>Krogstad’s 4th-order ETD</strong> (Krogstad 2005). Krogstad introduced a renowned 4th-order ETD RK (sometimes called ETD-RK4). The alt likely uses a modified formula (maybe a later improvement or a variation for better stability). <em>Order:</em> 4. <em>Stages:</em> 4. <em>Behavior:</em> High accuracy, but requires computing several φ-functions of the linear operator (which in diffusion context are handled analytically since linear part is simple $- \\sigma$). It’s more computationally heavy (4 model calls per step, plus overhead), but yields sharper results. One test found an <em>8-stage Minchev (see below)</em> and other 4+s methods among the top quality images, implying 4th+ order ETD methods shine in detail. <em>Use:</em> When you want <strong>very sharp edges and textures</strong> without going implicit. It excels at preserving fine details (a 4th order method has lower local error, meaning less detail lost each step). <em>Steps:</em> You can often get away with ~15–20 steps for decent results due to high order, but at cost of 4* that in model calls. For ultimate quality, 30 steps (120 evals) might surpass virtually any lower-order method’s 50 steps. *Schedulers:* <strong>Karras</strong> recommended – high-order integrators benefit from sigmoidal schedules because they handle large early steps well and then refine small steps well. A community result singled out <strong>res_4s_minchev</strong> (another 4th-order ETD) as producing “very good” quality at 95s vs others, implying any of these 4th order ETDs with a decent schedule will do great. Alt vs non-alt: you might try both if curious; differences would be subtle unless pushing near stability limits. <em>Pitfall:</em> Running a 4th-order method at too few steps (say 5-10) might result in <em>over-sharpening</em> or slight inaccuracies because each step assumes a smooth polynomial solution – if the model’s behavior is non-smooth between wide steps, error could manifest as minor artifacts or overshoot. Mitigate by using sufficient steps or adding a touch of noise (SDE mode) to regularize.<br><br>- <strong>res_4s_strehmel_weiner / alt</strong> – Strehmel & Weiner also proposed a 4th order scheme in late 80s. Similar story to Krogstad: an ETD4 integrator. The alt might be a variation or the “A/B” versions sometimes found in literature. <em>Order:</em> 4. <em>Use:</em> Same use case as Krogstad’s: high-accuracy needs. You might not notice much difference between Krogstad and Strehmel–Weiner results unless under extreme conditions; both target the same order. <em>No specific notes documented.</em> Try these if Krogstad’s yields any instabilities or just for thoroughness.<br><br>- <strong>res_4s_cox_matthews</strong> – Cox & Matthews also have a 4th-order ETD (ETDRK4, famous for reaction-diffusion equations). However, ETDRK4 in their paper is actually the classic scheme that’s widely used (and requires careful polynomial interpolation to avoid NaNs). It might be included here. <em>Order:</em> 4. <em>Behavior:</em> Very accurate but known to have potential NaN issues if not implemented with contour integrals (in PDE context). The RES4LYF implementation likely handles it fine for diffusion ODE. <em>Use:</em> Any scenario needing high precision. If multiple 4s are available, one might check which gives the best result for a given prompt (they should be similar). <em>Pitfalls:</em> If any ETD4 were to produce numerical overflow, it’d be Cox–Matthews’ (due to dividing by small exponentials in intermediate steps), but this is speculation – likely fine in this controlled setting, just something historically noted in papers.<br><br>- <strong>res_4s_cfree4</strong> – “cfree4” likely stands for <em>commutator-free</em> order-4 integrator. Commutator-free Lie integrators (by Celledoni, Owren, etc.) avoid calculating matrix commutators, making them efficient for certain high-dimensional problems. <em>Order:</em> 4. <em>Behavior:</em> These often preserve certain invariants or structures, and are stable. Possibly from Minchev & Munthe-Kaas (they described commutator-free schemes in a 2004 study). <em>Use:</em> Perhaps with <em>flow models or when combining transformations</em> (commutator-free might better handle dynamic conditioning?). For typical image generation, it likely behaves like any other 4th-order ETD. <em>No special notes</em>, included for completeness.<br><br>- <strong>res_4s_friedli</strong> – Friedli (1978) is credited with the first exponential RK methods. This could be an implementation of Friedli’s 4th-order method (he described one in his thesis). <em>Order:</em> possibly 4. <em>Behavior:</em> As a pioneering method, it may not be as optimized as later ones for stability, but it’s <em>exact for linear part</em>. Use if you’re curious historically. It will work as a 4th-order ETD integrator. <em>Pitfalls:</em> None noted, but earlier methods might have tighter stability constraints (so if at very large step sizes it diverges while Krogstad doesn’t, that’s why). Keep steps moderate.<br><br>- <strong>res_4s_minchev / res_4s_munthe-kaas</strong> – These reference <em>Minchev & Munthe-Kaas (2004)</em> who analyzed high-order exponential integrators. Possibly one is the method “ϕ4” from their work, and the other a variant or an 8-stage 4th-order method they recommended. <em>Order:</em> 4 (though some of their work also covers 6th order, but given “4s”, likely 4th order methods). <em>Behavior:</em> These might be fine-tuned for certain stability or efficiency aspects (Minchev & Munthe-Kaas looked at stiff order conditions). <em>Use:</em> They were highlighted in the comparative chart: <em>res_4s_minchev produced one of the top results (very good) while running in 95 seconds for 16 steps</em>. This suggests it’s both fast and high quality – anecdotally, a great choice for sharp outputs. It may introduce a tiny bit more noise or creativity than Krogstad’s (since different coefficient choices can slightly affect style). <em>Pitfalls:</em> None specific; if one is labeled “munthe-kaas” it might be more geared to Lie-group integration (relevant if the model had a Lie algebra structure, which here is not obvious – likely just naming credit). Both are safe picks for high-quality results.<br><br>- <strong>res_5s</strong> – <em>Class:</em> exponential explicit. <em>Idea:</em> A 5-stage, likely 5th-order integrator. Possibly a straightforward extension or one of Hochbruck–Ostermann’s methods. <em>Order:</em> If it’s consistent with naming, 5. But the docs don’t confirm, saying “most notably RES_5S implemented”. At least it’s higher than 4. <em>Behavior:</em> The doc highlights RES_5S as one of the <em>notable new explicit samplers</em>, implying it’s a star performer. It probably has very high local accuracy, meaning extremely sharp outputs if used right. However, being 5-stage, it’s heavy: 5 model calls/step. It might overfit the noise schedule a bit (5th-order can sometimes “assume” too much smoothness). <em>Use:</em> For <strong>absolute fidelity</strong> – e.g. photorealistic images where every micro-detail matters. Or when doing <em>very low denoise strength</em> (like 0.1 in img2img) where you need sampler accuracy to preserve details. <em>Steps:</em> You can try as low as 10–15 steps and often get away with it given the order (but that’s 50–75 model evals). Otherwise ~20 steps (100 evals) will be extremely clean. *Schedulers:* <strong>Karras</strong> strongly recommended – a high-order sampler paired with Karras’ smooth spacing is often ideal. If using SDE, maybe drop to 4s or 3s; 5s SDE might be overkill. <em>Pitfalls:</em> Diminishing returns beyond a point: you might not see much difference between 5s and 4s unless the scenario is complex or you push fewer steps. Also, 5th order methods can be less stable than 4th if pushing to very large step sizes or weird schedules (but with normal usage this is fine).<br><br>- <strong>res_5s_hochbruck-ostermann</strong> – Likely the 5th-order ETD RK by Hochbruck & Ostermann (they published a 5(4) pair in 2005). <em>Order:</em> 5. <em>Behavior:</em> Should be similar to res_5s if that wasn’t already HO’s method. Possibly one is the base RES_5S, and this explicitly credit HO’s coefficients (maybe slight differences). HO are experts in exponential integrators for parabolic PDEs, so their method might be particularly stable. <em>Use:</em> If using implicit guides or particularly stiff conditions, HO’s method could maintain stability. But in generation tasks, it should behave just as an excellent high-order sampler. <em>No special user guidance given beyond commitment to climate change if used with implicit refine</em> – translation: it’s very slow.<br><br>- <strong>res_6s, res_8s, res_10s, res_15s, res_16s</strong> – <em>Class:</em> exponential explicit. <em>Idea:</em> These are extremely high stage integrators, presumably aiming for very high order (perhaps up to 16th order for 16s!). They are not individually documented, implying they were added to experiment with how far one can push explicit solvers. <em>Order:</em> Not stated; presumably equals stages for the Gaussian quadrature integrators, but these are explicit – likely they use some schemes of those orders (maybe adapted from Butcher’s high-order RK or Taylor methods with automatic differentiation?). <em>Behavior:</em> Expect diminishing returns: as order increases, if the model’s true ODE isn’t smooth enough (due to neural network noise), super-high order doesn’t buy much. However, they will minimize error at any reasonable step size. <em>Use:</em> Largely experimental. You might use e.g. <code>res_8s</code> if generating at only ~8 steps (attempting “one step per order” ideal) – e.g. 8 steps with an 8th-order method might produce something workable where Euler would fail completely. If you’re trying crazy low-step generations (like 5–10 steps total for a 1024px image), these could be interesting. <em>Steps:</em> If using these at all, try a very low step count relative to normal (because if you do 50 steps with a 16th-order method, you’re really overshooting – numerical error will be negligible compared to model stochasticity). <em>Schedulers:</em> No data; presumably <strong>Simple</strong> or <strong>Linear</strong> might suffice because these integrators can handle it (maybe combine a linear schedule with a 16th-order method to basically solve the ODE near-exactly in few large jumps). <em>Pitfalls:</em> Enormous computation (16s = 16 model calls per step). Also, potential for numerical overflow if extreme (but likely fine here). The law of diminishing returns: beyond order ~8, the improvements are theoretical – the diffusion model and quantization noise introduce errors that numerical precision can’t overcome. So these are more for academic completeness.<br><br>- <strong>etdrk2_2s</strong> – <em>Class:</em> exponential (ETD). <em>Idea:</em> A specific 2nd-order ETD RK (probably ETD2, the simplest scheme also known as <em>exponential trapezoidal </em>or similar). <em>Order:</em> 2. <em>Behavior:</em> It’s like res_2s but likely a canonical implementation of Cox–Matthews ETDRK2. <em>Use:</em> If you want a well-tested stable 2nd order ETD. Not much to add beyond res_2s notes. <em>Steps:</em> 30+. <em>Pitfalls:</em> none beyond second-order limitations.<br><br>- <strong>etdrk3_a_3s / etdrk3_b_3s</strong> – Two variants of 3rd-order ETD RK. Possibly “A” and “B” from some paper (maybe related to different ways to compute intermediate matrix exponentials). <em>Order:</em> 3. <em>Use:</em> Same as res_3s: good general high-fidelity sampler. If one is labeled A vs B, one might prioritize stability over error or vice versa. <em>No doc guidance beyond names.</em> Trying both on a tough prompt to see if one yields fewer artifacts could be informative.<br><br>- <strong>etdrk4_4s / etdrk4_4s_alt</strong> – Cox & Matthews ETDRK4 and an alternative (perhaps the Kassam & Trefethen variant that addresses stability issues). ETDRK4 is 4th order and widely used, but one must carefully compute φ-functions to avoid numerical issues. The alt likely does just that. <em>Order:</em> 4. <em>Use:</em> High-quality sampling. <em>Pitfalls:</em> If any integrator here were to produce a NaN or diverge, it’d be ETDRK4 without stabilization at large step sizes. The alt presumably fixes it, so prefer alt unless you have reason not to. Again, this is a fine detail – in diffusion, matrix exponentials are easy (they’re just scalars e^{-λΔt}), so it might be moot and both work similarly. <br><br>- <strong>dpmpp_2s / dpmpp_sde_2s</strong> – <em>Class:</em> exponential explicit (though these could also be considered advanced single-step solvers). <em>Idea:</em> DPM++ 2S is the <strong>single-step</strong> second-order solver from Karras/LDMS (using a mid-point correction). DPM++ SDE 2S is the stochastic version designed for SDE (it adds noise in a midpoint manner). <em>Order:</em> 2 (for ODE version). <em>Behavior:</em> DPM++ 2S is known to be very accurate and tends to produce slightly smoother outputs than 2M since it recalculates mid-step rather than relying on previous steps. The SDE 2S variant will include randomness at each step like PLMS with noise, giving more diverse outputs. <em>Use:</em> These were many users’ go-to in Stable Diffusion 1.x era for good detail. Use DPMPP 2S when you want deterministic but high-quality results; use DPMPP SDE 2S if you want a dash of randomness for creative effects while still using a proper solver each step (less chaotic than full Euler SDE). <em>Steps:</em> ~20 is usually enough (the community often ran DPM++ 2S a Karras 20–30 steps for great images). *Schedulers:* <strong>Karras</strong> is practically assumed for DPM++ in many references – it was tuned for that. Both ODE and SDE versions benefit from Karras or at least some sigmoidal schedule. <em>Pitfalls:</em> The SDE version, if run for too many steps (say >100), won’t converge – it will keep injecting noise (“non-deterministic samples don’t converge at high steps”), so don’t overshoot steps expecting a stable output. The ODE version will converge but if you push steps very high, watch out for floating point precision issues (generally not a problem in 32-bit until thousands of steps).<br><br>- <strong>dpmpp_3s</strong> – <em>Class:</em> exponential explicit (or just explicit multi-stage). <em>Idea:</em> Possibly the 3rd-order single-step variant (if it exists – DPM-Solver++ had order 3 solver as well). <em>Order:</em> 3. <em>Behavior:</em> A bit more accurate than 2S at cost of extra eval. <em>Use:</em> Rarely needed; 2S was usually the sweet spot. But if you find 2S slightly lacking, 3S is here. <em>Steps:</em> 15–25. <em>Pitfalls:</em> None known.<br><br>- <strong>lawson2a_2s / lawson2b_2s</strong> – <em>Class:</em> exponential explicit. <em>Idea:</em> Two different 2-stage Lawson methods. Lawson (1967) introduced the idea of using $e^{At}$ factor for linear part; these might refer to two specific sets of coefficients from his paper or subsequent work labeled “Method 2a” and “2b”. <em>Order:</em> 2. <em>Behavior:</em> They likely differ in how they weight the stages (one might minimize error constant, another might be L-stable in linear case). <em>Use:</em> If you specifically want a stable 2nd order method that exactly integrates the linear part. For instance, if using an older model with known linear behavior, Lawson’s might do well. In practice, similar to etdrk2. <em>No strong recommendations in docs.</em> <br>- <strong>lawson4_4s</strong> – A 4-stage Lawson method, presumably 4th order. Possibly Lawson’s original 4th order scheme or a later development. <em>Order:</em> 4. <em>Use:</em> Similar role as ETDRK4. Use if you want an exponential integrator but maybe fewer φ-function computations (Lawson methods often turn the problem into needing standard RK on the nonlinear transformed system). <br>- <strong>lawson41-gen_4s / lawson41-gen-mod_4s</strong> – These cryptic names suggest a <em>generalized Lawson 4(1) method</em>, perhaps an embedded pair (4th order with 1st order error estimate) or generation method. The “mod” could be a modified version. <em>Use:</em> Possibly for steps adaptivity or error tracking (though adaptivity isn’t exposed in UI). Without adaptivity, they’d just behave as 4th-order. They might have been included for future-proofing. <br>- <strong>ddim</strong> – Actually listed under explicit samplers (as “ddim” with no stages). It’s a special case: <em>Denoising Diffusion Implicit Model</em> sampler by Song et al. (2020), which can be seen as a first-order ODE integrator that exactly matches the diffusion process when steps = original count. It’s not a classical RK but fits here as an explicit method. <em>Order:</em> effectively 1 (it’s like Euler method applied to a non-linear schedule, though it can be surprisingly accurate). <em>Behavior:</em> Deterministic, tends to give smoother, lower-contrast outputs compared to PLMS or DPM++. It’s sometimes used for style or when simulating how diffusion progresses. <em>Use:</em> When you want to replicate a known reference (like to produce the same result as a certain DDIM step count from another tool). Or if you desire its characteristic softer look. <em>Steps:</em> Historically people use 50 DDIM steps to mimic 100 DDPM steps, etc. You can start with 20–50. *Schedulers:* <strong>DDIM_uniform</strong> is logically the match (the sampler ignores the sigma input somewhat because it has its internal formula, but RES4LYF’s scheduler will influence it). Using DDIM with the “ddim_uniform” schedule best replicates standard DDIM. <em>Pitfalls:</em> Because it doesn’t correct error (no PC, no higher order), it may lose fidelity at low steps (under 20, e.g., can be muddy). Also, it lacks the ability to recover from mistakes – each step directly maps a fractional time step to latent, so any error carries forward.<br><br>- <strong>euler</strong> – Also listed under explicit. The simplest sampler: 1 model eval per step, first-order. In KSampler this is “Euler” (ancestral variant adds noise). In RES4LYF context, Euler is the probability flow ODE Euler solver (deterministic). <em>Order:</em> 1. <em>Behavior:</em> Fast but low accuracy. Tends to oversmooth or under-shoot details unless many steps are used. <em>Use:</em> For quick previews or if you intentionally want a simpler, possibly more varied result (with ancestral noise, Euler a is popular for its creativity – but in ODE mode here, it will be deterministic, so creativity only via prompt). <em>Steps:</em> ~30–100 depending on needed quality. It’s linear convergence, so doubling steps roughly halves error. *Schedulers:* <strong>Exponential</strong> works great for Euler – it covers the big stuff quickly and Euler doesn’t do well with many small steps at high noise anyway. That combination was the fastest in one test (16 steps in 24.5s). Euler with Karras is also used to refine quality at high step counts (Karras schedule reduces Euler’s error in later stages by making steps very small when they matter most). <em>Pitfalls:</em> Very low-step Euler (like <10) often yields bad outputs or very unrealistic ones. It lacks the corrections higher-order methods have, so structure may be less coherent (could manifest as repetitive detail artifacts or wrong anatomy at moderate steps). <br><br>### Hybrid Samplers (Predictor–Corrector / mixed methods)<br><br>**Integration idea:** Hybrids combine multiple methods – often an explicit predictor and an exponential or implicit corrector. They leverage strengths of both: explicit predictor for speed, then a partial correction to improve accuracy or stability. Notably, “PEC” means <strong>Predict, Evaluate, Correct</strong>, often referencing Norsett’s predictor-correctors.<br><br>- <strong>pec423_2h2s / pec433_2h3s</strong> – <em>Class:</em> hybrid (probably Predictor–Explicit–Corrector). The notation suggests something like: 4-2-3 stands for a predictor of order 4, corrector of order 3, etc. Actually, <strong>PEC423</strong> might mean a 2-step predictor, 4th order predictor, 2nd order corrector, 3rd overall? It’s not entirely clear, but likely these implement a predictor followed by a corrector iteration using half the stage count “h”. For example, pec423_2h2s might do a 2-stage predictor, then another 2-stage corrector (so hybrid 2+2 = effectively 4 stages?). <em>Behavior:</em> Hybrids of this sort can improve stability without going fully implicit. <em>Use:</em> Possibly designed for cases where explicit alone was slightly insufficient. They might allow a larger step size or better detail with fewer overall model calls by reusing evaluations. <em>No direct documentation</em>; treat as advanced if others fail. Use a moderate scheduler (Normal or linear) so that the predictor-corrector interplay is well-behaved. <br>- <strong>abnorsett2_1h2s, abnorsett3_2h2s, abnorsett4_3h2s</strong> – These appear to embed AbNorsett multistep methods into hybrid forms. E.g., “abnorsett2_1h2s” might mean an AbNorsett 2-step method with 1 hybrid iteration using a 2-stage method. Possibly a predictor (AbNorsett) then one correction via 2-stage implicit or explicit. It’s highly specific and not detailed in docs. <em>Use:</em> Only if you deliberately want to experiment with multi-step + small correction. <br>- <strong>lawson42-gen-mod_1h4s</strong> through <strong>lawson45-gen-mod_4h4s</strong> – These look like a series of Lawson methods combined with something (“gen-mod” again). The prefix suggests maybe an embedded family Lawson4.2, 4.3, 4.4, 4.5 which are then used in hybrid form (h?). The numbers before h (1,2,3,4h4s) likely mean how many explicit half-steps or something. Honestly, these are not documented and likely experimental. <em>Use:</em> They might aim to achieve very high order or special stability by combining multiple 4-stage integrator passes in one step. Unless you’re an integration researcher, you can probably skip these; or if you are, you might consult the code for how they’re constructed. (In practice: run some tests with a fixed seed across these to see differences. Without official notes, one cannot say which scenarios they benefit.)<br><br>In summary, <strong>Hybrid samplers</strong> in RES4LYF are niche. The typical user will rarely need to select these manually – they exist perhaps to support internal “chains” or as stepping stones to fully implicit methods. If you do use them, note that they internally disable ComfyUI’s noise addition in certain modes (unsampling/resampling contexts) to manage the two-phase integration properly. Always keep <em>SAMPLER_MODE = standard</em> unless you specifically are doing an unsampling workflow.<br><br>### Linear (Explicit ERK) Samplers<br><br>**Integration idea:** These are standard Runge-Kutta methods without exponential factors. They treat the diffusion ODE as is, explicitly. They are included partly to leverage known properties (like symplecticness, SSP, etc.) and partly for completeness (one might discover a classical RK yields a particular aesthetic). All are ODE (non-SDE) samplers by default, though you can turn on SDE mode to add noise post-step.<br><br>- <strong>ralston_2s / 3s / 4s</strong> – <em>Class:</em> explicit RK. <em>Idea:</em> Ralston’s methods – optimized RK with minimum error bounds for given stages. 2s is Ralston’s 2-stage second-order method. 3s is Ralston’s 3-stage third-order. 4s is Ralston’s 4-stage fourth-order method. <em>Behavior:</em> These prioritize even error distribution. They tend to slightly undershoot overshoot, producing very stable integration for their order (hence minimum error constants). For image generation, that means reliable, if maybe a touch less “sharp” than some other RK of same order (because they avoid aggressive slopes that could overshoot). <em>Use:</em> Great for <strong>stable, artifact-free results</strong> if you match steps appropriately. For instance, if you want a well-behaved 3rd-order sampler, Ralston 3s is a good pick. <em>Steps:</em> ~20 for 2s, ~15-20 for 3s, ~10-15 for 4s (because higher order needs fewer steps). *Schedulers:* <strong>Normal</strong> suits Ralston’s philosophy (uniform improvement each step). Also Beta or linear_quadratic might pair nicely to further minimize any erratic change. <em>Pitfalls:</em> None documented. Ralston’s methods are not specialized for stiffness, so if you push step count very low, they’ll fail like any explicit RK. But within normal use, they’re quite robust.<br><br>- <strong>midpoint_2s</strong> – <em>Class:</em> explicit RK. <em>Idea:</em> The explicit midpoint method (also known as RK2 or “modified Euler”). <em>Order:</em> 2 (and symplectic). It’s the simplest Gauss–Legendre collocation as Wikipedia notes. <em>Behavior:</em> Slightly better than Euler; tends to damp less than Heun’s method. It’s actually symplectic (though that’s more relevant for conservative systems). In diffusion, it’s just a standard second-order that might preserve some geometric structure. <em>Use:</em> If Euler was too crude but you want minimal extra cost. It might give a bit more contrast or structure preservation than Heun because it’s the implicit midpoint’s explicit twin (which is symplectic, though explicit midpoint itself isn’t symplectic, implicit is – so scratch that, explicit midpoint is just another term for modified Euler). <em>Steps:</em> ~30. <em>Pitfalls:</em> None unique; it’s just an ordinary method seldom uniquely best.<br><br>- <strong>heun_2s / 3s</strong> – <em>Class:</em> explicit RK. <em>Idea:</em> Heun’s method typically refers to the explicit trapezoidal rule (2-stage, 2nd order) – that’s heun_2s. Heun’s 3s likely refers to a less-common 3-stage method by Heun or in his style (maybe an older 3rd order method). <em>Behavior:</em> Heun 2s (also called “improved Euler”) is a bit more diffusive than midpoint: it averages the slope, which often yields smoother, slightly less sharp results than midpoint method. But it is good for avoiding overshoot. The 3-stage Heun might be akin to Kutta’s 3rd order or some variant (the exact method is unclear from name, possibly identical to Ralston’s 3rd since Heun did propose some higher order in early 1900s). <em>Use:</em> For <strong>stable progression with fewer ringing artifacts</strong>. Good if your images had some repetitive noise with other methods – Heun’s averaging might mitigate that. <em>Steps:</em> ~30 for 2s, ~20 for 3s. <em>Schedulers:</em> Normal or Karras. (Heun’s doesn’t have special scheduling needs.) <em>Pitfalls:</em> Slightly more blurring relative to midpoint or Ralston (because it damps changes more). Not a big difference at proper steps though.<br><br>- <strong>houwen-wray_3s</strong> – <em>Class:</em> explicit RK. <em>Idea:</em> Van der Houwen & Wray’s 3-stage, 3rd-order method. They likely designed it with some property (like FSAL or minimized error). <em>Behavior:</em> If it’s the one I suspect, it might be the <em>same</em> as Ralston 3rd or similar. Possibly they introduced an alternative third-order with maybe lower error in some scenario. <em>Use:</em> Hard to differentiate usage without specifics – treat as another option if other 3-stage methods produce subtle differences. <em>No pitfalls known.</em> <br><br>- <strong>kutta_3s</strong> – <em>Class:</em> explicit RK. <em>Idea:</em> Kutta’s original 3-stage 3rd-order method from 1901[1]. <em>Behavior:</em> It’s a historically important method but has a larger error constant than Ralston’s 3-stage. It might overshoot slightly more (leading to a bit more contrast?). <em>Use:</em> Nostalgic or experimental; probably not outperforming modern optimized ones. <em>Pitfalls:</em> none special, it’s a solid method.<br><br>- <strong>ssprk3_3s</strong> – <em>Class:</em> explicit RK. <em>Idea:</em> The <strong>Shu–Osher 3-stage SSP</strong> method (a.k.a. RK3 with TVD properties). <em>Order:</em> 3. <em>Behavior:</em> It guarantees no new extrema (for convex combinations) if time-step is within certain limits. For diffusion, this likely translates to <em>no overshoot in pixel intensities beyond what’s expected</em>, giving very stable images. It’s known to be the optimal 3-stage 3rd-order SSP method. <em>Use:</em> If you encountered weird artifacts like stripe patterns or pixel ringing with other samplers, SSPRK3 can help avoid that due to its strong stability preserving nature. It might produce a bit “flatter” image (less micro-contrast) but very orderly. <em>Steps:</em> ~20–30. *Schedulers:* <strong>Normal</strong> is a good match (monotonic decay). <em>Pitfalls:</em> Only that it might be slightly more conservative (could hide some texture by smoothing it out, because it avoids overshoot). But results will be consistent.<br><br>- <strong>ssprk4_4s</strong> – <em>Class:</em> explicit RK. <em>Idea:</em> A 4-stage attempt at SSP 4th order. Note: truly SSP 4th-order explicit RK requires at least 8 stages (Shu proved no 4-stage can be SSP of order 4 except trivial zero step-size). Possibly this method is one that maximizes SSP coefficient but isn’t fully SSP for all steps. In practice, it might be just a decent 4th-order method that tries to avoid overshoot. <em>Behavior:</em> Should be high accuracy and fairly stable, but not absolutely monotonic like the 3-stage case. <em>Use:</em> If you want 4th-order but worry about artifacts, try it. <em>Pitfalls:</em> If expecting full SSP, note it’s not guaranteed beyond a certain threshold. But likely not an issue in our context.<br><br>- <strong>rk38_4s</strong> – <em>Class:</em> explicit RK. <em>Idea:</em> The “3/8-rule” Runge-Kutta, which is an alternative 4th-order method to the classic RK4. It has a different distribution of error (slightly better phase properties for oscillatory problems). <em>Behavior:</em> Both RK4 and RK3/8 are 4-stage order 4, but RK3/8 tends to spread error more evenly. It might produce marginally different image details – one might give a bit more contrast, the other a bit more smoothness. The differences are subtle. <em>Use:</em> You can switch between <code>rk4_4s</code> (classic) and <code>rk38_4s</code> if you’re chasing tiny differences in output style; otherwise, treat them as the same class. <em>No strong use guidance in docs, just included.</em> <br>- <strong>rk4_4s</strong> – <em>Class:</em> explicit RK. <em>Idea:</em> The <em>classic Runge-Kutta</em> (a.k.a. RK4 or Kutta’s 4th-order from 1901). <em>Order:</em> 4. <em>Behavior:</em> The gold standard of generic integrators – good accuracy, stability not bad but not L-stable. It tends to produce crisp results when enough steps are given. Indeed, in the sampler comparison, a 7-stage variant of RK (Dormand-Prince) gave best quality; RK4 is a step below that but still excellent. <em>Use:</em> Good general-purpose if you don’t want to fuss with exponential stuff. It’s likely to yield results similar to DPM++ 2M at 2x steps, etc. <em>Steps:</em> ~15–25. <em>Schedulers:</em> Karras recommended for fewer steps (since RK4 can handle larger steps early and refine later with Karras’s small end steps). <em>Pitfalls:</em> Slightly less stable at very large step sizes – e.g. if you try 5 steps with RK4 it might overshoot more than an implicit method would. But at reasonable steps, no problem.<br><br>- <strong>rk5_7s</strong> – <em>Class:</em> explicit RK. <em>Idea:</em> Possibly the <em>Nyström 5th-order 7-stage</em> or something (though Nyström is for second-order eqns). More likely it is <em>Butcher’s 7-stage 5th-order</em> or some other classic 5th-order method (maybe Kutta’s 5th or Carver’s). Without doc, we assume a standard 5th-order one. <em>Behavior:</em> Very accurate, but 7 stages. Unless you need the error reduction, Dormand-Prince or Tsitouras pairs might be more effective (since they give error estimate). <em>Use:</em> High accuracy explicit without adaptivity. If you want slightly better than RK4 and don’t mind extra calls. <br>- <strong>rk6_7s</strong> – <em>Class:</em> explicit RK. <em>Idea:</em> Possibly the <em>Chawla 6th-order 7-stage</em> or similar method (there are known RK6(7) designs). The comparison chart named “Linear rk6_7s” as <em>best quality and prompt adherence</em> at the cost of 167 seconds. That indicates this is an extremely good integrator (6th order) that, given enough compute, nails the details. <em>Behavior:</em> The image came out best-of-group in that test – meaning extremely sharp and matching prompt, but it took 167s (7 stages, 16 steps = 112 model calls) so not surprising. <em>Use:</em> If you want the absolute best result and can wait ~5-7x longer than a normal sampler. Good for final renders or when upscaling with minimal denoise. <em>Schedulers:</em> They used Normal in that test, interestingly – implying even linear spacing with a high-order method gave top quality. Karras might push it even further, but either should be excellent. <em>Pitfalls:</em> Just the huge time. Nothing weird expected – high order explicit is fine if you give it the steps to chew on.<br><br>- <strong>bogacki-shampine_4s / 7s</strong> – <em>Class:</em> explicit RK. <em>Idea:</em> Bogacki–Shampine is known for an embedded 3(2) method (3-stage 3rd order with an extra output for 2nd order error). But here “4s” and “7s” suggests possibly the <em>embedded 5(4) method by Bogacki & Shampine</em>, which uses 7 stages (they did publish a 7-stage 5th-order pair as well). The 4s might be the simpler BS3 method repurposed or another method by them. <em>Behavior:</em> The 7s likely refers to a 5th-order method with an embedded 4th for adaptivity (like BS5(4)). If adaptivity isn’t used, it’s effectively a very good 5th-order method. <em>Use:</em> If you want Dormand–Prince-like performance but maybe a different error distribution. Possibly yields similar quality as RK6_7s or DP. <em>Steps:</em> ~10–20. <br>- <strong>dormand-prince_6s / 13s</strong> – <em>Class:</em> explicit RK. <em>Idea:</em> The famous Dormand–Prince pairs. “6s” likely is the DP5(4) – often described as a 7-stage method but with FSAL it requires 6 distinct stages[2] (the last evaluation reuses the first’s result). Indeed DP 5th order has 7 tableau rows but effectively 6 function evaluations. The “13s” is likely the DP8(7) pair (13 stages) offering 8th order accuracy. <em>Behavior:</em> Dormand–Prince 5th order (the one behind <code>ode45</code>) is very efficient and produces very accurate results per computation – ideal for smooth problems. Here, one user found DP6s (which they likely ran as <code>dormand-prince_6s</code>) gave the absolute best quality among dozens of samplers at 30–40 steps (particularly a variant was singled out as “best quality”). So DP5 is a top-tier method for image diffusion as well. The 13s (8th order) is even more precise but with massive overhead. <em>Use:</em> For ultimate quality. The DP5(4) method will reach high quality in fewer steps than RK4. The DP8(7) (if indeed that) might be overkill, but you could generate an image in maybe 8 steps with basically no numerical error – any flaws are purely model/semantic. <em>Steps:</em> DP5(4): ~15–20. DP8(7): you could try as low as 8–10 steps and likely be fine due to 8th order, though the model’s own uncertainty might then dominate. <em>Schedulers:</em> DP is designed to work with an adaptive controller usually. Without that, probably <strong>Normal or linear</strong> (to not confuse it with drastically uneven step sizes). In the test, a linear schedule with DP gave superb results. Karras might be okay too but adaptivity is where DP shines, which we aren’t using here, so any reasonable schedule works. <em>Pitfalls:</em> Long runtime for the 13s. Also, extremely high order can amplify floating-point issues if you push steps too far – but with 8th order you wouldn’t need many steps anyway. Practically, nothing concerning beyond speed.<br><br>- <strong>tsi_7s</strong> – <em>Class:</em> explicit RK. *Idea:* <strong>Tsitouras 5(4) method</strong> with 7 stages and FSAL. Tsitouras (2011) created a more optimal 5th-order RK than Dormand–Prince, often yielding smaller error. <em>Behavior:</em> It’s an improved ode45 – usually about 20-50% error reduction at same step size, or can take slightly larger steps for same error. In diffusion terms, it should produce equally good or slightly sharper images compared to DP5 for the same number of steps. It also tends to be more stable for certain tricky ODEs. <em>Use:</em> A great choice if you want high-order but maybe shave off a few steps vs Dormand–Prince. <em>Steps:</em> ~15 (Tsit5 can achieve what DP5 might in ~18-20 steps typically, though this is heuristic). <em>Schedulers:</em> Normal or Karras; Tsitouras method is typically used adaptively, but fixed-step is fine. <em>Pitfalls:</em> None; Tsitouras designed it to avoid some pitfalls of Dormand–Prince (like better stability in some cases). So it’s arguably one of the best fixed 5th-order methods to use here.<br><br>- <strong>euler (again)</strong> – Already covered in explicit category above.<br><br>*(Glossary note: Many of these linear RK methods (Ralston, Kutta, Bogacki-Shampine, Dormand-Prince, etc.) come from classical ODE development. See the <strong>Glossary</strong> at the end for brief definitions and significance of each.)*<br><br>### Diagonally Implicit Samplers (DIRK methods)<br><br>**Integration idea:** Diagonally implicit RK have an implicit formula but only on the diagonal of the Butcher tableau, so each stage can be solved sequentially (not fully coupled as in Gauss-type). This makes them easier to solve (each stage is like a single implicit equation often linear in diffusion case), and they can handle <em>stiffer</em> equations than explicit methods. They usually are A-stable (stable for any step size in linear systems) and some are even L-stable (damp out high-frequency modes strongly). They cost more per step (each stage requires solving an implicit equation – in diffusion ODE context, that’s actually straightforward, maybe solved analytically; in practice likely they iterate or just apply formula since our ODE is simple).<br><br>- <strong>irk_exp_diag_2s</strong> – <em>Class:</em> diagonally implicit (kind of hybrid). <em>Idea:</em> Described as “features an exponential integrator” in the README. Possibly an implicit RK that also uses an exponential factor. It might be a 2-stage DIRK with exponential coefficient – maybe meaning it solves linear part implicitly (which would just be exact anyway) and handles non-linear part explicitly? Without clarity, treat it as an *advanced stable 2-stage method combining implicit and exponential techniques.* <em>Order:</em> Likely 2. <em>Behavior:</em> Very stable and forgiving on step size. You could probably run fewer steps without blow-up. <em>Use:</em> If explicit 2-stage gives weirdness, try this; it will be slower but rock-solid.<br>- <strong>kraaijevanger_spijker_2s</strong> – <em>Class:</em> DIRK. <em>Idea:</em> A 2-stage DIRK from Kraaijevanger & Spijker (early 1990s), likely focusing on large stability region or SSP. The Butcher tableau is given in Wikipedia. <em>Order:</em> Possibly 2 (the snippet didn’t state, but likely second-order). <em>Stability:</em> They probably optimized for <em>A-stability with some SSP property</em>. Indeed, the text around it hints at positivity conditions (though those lines apply to P&R method). <em>Behavior:</em> Should be unconditionally stable (A-stable) and maybe strongly damping. <em>Use:</em> For scenarios where explicit methods show oscillation or overshoot (like bright/dark edges overshooting). It will march firmly towards solution without ringing. <em>Steps:</em> Because it’s implicit and stable, you can potentially use larger step sizes – maybe ~10–15 steps for a decent result where explicit needed 20. But note each “step” has 2 implicit solves (though likely cheap here). <em>Schedulers:</em> Could even try <strong>Linear</strong> or <strong>Beta</strong> – implicit methods can handle a linear schedule without trouble because they are A-stable, meaning even a big initial step won’t destabilize. <em>Pitfalls:</em> Implicit methods are slower (in actual compute); also sometimes they overly damp fine details (L-stability can erase tiny features along with noise). Use them when stability is more important than preserving every micro-detail.<br>- <strong>qin_zhang_2s</strong> – <em>Class:</em> DIRK. <em>Idea:</em> Qin & Zhang’s 2-stage, 2nd-order DIRK. It’s noted as <em>symplectic</em> in Wikipedia (the table snippet calls it symplectic DIRK). <em>Order:</em> 2. <em>Stability:</em> symplectic means it preserves volume in phase space – in diffusion that doesn’t directly apply, but it suggests it’s not overly damping (contrasted with L-stable methods). <em>Behavior:</em> Likely less aggressive damping of high frequencies, so it might preserve texture better than an L-stable method. <em>Use:</em> If you want some stability benefits of implicit, but still keep some liveliness. <em>Steps:</em> ~15–20. <em>Pitfalls:</em> None documented. It’s a safe method. Possibly not L-stable (since symplectic usually implies some oscillations are preserved), so not as good if you need maximum damping.<br>- <strong>pareschi_russo_2s / alt_2s</strong> – <em>Class:</em> DIRK. <em>Idea:</em> Pareschi & Russo’s 2-stage method (2000s) focusing on stiff source terms (they work on kinetic equations). Wikipedia gives a family with a parameter x; stability requires x≥1/4 for A-stability, and a particular x for L-stability. The “alt” probably picks the other root of that polynomial for L-stability. So one is likely L-stable, the other maybe only A-stable but perhaps more accurate. <em>Order:</em> 2. <em>Behavior:</em> The alt with x being one of the roots (there were two possible x that yield L-stability in text) means the alt is L-stable (maximally damping). The other might choose x=1/2 (for example) to maximize stability interval but not fully L-stable. <em>Use:</em> If you want absolutely no ghosting or ringing, use the L-stable variant (maybe alt). If you want a bit less damping, use the base. They’re good for stiff situations, e.g. when using huge guidance or very abrupt noise reduction where explicit blows up. <em>Steps:</em> Could do fewer steps thanks to L-stability. <em>Pitfalls:</em> L-stable methods can sometimes overly blunt details, as mentioned. But if you plan to add a little noise (SDE) or do a refinement pass after, that’s fine.<br><br>- <strong>crouzeix_2s / 3s / 3s_alt</strong> – <em>Class:</em> DIRK. <em>Idea:</em> Michel Crouzeix’s 2-stage 3rd-order DIRK[3] and 3-stage 4th-order DIRK (with a parameter α in the 3-stage; alt likely uses the alternative α solution). <em>Order:</em> 2s is actually 3rd order (surprisingly high order for 2 stages, but indeed Crouzeix found a 3rd-order DIRK)[3]. 3s is 4th order. <em>Stability:</em> Probably A-stable (Crouzeix was big on stability of RK). These might not be L-stable though. <em>Behavior:</em> High order implicit with fewer stages: 2-stage 3rd order is economical. It might be a great choice if you want a bit more accuracy than Radau IIA 2-stage (which is order 3 too) but simpler? The 3-stage 4th order can possibly compete with explicit 4th order but with better stability. <em>Use:</em> Very solid all-rounders for those who aren’t shy about implicit. For example, if a prompt tends to create slight color ringing, crouzeix methods will suppress that but still achieve high accuracy. <em>Steps:</em> 2s (3rd order) can do with ~15 steps; 3s (4th order) maybe ~10–15. <em>Pitfalls:</em> Not much – Crouzeix’s methods are well-behaved. The alt likely corresponds to a different α in the 4th order family (one root might give slightly better stability at the cost of something else). Without detail, you can test which yields a nicer image; differences likely minor.<br><br>### Fully Implicit Samplers (Gauss, Radau, Lobatto collocation)<br><br>**Integration idea:** Fully implicit Runge-Kutta (collocation methods) treat all stages as coupled. They often require solving a system of equations at each step (which is computationally heavy, but in our context with a 1D time ODE, it’s more feasible – essentially it might iterate a fixed-point or do a small Newton iteration per step). These methods have superior accuracy for a given number of stages (Gauss-Legendre attains order 2s) and excellent stability (often A-stable or L-stable). They are the “Rolls-Royce” of integrators – expensive but high-performing.<br><br>- <strong>gauss-legendre_2s, 3s, 4s, 5s (+ variants)</strong> – <em>Class:</em> fully implicit (collocation). <em>Idea:</em> Gauss–Legendre collocation at Gaussian nodes yields symplectic, A-stable integrators of order 2s. 2s is order 4, 3s order 6, 4s order 8, 5s order 10. The variants “4s_alternating_a”, “4s_ascending_a”, “4s_alt” etc., likely refer to different ways to solve or different root ordering of polynomial (not affecting results) or might correspond to different formulations but same mathematics. Possibly some addressing of the fact that multiple equivalent coefficient sets exist due to symmetry. <em>Behavior:</em> These are extremely accurate and stable. Being symplectic means they won’t artificially damp or grow energy – in diffusion, they just accurately trace the probability flow ODE with minimal numerical diffusion or dissipation. <em>Use:</em> When <strong>maximum accuracy per stage</strong> is needed. Note the doc joke: pairing res_5s predictor with gauss-legendre_5s corrector for “ultimate image quality”. That underscores GL’s prowess (but also its slowness). Even standalone, gauss-legendre_3s or 4s will produce excellent results; you might not see much difference from an explicit 4s except at lower step counts or tricky prompts, but it’s there. <em>Steps:</em> You can push to very low steps: e.g. 3s (6th order) might do fine in 10 steps (like an RK6 would), but with more stability. Or use normal step counts and expect slightly better detail. *Schedulers:* <strong>Any</strong> – Gauss methods are A-stable for all s, so even aggressive schedules won’t break them. If using very few steps, maybe avoid a schedule that starves the initial noise too quickly (though A-stability means they can handle it). Normal or Karras are fine. <em>Pitfalls:</em> The solve: each step requires solving s nonlinear equations. With our neural network, presumably they do a few functional iterations. This might increase latency per step, especially at high s. Also, while symplectic, they are not L-stable – they don’t aggressively damp. That means if there are very high-frequency components (which in an image might be tiny noise), they won’t kill them as quickly as Radau IIA would. Sometimes that’s good (preserve texture), sometimes not (maybe residual noise remains). In practice, the difference is minor and can be controlled by the scheduler or a tiny denoise at end.<br><br>- <strong>radau_ia_2s, 3s; radau_iia_2s, 3s, 5s, 7s, 9s, 11s (+ alt)</strong> – <em>Class:</em> fully implicit (Radau collocation). <em>Idea:</em> Radau collocation points including one end of interval. Radau IIA includes the right end, yielding L-stable methods of order 2s-1. Radau IA includes the left end, also order 2s-1, but typically not L-stable (only A-stable). <em>Order:</em> 2s-1 (Radau IIA 3s = 5th order, 5s = 9th order, etc.). <em>Stability:</em> Radau IIA are L-stable (great for stiff problems). Radau IA share stability function but not L-stable due to different error propagation. The alt might be an alternate coefficient set (maybe swapping some stage order, since the collocation polynomial might not be unique?). Possibly alt uses a different root ordering or addresses a numerical issue (like ill-conditioning for high s). <em>Behavior:</em> L-stability means these methods <em>annihilate</em> any high-frequency error by the next step. So images come out very clean – no lingering noise or tiny wobbles. They also handle abrupt changes in dynamics well. Their accuracy is slightly lower than Gauss (order 2s-1 vs 2s), but still very high. E.g., radau_iia_3s is 5th order (compared to Gauss 3s 6th order), but radau will damp more. <em>Use:</em> Excellent for tricky cases like <em>extremely high CFG scales or weird conditioning</em> where other samplers freak out – Radau will plow through stably. Also good for final-stage img2img where you want to ensure no residual noise. <em>Steps:</em> Because of L-stability, you can even use fewer steps than an explicit method of similar order – e.g. radau_iia_3s (order 5) might perform well at 10–15 steps where an explicit 5th order might still leave slight noise. For super accuracy, use as many steps as needed to reduce error (they’ll do it monotically). *Schedulers:* <strong>Any</strong>, but pairing with <em>exponential or tangent schedules</em> might be redundant – Radau already handles things; a simple schedule is fine. However, in practice Karras or normal both work. One user ran radau_iia_7s (order 13) with normal schedule and got “very good” results (very sharp and consistent) albeit slow. <em>Pitfalls:</em> The heavy computational cost (solving 7 or 9 simultaneous eq per step for high s). And the possibility of <em>over-damping</em>: if you rely on a bit of randomness or noise, Radau might remove it too effectively, leading to images that are perhaps too smooth or a tad less detailed than an explicit counterpart. The fix could be to add SDE noise or use a slightly less damped method for final few steps. But as a default, it’s hard to fault Radau IIA for stability-critical tasks.<br><br>- <strong>lobatto_iiia/b/c/d/star (2s, 3s, 4s)</strong> – <em>Class:</em> fully implicit (Lobatto collocation). <em>Idea:</em> Lobatto methods include both interval endpoints in collocation. They typically have order 2s-2 (slightly less than Radau/Gauss). There are several variants (IIIA, IIIB, IIIC, IIIC*, IIID) depending on how collocation conditions are set (they differ in stability/structure). Many Lobatto methods are symmetric (good for reversible systems) and some are <em>embedded</em> (IIIA & IIIB are dual that can be combined for higher order or error estimation). E.g., Lobatto IIIC* might refer to an variant with certain simplifications. <em>Order:</em> For instance, Lobatto IIIA 2s is order 2, 3s order 4, 4s order 6; similarly for IIIB, IIIC. They are usually A-stable but not L-stable (some are only A-stable for limited step sizes). They are often <strong>symplectic</strong> and used in conservative systems. <em>Behavior:</em> In diffusion context, symplecticness isn’t needed (system isn’t conservative). So Lobatto might not offer special advantages except being implicit and symmetric. Symmetry can sometimes mean better error distribution and perhaps slightly less bias – maybe preserving mean color or something? Not documented though. <em>Use:</em> Not a go-to for diffusion, but included. If you want to experiment, you might try combining a Lobatto IIIA and IIIB as predictor-corrector (since they are dual). But that’s beyond normal usage. <em>Pitfalls:</em> Lower order than Gauss for same stages and no L-stability, so they may be slightly worse than Radau or Gauss in our tasks. Possibly they were included because they come out of the box when implementing collocation general solutions. So use them if curious, but expect radau_iia or gauss-legendre of same stage count to usually outperform them in either stability or accuracy.<br><br>The <strong>matrix below</strong> summarizes recommended pairings between sampler families and schedulers, based on documentation or credible tests:<br><br>| <strong>Sampler →</strong> | <strong>Multistep (RES, DPM)</strong> | <strong>Exponential (RES, ETD)</strong> | <strong>Linear RK</strong> | <strong>DIRK (Diag. Implicit)</strong> | <strong>Fully Implicit</strong> |<br>|-----------------------------|----------------------------------|-------------------------------|-------------------------|---------------------------|---------------------------|<br>| <strong>Scheduler ↓</strong> | <em>res_2m, dpmpp_2m, etc.</em> | <em>res_4s, etdrk4, etc.</em> | <em>rk4, rk6, dp5, etc.</em> | <em>pareschi2, crouzeix3, etc.</em> | <em>radau, gauss, etc.</em> |<br>| <strong>Normal (Linear)</strong> | Documented 👍<br>*(balanced for any)* | Ok (safe choice) | Documented 👍 | Good default (A-stable methods handle it) | Good default (A-stable) |<br>| <strong>Karras (EDM)</strong> | Great for DPM++ (noted)<br>Works well for RES | 👍 (refines detail at end) | 👍 (used in high-quality tests) | 👍 (these can handle fine refinements) | 👍 (these benefit from small end-steps) |<br>| <strong>Exponential</strong> | 👍 speeds convergence | Maybe overkill (duplication of concept) | Can cause jumpy start for high-order (use if <10 steps) | Fine (L-stable can handle big early drop) | Fine (A-stable can handle it) |<br>| <strong>SGM / DDIM Uniform</strong> | If matching original training (rare) | Suited if using SDE models | Not typical (but harmless) | Not needed (explicit scheduling fine) | Not needed |<br>| <strong>Linear-Quadratic</strong> | 🤷 No data (works generically) | 🤷 No specific note | Possibly helps (split coarse/fine) | Good for semi-stiff (some damping early, gentle later) | Not required (implicit already smooth) |<br>| <strong>KL Optimal</strong> | 👍 If minimizing step error is goal (multistep could benefit from optimal spacing) | Could help high-order too, but less needed | 👍 Theoretical best for given steps | 👍 Ensures implicit uses its power optimally | 👍 (Though these have tiny error anyway) |<br>| <strong>Bong_Tangent</strong> | 👍 Popular with RES 2m (improves detail) | Possibly redundant (exponential + tangent might double dip) | Could improve detail on low-order RK | Can use if want extra refinement pass effect | Probably not needed (implicit already “two-sided” in some stability sense) |<br>| <strong>Beta / Beta57</strong> | Fine (front-load or back-load noise drop as needed) | Fine, but exponential sched might conflict; Beta(0.5,0.7) used if desired look | Use if aiming for training-like noise profile or special needs | Fine (these methods aren’t sensitive to schedule shapes much) | Fine |<br><br>Key: “👍” indicates a documented or empirically supported good pairing; “🤷” means no info or not significant either way.<br><br>*(Example use of matrix: If using a <strong>Multistep</strong> sampler, <strong>Exponential</strong> scheduler is a documented good choice for speed; if using a <strong>Fully Implicit</strong> sampler, any scheduler works, but no special pairing is required due to strong stability.)*<br><br>## Notable Parameters in RES4LYF<br><br>Finally, a few custom parameters in the RES4LYF interface deserve explanation, as they differ from standard KSampler:<br><br>- <strong>ETA (η):</strong> This controls the amount of noise added each step in SDE mode. η = 0 means no noise (ODE mode); higher η means more random perturbation per step. There are multiple <em>NOISE_MODE_SDE</em> options that interpret η differently (listed in increasing strength). For most modes, η ≥ 1 triggers an internal rescaling to avoid blow-up, except “exp” mode which tolerates very large η. Typical safe range: 0 (deterministic) up to ~0.8 for mild creativity; 0.8–1.0 is already quite strong; above 1.0 only in “exp” mode or if you specifically want extreme variance. Keep η low if you want convergence; high η can prevent the image from settling (it “keeps adding noise” each step as one user noted).<br>- <strong>Shift & Base Shift:</strong> These relate to <strong>rectified flow models</strong> like SD3.5, Aura, Flux. “Shift” is akin to <code>max_shift</code> in Flux, controlling how much latent shift occurs per step for motion dynamics. “Base_shift” is specifically for Flux to offset its inherent motion. In static image generation, these are usually set to -1 (disabled) unless using those model types. If using e.g. WAN or HiFlow models that have a temporal component, you might set a shift (like 0.2 or so) to induce consistent movement across frames. For normal SD, <strong>leave Shift = -1</strong> (no shift). If you do play with it, “exponential” vs “linear” shift scaling can be chosen – exponential is default for Flux (smooth ramping) and generally gives better results in guided motion.<br>- **Steps vs Steps_to_run:** <strong>Steps</strong> is the total planned steps. <strong>Steps_to_run</strong> allows you to execute only a portion of those steps. For example, Steps = 30 and Steps_to_run = 10 will run 10 steps out of 30, and then stop, presumably allowing a second sampler to continue from 10 to 30. This is useful in <strong>chained sampling</strong> workflows: you can have one sampler do the first part (coarse) and another do the rest (fine) without resetting the noise. If Steps_to_run is set to -1 (default), it will run all steps. In practice: to have Sampler A do N₁ steps and Sampler B do N₂ steps (with total = N₁+N₂), set both Samplers’ Steps = N₁+N₂, then set Steps_to_run = N₁ for the first, and -1 for the second (which will then run the remaining N₂). This is how users combine different samplers sequentially. If not chaining, leave Steps_to_run = -1 (so it uses the Steps value fully). <em>Safe ranges:</em> -1 or 1 to Steps; must not exceed Steps. This feature is advanced; the UI doesn’t deeply document it beyond usage in forum examples.<br>- <strong>Sampler Mode (Standard/Unsample/Resample):</strong> By default use <strong>Standard</strong>. <em>Unsample</em> mode is for special <em>noise inversion</em> workflows: it disables adding model noise, letting you diffuse upwards (from a low-noise image to a high-noise one = “unsampling”) in order to then resample back down. <em>Resample</em> is the second phase after unsample – it similarly suppresses noise addition so that you only apply denoising that was deferred. In short, these are for when you do two-stage generation: first unsample (create a noise-expanded latent that still contains the image structure), then resample (bring it back down to low noise in a refined way). If you’re not doing those elaborate techniques, keep it on Standard (which ensures standard noise addition in normal SDE and normal denoising).<br>- <strong>Implicit Predictor (Implicit_Sampler_Name & Implicit_Steps):</strong> These let you do <strong>one sampler inside another</strong> as a predictor-corrector pair. For example, explicit = Euler, implicit = Radau, with implicit_steps = 2 means each step, it will first do Euler prediction then do 2 iterations of Radau corrector. This can dramatically improve quality on some models (e.g. SD3.5 Medium, as doc says it reduces artifacts). However, runtime multiplies by (1 + implicit_steps <em>implicit_stages). Typically “gains diminish after 2–3 implicit steps”. If you use this, choose a fast explicit predictor (Euler or RES_2m) so you don’t double heavy costs. For everyday use, this is overkill; but if you hit a scenario with subtle artifacts that only vanish with implicit refinement, consider using 1–2 implicit steps of a Gauss or Radau method (with same stage count as predictor or higher). Always set Implicit_Sampler_Name to “none” (or leave default) if not using this; implicit_steps = 0 or none is effectively off.</em><em><br></em><em><br></em><em>## Reproducible Test Recipes</em><em><br></em><em><br></em><em>To help you validate these samplers and schedulers yourself, here are a few example setups. Each uses a fixed seed and known model, so you can reproduce identical results. After generating, you can compare outputs to see the effect of different samplers or step counts.</em><em><br></em><em><br></em>*Setup:** We use two models – <em>Stable Diffusion XL (SDXL)</em> for photo-realistic prompts, and <em>Qwen-Image (Qwen)</em> for artistic illustration prompts. We give a prompt, and test a grid of 4 samplers × 3 schedulers × a few step counts. Due to length, we suggest doing this systematically by changing one factor at a time.<br><br>1. <strong>Prompt A (Portrait, Photorealistic)</strong> – <em>“Ultra-detailed studio portrait of an elderly man with wrinkled skin, dramatic lighting, 85mm lens”</em>; Model: <strong>SDXL</strong>; CFG scale = 7.5; Seed = 12345.<br> - Run with <strong>RES_3M</strong> + <strong>Karras</strong> at 20, 30, 40 steps. Then with <strong>Euler</strong> + <strong>Exponential</strong> at 20, 30 steps. Observe: RES_3M should achieve similar or better detail at 20 steps as Euler at 40 steps. The Euler+Exponential 20-step may look decent globally but lack pore-level detail; RES_3M 20-step should have sharper micro-details (e.g. in the wrinkles) and fewer weird artifacts.<br> - (Optional) Also try <strong>Radau_IIA_3s</strong> + <strong>Normal</strong> at 20 steps. It should be extremely clean (no noise speckles in shadows), but might be slightly less contrasty than RES_3M (due to heavy damping). This validates the stability vs detail trade-off.<br><br>2. <strong>Prompt B (Product Shot)</strong> – <em>“A shiny red sports car on a reflective floor, studio lighting, 8K photography”</em>; Model: <strong>SDXL</strong>; CFG = 8; Seed = 22222.<br> - Compare <strong>RK4_4s</strong> vs <strong>SSPRK3_3s</strong> vs <strong>Gauss-Legendre_4s</strong>, all with <strong>Normal</strong> scheduler, 25 steps.<br> - The car’s edges and reflections are a good test. RK4 might produce very crisp reflections but possibly some aliasing. SSPRK3 might slightly smooth out the reflection but ensure no ringing. Gauss-Legendre_4s will likely give the best of both: crisp and no ringing, owing to its high order and stability. These differences might be subtle, so zoom in on reflections or specular highlights to spot overshoot or oscillations.<br> - This confirms classical RK vs SSP vs fully implicit differences in a controlled way.<br><br>3. <strong>Prompt C (Landscape, Smooth gradients)</strong> – <em>“A serene sunset over a lake, smooth water, mountains silhouette, minimalistic”</em>; Model: <strong>Qwen-Image</strong>; CFG = 7; Seed = 33333.<br> - Try <strong>Euler</strong> + <strong>Bong_Tangent</strong> vs <strong>Heun_2s</strong> + <strong>Normal</strong> vs <strong>Crouzeix_2s</strong> + <strong>Normal</strong>; 50 steps each (Qwen tends to need more steps).<br> - Euler + Bong_Tangent: Should converge quickly and the tangent schedule will refine the gradient transitions – you might see a bit of banding reduction due to the backward pass. Heun 2s normal: stable, smooth, likely very similar result but maybe slightly softer hues. Crouzeix 2s: as a 3rd-order implicit, should produce an almost identical image to Heun but with maybe even less banding (since it’s effectively like doing a subtle implicit smoothing).<br> - This shows how a clever scheduler (bong_tangent) can emulate some benefits of implicit refinement for explicit Euler.<br><br>4. <strong>Prompt D (Abstract/Creative)</strong> – <em>“An abstract swirl of colorful paint, dynamic movement, 4K render”</em>; Model: <strong>Qwen-Image</strong>; CFG = 11 (high to induce challenge); Seed = 44444.<br> - Use <strong>DPMPP_SDE_2S</strong> + <strong>Karras</strong> at 30 steps, vs <strong>DPMPP_2M</strong> + <strong>Karras</strong> at 30 steps, vs <strong>RES_2M</strong> + <strong>Karras</strong> at 30 steps.<br> - With SDE 2S, each run will differ; do 2–3 and pick a representative. You’ll see more wild variations in the swirl patterns (the SDE sampling explores more) – it might even add small paint splatters or deviations. DPMPP_2M (ODE) will be consistent and probably a bit less chaotic in patterns. RES_2M likely gives similar structure to DPMPP_2M but perhaps sharper edges (since RES was refined).<br> - This highlights ODE vs SDE: SDE adds creativity at the cost of consistency.<br><br>5. <strong>Prompt E (Inpainting scenario)</strong> – <em>Use a partially masked image (like a photo with an object removed) and prompt: “fill the mask with matching background pattern”.</em><br> - Model: <strong>SDXL</strong> or relevant inpainting model; use <strong>Implicit Predictor</strong>: e.g. set Sampler = <strong>Euler</strong> explicit, Implicit_Sampler = <strong>Lobatto_IIIC_3s</strong>, Implicit_steps = 1, Steps = 50.<br> - Also run normal Euler 50, and maybe Euler 50 with a standard denoise=0.8 vs Euler+Lobatto with denoise=0.8.<br> - The implicit refinement might better maintain continuity at mask boundaries (reducing any seam or brightness mismatch, since it’s effectively doing a small implicit smoothing on the result each step). It’s hard to quantify without visual, but look closely at the filled area boundaries: the hybrid method should integrate the conditions a bit more consistently (less jitter or brightness difference).<br> - This demonstrates the use of implicit corrector in a practical task (inpainting requires stability to not hallucinate edges where mask meets content).<br><br>For each test, ensure you fix the seed and only change the intended variables, so differences come purely from the sampler or scheduler. You can then magnify specific regions to spot differences (e.g. graininess in dark areas to test stability, or edge sharpness to test accuracy).<br><br>## FAQ (Frequently Asked Questions)<br><br>**Q: Why does SDE mode often look more <em>creative</em> or varied than ODE mode?** <br>**A:** Because SDE sampling adds random noise after every step. This “continuous noise injection” means the sampler isn’t following one deterministic path; it’s exploring around it. Thus, even with the same seed, SDE can produce different outputs on different runs (the seed just sets initial noise, thereafter each step injects new randomness). At high step counts, ODE will converge to a fixed image (no new noise, just refining), whereas SDE will <em>never fully converge</em> – it keeps perturbing details. This can yield more interesting or detailed textures (the noise may push the model into new ideas), but it can also prevent the image from settling, sometimes resulting in a grainy look if overdone. It’s useful when you want <em>variety</em> or a touch of chaos – e.g. for creative art or when you deliberately want each outcome unique. Dial η (eta) to adjust how exploratory it is. Low η (0.1–0.3) gives mild variety, high η (~0.8 or 1) gives wild changes and risk of not converging to the prompt subject at all.<br><br>**Q: Some exponential integrator samplers (RES_S) feel like they produce <em>sharper edges</em> than others – is that real or placebo?** <br>**A:** It can be real. Exponential integrators can reduce numerical diffusion. In explicit Euler or low-order RK, each step might blur details a bit (numerical error smears high-frequency info). High-order exponential methods (like res_4s_minchev, etdrk4, etc.) have much lower local error and treat the linear decay exactly, so they preserve sharp transitions more faithfully. For example, a fine edge or thin line in the image might slightly fade with a basic solver at low steps, but remain crisp with a higher-order solver given the same step count. Community tests did find certain high-order samplers delivered noticeably sharper results. However, beyond a point, the model’s own learned blur vs sharpness dominates – a sampler can’t add detail the model wouldn’t produce with enough steps. It’s about how quickly it gets to that detail. So yes, a better integrator reaches a sharp result in fewer steps, which looks sharper for a given step count.<br><br>**Q: Why do implicit methods consume more GPU time?** <br>**A:** Implicit solvers require solving equations for each step. In practice, RES4LYF likely uses iteration or multiple model calls per step for implicit stages. For example, Gauss-Legendre 4s has 4 coupled stages – one way to solve is to iterate the network 4 times per step until consistency (or use a direct formula if linear, but our model’s score function is non-linear in latent so it likely iterates). This means more forward passes. Even if it’s just a small fixed number of iterations, that’s extra cost. Also, certain implicit methods might be implemented by internally calling the model function multiple times per stage for root-finding. So while explicit 4s = 4 evals, implicit 4s could be, say, 8–12 evals total if using Newton’s method or fixed-point iterations. The trade-off is you often can take larger steps (so need fewer steps). If you find implicit samplers too slow, consider using the <strong>Implicit Predictor</strong> with a small number of implicit refinements on an explicit sampler – that way you target the expensive implicit work where it’s most needed (e.g. one refinement per step to clean up artifacts, instead of fully implicit each step).<br><br>**Q: With so many samplers, how do I choose quickly?** <br>**A:** Identify your priority: <strong>Speed</strong> → try multistep (RES or DPM) with an aggressive scheduler (Exponential or BongTangent). <strong>Quality</strong> → try a multi-stage sampler (3s or 4s) with Karras. <strong>Robustness</strong> (no weird artifacts) → try an implicit or hybrid sampler (Radau, Crouzeix, etc.) – they are very reliable. If unsure, the <strong>RES samplers</strong> were created as balanced choices (fast and good quality). For example, <em>RES_3M + Karras 20 steps</em> is a strong general choice. As you gain experience, you might have favorites for certain styles (some report <em>Euler a</em> with tangent schedule for sketchy art, <em>DPM++ 2M Karras</em> for photorealism, etc., largely based on anecdote). Our guide’s cheat-sheet and individual notes can steer you: e.g. for smooth skin – maybe an implicit method; for wild textures – maybe SDE plus high order. It’s encouraged to do small A/B tests: generate 4 images, 2 with sampler X and 2 with sampler Y (same seed for each pair), to directly compare. The differences can be subtle but meaningful for discerning eyes.<br><br>**Q: Are higher-order samplers always better? Why not just use the highest (like 16s) always?** <br>**A:** Not always – diminishing returns and practical limits. A 16th-order method can take huge steps with small error <em>in theory</em>, but our diffusion model isn’t a simple equation – if you take steps too large, you might skip over nuances of the model’s response. Also, floating-point round-off error grows with very high order (summing many stages). There’s also the compute cost: 16s means 16 network evaluations per step; you could instead do, say, 8 steps of a 4th order method (=32 evals) versus 2 steps of a 16th order (=32 evals) – perhaps similar cost, but the latter might not capture the model’s non-linear evolution as well with just 2 big steps. In practice, methods up to 4th or 5th order seem to hit a sweet spot for diffusion, and beyond that, model uncertainty overtakes numerical error. <strong>RES4LYF’s author notes much is experimental</strong>, so 15s/16s are likely there for completeness. Use them if you’re curious or in very special cases (maybe extremely low step count runs), otherwise something like RES_5S or RK6_7s is usually sufficient to get nearly error-free results if you give it moderate steps. Always weigh time vs quality: a 8th-order with half the steps might still be slower than a 4th-order with full steps, depending on stage count. <br><br>**Q: I got NaNs using an ETD sampler with a high η noise – what happened?** <br>**A:** Likely you hit the scenario the docs warned about: for most SDE noise modes, η ≥ 1 triggers internal scaling to avoid NaNs. Possibly using “exp” noise mode bypasses that, allowing very large noise injection which can push the latent into unstable territory (model outputs inf). Or it could be that an exponential integrator without stabilization had an issue (Cox-Matthews ETDRK4 is known to potentially produce large intermediate values if not careful). Solution: reduce η below 1, or if you need η high, switch noise mode if available (e.g. use “exp” mode which expects large η but watch out as it allows more freedom). If it’s an integrator issue, try using the “alt” version of that sampler which might be a more stable formulation. Also, make sure your CFG scale isn’t outrageously high; a combination of high CFG and SDE noise could cause instabilities independent of the integrator.<br><br>**Q: What is “BONGMATH” actually?** <br>**A:** “Bongmath” isn’t a standard term in numerical analysis – it’s more of an internal nickname by the author for the novel noise scaling and scheduling tricks in RES4LYF. Specifically, it refers to the custom SDE implementation that allows using SDE sampling on models not originally trained for it (like stable diffusion), by adjusting how noise is injected and scaled, plus things like the bong_tangent schedule that do forward-backward denoising. It’s a playful term, but practically it means: thanks to some math, we can do fancy stuff (like bidirectional denoising, unsampling, etc.) that standard samplers don’t support. For users, it means you have extra capabilities – but you don’t need to understand the math under the hood. Just know that enabling SDE mode and using Bong_tangent, etc., are part of the “bongmath” toolkit that gives potentially superior results (the README calls it “power of bongmath” with an example of RES sampler outperforming UniPC). So, it’s the secret sauce behind many of these advanced samplers.<br><br>---<br><br>Below we include two CSV tables as an appendix: one summarizing all the samplers with their documented properties, and one for the schedulers. Use them as a quick reference. Finally, see the <strong>Glossary</strong> for definitions of the technical terms and method names we’ve cited. Happy experimenting with RES4LYF – may your generations be ever more precise and imaginative!<br><br>### Appendix: Samplers Table (CSV)<br><br>The columns are: <code>sampler_name, class_family, documented_order, stages_model_calls, notable_behavior, recommended_steps, suggested_scheduler, sources</code>. (Note “documented_order” is based on RES4LYF docs or known theory; if blank or <code>?</code>, it’s not specified explicitly).<br><br>```csv<br>sampler_name, class_family, documented_order, stages_model_calls, notable_behavior, recommended_steps, suggested_scheduler, sources<br>res_2m, multistep (explicit), \"not stated (est. 2)\", \"1 per step\", \"Refined DPM++ 2-step; very accurate for low cost\", \"20-30\", \"Karras or BongTangent\", \"\"<br>res_3m, multistep (explicit), \"not stated (est. 3)\", \"1 per step\", \"Refined 3-step; achieves high quality in few steps\", \"20\", \"Karras\", \"\"<br>dpmpp_2m, multistep (explicit), 2, \"1 per step\", \"DPM++ 2M from literature; reliable baseline\", \"20-50\", \"Karras\", \"\"<br>dpmpp_3m, multistep (explicit), 3, \"1 per step\", \"DPM++ 3M (third-order multi-step variant)\", \"15-30\", \"Karras\", \"\"<br>abnorsett_2m, multistep (explicit), \"not specified\", \"1 per step\", \"Multi-step (2) possibly Adams/Norsett method\", \"30?\", \"Normal\", \"\"<br>abnorsett_3m, multistep (explicit), \"not specified\", \"1 per step\", \"Multi-step (3) Norsett method\", \"30?\", \"Normal\", \"\"<br>abnorsett_4m, multistep (explicit), \"not specified\", \"1 per step\", \"Multi-step (4) Norsett method\", \"30-40?\", \"Normal\", \"\"<br>deis_2m, multistep (explicit), 2, \"1 per step\", \"DEIS (Diffusion Exp. Integrator) 2-step variant\", \"20-40\", \"SGM or Karras\", \"\"<br>deis_3m, multistep (explicit), 3, \"1 per step\", \"DEIS 3-step variant\", \"20-30\", \"SGM or Karras\", \"\"<br>deis_4m, multistep (explicit), 4, \"1 per step\", \"DEIS 4-step variant\", \"15-30\", \"SGM or Karras\", \"\"<br>res_2s, exponential explicit, \"not stated (est. 2)\", 2, \"2-stage Lawson/ETD; base version\", \"30+\", \"Normal\", \"\"<br>res_2s_stable, exponential explicit, \"2\", 2, \"Stability-optimized 2-stage\", \"30+\", \"Normal\", \"\"<br>res_2s_rkmk2e, exponential explicit, 2, 2, \"Runge-Kutta-MuntheKaas 2-stage variant\", \"30+\", \"Normal\", \"\"<br>res_3s, exponential explicit, \"not stated (est. 3)\", 3, \"3-stage ETD (base version)\", \"20-30\", \"Karras\", \"\"<br>res_3s_non-monotonic, exponential explicit, 3, 3, \"3-stage variant allowing overshoot (more accurate)\", \"20-30\", \"Karras\", \"\"<br>res_3s_alt, exponential explicit, 3, 3, \"Alternate 3-stage coefficients\", \"20-30\", \"Karras\", \"\"<br>res_3s_cox_matthews, exponential explicit, 3, 3, \"ETDRK3 by Cox & Matthews (2002)\", \"20-30\", \"Normal\", \"\"<br>res_3s_lie, exponential explicit, 3, 3, \"Lie-Trotter type 3-stage integrator\", \"20-30\", \"Normal\", \"\"<br>res_3s_strehmel_weiner, exponential explicit, 3, 3, \"3-stage ETD by Strehmel & Weiner (1987)\", \"20-30\", \"Normal\", \"\"<br>res_4s_krogstad, exponential explicit, 4, 4, \"ETDRK4 by Krogstad (2005)\", \"15-25\", \"Karras\", \"\"<br>res_4s_krogstad_alt, exponential explicit, 4, 4, \"Alternate Krogstad 4th-order\", \"15-25\", \"Karras\", \"\"<br>res_4s_strehmel_weiner, exponential explicit, 4, 4, \"4-stage ETD by Strehmel & Weiner\", \"15-25\", \"Karras\", \"\"<br>res_4s_strehmel_weiner_alt, exponential explicit, 4, 4, \"Alternate S-W 4th-order\", \"15-25\", \"Karras\", \"\"<br>res_4s_cox_matthews, exponential explicit, 4, 4, \"ETDRK4 by Cox & Matthews\", \"15-25\", \"Karras\", \"\"<br>res_4s_cfree4, exponential explicit, 4, 4, \"4-stage commutator-free integrator\", \"15-25\", \"Normal\", \"\"<br>res_4s_friedli, exponential explicit, 4, 4, \"Friedli (1978) 4th-order ETD\", \"15-25\", \"Normal\", \"\"<br>res_4s_minchev, exponential explicit, 4, 4, \"4-stage by Minchev (2004); high accuracy\", \"15-20\", \"Karras\", \"\"<br>res_4s_munthe-kaas, exponential explicit, 4, 4, \"4-stage by Munthe-Kaas; similar to above\", \"15-20\", \"Karras\", \"\"<br>res_5s, exponential explicit, 5, 5, \"5-stage high-order (notable sampler)\", \"10-20\", \"Karras\", \"\"<br>res_5s_hochbruck-ostermann, exponential explicit, 5, 5, \"Hochbruck-Ostermann 5th-order (SIAM 2005)\", \"10-20\", \"Karras\", \"\"<br>res_6s, exponential explicit, \"not stated (~6)\", 6, \"6-stage (exp.); experimental high order\", \"8-15\", \"Karras\", \"\"<br>res_8s, exponential explicit, \"not stated (~8)\", 8, \"8-stage (exp.); very high order\", \"6-12\", \"Linear or Simple\", \"\"<br>res_10s, exponential explicit, \"not stated (~10)\", 10, \"10-stage (exp.); extremely high order\", \"5-10\", \"Linear\", \"\"<br>res_15s, exponential explicit, \"not stated (~15)\", 15, \"15-stage (exp.); extremely high order\", \"<10\", \"Linear\", \"\"<br>res_16s, exponential explicit, \"not stated (~16)\", 16, \"16-stage (exp.); extremely high order\", \"<10\", \"Linear\", \"\"<br>etdrk2_2s, exponential explicit, 2, 2, \"Exponential TD RK2 scheme\", \"30+\", \"Normal\", \"\"<br>etdrk3_a_3s, exponential explicit, 3, 3, \"ETDRK3 scheme (variant A)\", \"20-30\", \"Normal\", \"\"<br>etdrk3_b_3s, exponential explicit, 3, 3, \"ETDRK3 scheme (variant B)\", \"20-30\", \"Normal\", \"\"<br>etdrk4_4s, exponential explicit, 4, 4, \"ETDRK4 (original Cox-Matthews 4th order)\", \"15-25\", \"Normal\", \"\"<br>etdrk4_4s_alt, exponential explicit, 4, 4, \"ETDRK4 (stabilized/alt version)\", \"15-25\", \"Normal\", \"\"<br>dpmpp_2s, explicit (single-step), 2, 2, \"DPM++ 2S (2nd-order single-step)\", \"20-30\", \"Karras\", \"\"<br>dpmpp_sde_2s, explicit (single-step), 2, 2, \"DPM++ 2S for SDE (stochastic)\", \"20-30\", \"Karras\", \"\"<br>dpmpp_3s, explicit (single-step), 3, 3, \"DPM++ 3S (3rd-order single-step)\", \"15-25\", \"Karras\", \"\"<br>lawson2a_2s, exponential explicit, 2, 2, \"Lawson 2-stage method (variant a)\", \"30+\", \"Normal\", \"\"<br>lawson2b_2s, exponential explicit, 2, 2, \"Lawson 2-stage (variant b)\", \"30+\", \"Normal\", \"\"<br>lawson4_4s, exponential explicit, 4, 4, \"Lawson 4-stage 4th-order\", \"15-25\", \"Normal\", \"\"<br>lawson41-gen_4s, exponential explicit, 4, 4, \"Lawson 4(1) general method\", \"15-25\", \"Normal\", \"\"<br>lawson41-gen-mod_4s, exponential explicit, 4, 4, \"Modified Lawson 4(1)\", \"15-25\", \"Normal\", \"\"<br>lawson42-gen-mod_1h4s, hybrid (Lawson + explicit), \"≈4\", \"1 half-step + 4-stage\", \"Hybrid Lawson method\", \"?\", \"Normal\", \"\"<br>lawson43-gen-mod_2h4s, hybrid, \"≈4\", \"2 half-steps + 4-stage\", \"Hybrid Lawson method\", \"?\", \"Normal\", \"\"<br>lawson44-gen-mod_3h4s, hybrid, \"≈4\", \"3 half-steps + 4-stage\", \"Hybrid Lawson method\", \"?\", \"Normal\", \"\"<br>lawson45-gen-mod_4h4s, hybrid, \"≈4\", \"4 half-steps + 4-stage\", \"Hybrid Lawson method\", \"?\", \"Normal\", \"\"<br>pec423_2h2s, hybrid (PEC), \"≈3\", \"2-stage predictor + 2-stage corrector\", \"Predictor-explicit-corrector scheme\", \"?\", \"Normal\", \"\"<br>pec433_2h3s, hybrid (PEC), \"≈3\", \"2-stage pred + 3-stage corr\", \"Predictor-explicit-corrector scheme\", \"?\", \"Normal\", \"\"<br>abnorsett2_1h2s, hybrid, \"≈2\", \"AbNorsett + 2-stage\", \"AbNorsett predictor + RK corrector\", \"?\", \"Normal\", \"\"<br>abnorsett3_2h2s, hybrid, \"≈3\", \"AbNorsett + 2-stage\", \"AbNorsett predictor + RK corrector\", \"?\", \"Normal\", \"\"<br>abnorsett4_3h2s, hybrid, \"≈4\", \"AbNorsett + 2-stage\", \"AbNorsett predictor + RK corrector\", \"?\", \"Normal\", \"\"<br>irk_exp_diag_2s, hybrid (implicit-exp), 2, 2, \"Implicit RK with exponential treatment\", \"20-30\", \"Normal\", \"\"<br>euler, explicit (RK1), 1, 1, \"Euler (deterministic)\", \"30-100\", \"Exponential for speed; Karras for quality\", \"\"<br>ddim, explicit (ODE-like), \"~1\", \"1 (per step)\", \"DDIM deterministic sampler (ODE)\", \"50\", \"DDIM_uniform\", \"\"<br>gauss-legendre_2s, fully implicit, 4, 2, \"Gauss collocation (A-stable, symplectic)\", \"20\", \"Any (A-stable)\", \"\"<br>gauss-legendre_3s, fully implicit, 6, 3, \"Gauss collocation (order 6)\", \"15\", \"Any\", \"\"<br>gauss-legendre_4s, fully implicit, 8, 4, \"Gauss collocation (order 8)\", \"10-12\", \"Any\", \"\"<br>gauss-legendre_4s_alternating_a, fully implicit, 8, 4, \"Gauss 4s alt formulation A\", \"10-12\", \"Any\", \"\"<br>gauss-legendre_4s_ascending_a, fully implicit, 8, 4, \"Gauss 4s alt formulation A (order asc)\", \"10-12\", \"Any\", \"\"<br>gauss-legendre_4s_alt, fully implicit, 8, 4, \"Gauss 4s alternate (likely same order)\", \"10-12\", \"Any\", \"\"<br>gauss-legendre_5s, fully implicit, 10, 5, \"Gauss collocation (order 10)\", \"8-10\", \"Any\", \"\"<br>gauss-legendre_5s_ascending, fully implicit, 10, 5, \"Gauss 5s alt ordering\", \"8-10\", \"Any\", \"\"<br>radau_ia_2s, fully implicit, 3, 2, \"Radau IA collocation (order 3)\", \"20\", \"Any\", \"\"<br>radau_ia_3s, fully implicit, 5, 3, \"Radau IA (order 5)\", \"15-20\", \"Any\", \"\"<br>radau_iia_2s, fully implicit, 3, 2, \"Radau IIA (L-stable, order 3)\", \"20\", \"Any\", \"\"<br>radau_iia_3s, fully implicit, 5, 3, \"Radau IIA (L-stable, order 5)\", \"15-20\", \"Any\", \"\"<br>radau_iia_3s_alt, fully implicit, 5, 3, \"Radau IIA alt (maybe different starting guess)\", \"15-20\", \"Any\", \"\"<br>radau_iia_5s, fully implicit, 9, 5, \"Radau IIA (order 9)\", \"10-15\", \"Any\", \"\"<br>radau_iia_7s, fully implicit, 13, 7, \"Radau IIA (order 13)\", \"8-12\", \"Any\", \"\"<br>radau_iia_9s, fully implicit, 17, 9, \"Radau IIA (order 17)\", \"6-10\", \"Any\", \"\"<br>radau_iia_11s, fully implicit, 21, 11, \"Radau IIA (order 21)\", \"5-8\", \"Any\", \"\"<br>lobatto_iiia_2s, fully implicit, 2, 2, \"Lobatto IIIA (order 2)\", \"30\", \"Any\", \"\"<br>lobatto_iiia_3s, fully implicit, 4, 3, \"Lobatto IIIA (order 4)\", \"20\", \"Any\", \"\"<br>lobatto_iiia_4s, fully implicit, 6, 4, \"Lobatto IIIA (order 6)\", \"15-20\", \"Any\", \"\"<br>lobatto_iiib_2s, fully implicit, 2, 2, \"Lobatto IIIB (order 2)\", \"30\", \"Any\", \"\"<br>lobatto_iiib_3s, fully implicit, 4, 3, \"Lobatto IIIB (order 4)\", \"20\", \"Any\", \"\"<br>lobatto_iiib_4s, fully implicit, 6, 4, \"Lobatto IIIB (order 6)\", \"15-20\", \"Any\", \"\"<br>lobatto_iiic_2s, fully implicit, 2, 2, \"Lobatto IIIC (order 2)\", \"30\", \"Any\", \"\"<br>lobatto_iiic_3s, fully implicit, 4, 3, \"Lobatto IIIC (order 4)\", \"20\", \"Any\", \"\"<br>lobatto_iiic_4s, fully implicit, 6, 4, \"Lobatto IIIC (order 6)\", \"15-20\", \"Any\", \"\"<br>lobatto_iiic_star_2s, fully implicit, 2, 2, \"Lobatto IIIC* (variant) (order 2)\", \"30\", \"Any\", \"\"<br>lobatto_iiic_star_3s, fully implicit, 4, 3, \"Lobatto IIIC* (order 4)\", \"20\", \"Any\", \"\"<br>lobatto_iiid_2s, fully implicit, 2, 2, \"Lobatto IIID (order 2)\", \"30\", \"Any\", \"\"<br>lobatto_iiid_3s, fully implicit, 4, 3, \"Lobatto IIID (order 4)\", \"20\", \"Any\", \"\"</p>\n<h3 id=\"appendix:-schedulers-table-(csv)\">Appendix: Schedulers Table (CSV)</h3>\n<p>Columns: scheduler_name, type_formula, intuitive_behavior, use_case, sources.</p>\n<p>scheduler_name, type_formula, intuitive_behavior, use_case, sources<br>normal, linear sigma decay, Uniform noise decay each step, Balanced default for general use, \"\"<br>sgm_uniform, SDE sigma(t) uniform, Steps evenly spaced in SDE time, Good for score-based models (VP/VE), \"\"<br>karras, sigma_max to min with power-law (ρ), Smooth non-linear decay (dense at end), High-quality/detail images, \"\"<br>exponential, σ_i = σ_max <em>(σ_min/σ_max)^(i/(N-1)), Rapid early decay, Fast generation / fewer steps, \"\"</em><em><br></em><em>ddim_uniform, t steps uniform, Matches DDIM time spacing, When using DDIM sampler for authenticity, \"\"</em><em><br></em><em>beta, β distribution curve (α,β params), Non-linear (shape-determined) decay, Special effects or mimic training schedule, \"\"</em><em><br></em><em>normal (duplicate entry?), Gaussian error? (maybe refers to EDM normal), </em>Potentially akin to “PNDM”*, (Often paired with AlignYourSteps/skip) – likely similar to linear, \"\"<br>linear_quadratic, piecewise: linear then quadratic, Fast initial drop, gentle finish, Complex scenes (prevent overdenoise late), \"\"<br>KL_optimal, derived from KL divergence optimality, Uneven but theoretically minimal KL error, To maximize fidelity at fixed steps, \"\"<br>bong_tangent, two-stage tangent function, Forward-then-back half-interval noise, Ultra-sharp details, bidirectional correction, \"\"<br>beta57, β schedule with α=0.5, β=0.7, Front-loaded slight curve, As alternative to default linear, \"\"<br>simple, basic fixed spacing (legacy), Simplified decay (maybe ~linear), Quick tests / baseline, \"\"</p>\n<p><em>(Note: The “normal” entry appears twice in different contexts (one likely the Standard linear, the other maybe a normal distribution spacing from AlignYourSteps). In ComfyUI BasicScheduler docs, “normal” is linear. The duplicate mention in sources like A1111 suggests it’s same. We treat it as linear.)</em></p>\n<h3 id=\"glossary\">Glossary</h3>\n<p>· <strong>ODE vs SDE:</strong> An <strong>ODE (Ordinary Differential Equation)</strong> sampler (like probability flow ODE in diffusion) deterministically integrates the denoising equation with no added randomness. An <strong>SDE (Stochastic DE)</strong> sampler includes a noise term (like the reverse diffusion SDE), adding randomness each step. ODE mode converges to one solution given a seed; SDE mode yields a distribution of solutions and can explore more. SDE often requires special noise scaling (“Bongmath”) to work well on models not originally continuous-time.</p>\n<p>· <strong>RK (Runge–Kutta):</strong> A family of one-step integrators that evaluate the derivative (here, model’s output) multiple times per step with different weights. “Explicit” RK means later stage outputs don’t depend on future ones (straightforward to compute sequentially). <strong>Implicit</strong> RK means stages satisfy equations including later stages – requiring solving equations (harder but more stable).</p>\n<p>· <strong>Diagonally Implicit RK (DIRK):</strong> A subset of implicit RK where the coefficient matrix is lower-triangular – each stage only depends on previous and current stage, not future. Solvable stage-by-stage (usually via one nonlinear solve per stage). These are often A-stable and sometimes L-stable. E.g. Kraaijevanger–Spijker 2s, Crouzeix’s methods.</p>\n<p>· <strong>Gauss–Legendre methods:</strong> Implicit RK with collocation at Gaussian quadrature nodes. They achieve the maximum possible order 2s for s stages. They are A-stable and symplectic (no numerical damping or growth). Great accuracy but fully implicit.</p>\n<p>· <strong>Radau IIA methods:</strong> Collocation at Radau nodes including the right end. Order 2s-1, L-stable (damps out all high-frequency errors). Often used for very stiff ODEs due to strong stability. Radau IA uses left end, also order 2s-1 but not L-stable.</p>\n<p>· <strong>Lobatto IIIA/B/C/D:</strong> Collocation including both interval ends. Order 2s-2 typically. They come in pairs (IIIA & IIIB are dual and symplectic when coupled, IIIC is another variant, IIID perhaps dual to C?). They are A-stable, some are symmetric (good for reversible systems).</p>\n<p>· <strong>Lawson method:</strong> Technique by J. Lawson (1967) for stiff ODEs: use an integrating factor $e^{At}$ for the linear part, then do an explicit RK on the transformed system. E.g. $y' = Ay + N(y)$ becomes $z = e^{-At}y$, integrate z explicitly. Many exponential integrators (ETD) are based on this idea.</p>\n<p>· <strong>ETD (Exponential Time Differencing) RK:</strong> A class of integrators (Cox & Matthews 2002, Krogstad 2005, etc.) that handle the linear part of the ODE exactly via matrix exponentials and the nonlinear part via RK expansions. They use functions φ_k (matrix exponential integrals) to achieve high order. Popular ETD schemes: ETDRK2, ETDRK4, etc. Require careful computation to avoid NaNs (Cox-Matthews noted instability and proposed contour integrals to compute φ functions).</p>\n<p>· <strong>SSPRK (Strong Stability Preserving RK):</strong> Explicit RK methods that preserve monotonicity/TV stability under certain time-step constraints. Shu-Osher developed SSPRK3 (3rd order, 3 stages) which is extensively used in CFD because it won’t create new extrema if none exist initially (for small enough step). It’s basically the best compromise between order and TVD stability (no 3rd-order 2-stage exists, and 4th-order explicit SSP requires many stages).</p>\n<p>· <strong>DPM++ (Denoising Diffusion Probabilistic Models++)</strong>: A family of samplers from diffusion research (especially by Chensiang (Liu) and team). DPM++ 2S, 2M, etc., are improvement over DDIM and PLMS solvers, using mid-point correctors and higher order. They were integrated in many UIs as they improved speed and quality. “M” stands for multi-step (uses previous points), “S” for single-step (higher order in one step).</p>\n<p>· <strong>DEIS (Diffusion Exponential Integrator Sampler):</strong> From Zhang & Chen 2022. Uses an exponential integrator viewpoint to create fast samplers. Introduced with ρ-th order versions and shown to be effective at low steps. Essentially applies an exponential time differencing approach to diffusion ODE/SDE.</p>\n<p>· <strong>DDIM (Denoising Diffusion Implicit Models):</strong> Sampler by Song et al. that solves a non-stochastic version of diffusion, allowing deterministic generation and accelerated sampling. It’s effectively an ODE (probability flow ODE) solver, equivalent to a particular first-order method (like implicit Euler in noise space). The “ddim_uniform” scheduler ensures time steps are evenly spaced, replicating how the DDIM paper defined it.</p>\n<p>· <strong>Heun’s method:</strong> Also known as explicit trapezoidal or improved Euler. A 2-stage RK: first stage Euler predictor, second stage corrector using the average of slopes. Second order accurate. Minimizes error for 2-stage along with Midpoint and Ralston’s being alternatives.</p>\n<p>· <strong>Ralston’s methods:</strong> Ralston (1962) provided coefficients for RK methods that minimize the error bounds. For 2-stage, his method uses c=2/3 with weights (2/3, 0) and yields error constant minimized. For 3-stage and 4-stage, he proposed specific sets also aimed at minimizing error norm.</p>\n<p>· <strong>Bogacki–Shampine:</strong> A 3-stage RK that’s third-order with an embedded second-order (used in ode23). Also a 7-stage 5th-order pair was later given by the same authors. They are known for reliable adaptive step solvers for moderate accuracy.</p>\n<p>· <strong>Dormand–Prince:</strong> A 7-stage RK with 5th order accuracy and an embedded 4th order (the classic “ode45” method). It’s efficient because it chooses coefficients such that the higher order solution is very accurate and the lower order is just used for error. Also a 13-stage 8th order DP method exists (DOP853).</p>\n<p>· <strong>Krogstad (ETD4):</strong> Krogstad (2005) introduced a 4th-order exponential integrator that became a standard because it was easier to implement than Cox–Matthews ETD4 while maintaining accuracy. It’s used often in solving semi-linear PDEs.</p>\n<p>· <strong>Cox–Matthews (ETD3/4):</strong> The 2002 paper that laid groundwork for modern ETD integrators. ETD4 in that paper (with 4 exponentials computed) became widely used but had stability issues that were later solved by others (like Kassam & Trefethen 2005).</p>\n<p>· <strong>Minchev & Munthe-Kaas:</strong> In 2004, they did analysis of exponential integrators and introduced commutator-free schemes and considered order conditions to high order. They likely provided examples of 4th order schemes and beyond, which are implemented as res_4s_minchev, etc.</p>\n<p>· <strong>Hochbruck–Ostermann:</strong> Mathematicians who published reviews and new exponential integrators in mid-2000s (and beyond). They, for example, showed in 2005 that certain high-order ETD schemes aren’t stable for some stiff problems and developed stabilized versions (like energy-decay preserving ETD3, etc.). The res_5s_hochbruck-ostermann probably references a scheme from one of their papers, maybe a 5th-order method adapted for diffusion equations.</p>\n<p>· <strong>Crouzeix:</strong> Known for work in RK methods, he found some diagonally implicit schemes, including the 2-stage 3rd-order and 3-stage 4th-order we have here<a target=\"_blank\" rel=\"noopener ugc\" href=\"https://en.wikipedia.org/wiki/List_of_Runge%E2%80%93Kutta_methods#:~:text=Crouzeix%27s%20two,Implicit%20Runge%E2%80%93Kutta%20method\">[3]</a>. Those are remarkable because usually a 2-stage can’t be 3rd order unless implicit. Crouzeix’s methods are often A-stable but not L-stable.</p>\n<p>· <strong>Kraaijevanger–Spijker:</strong> These authors studied stability of RK. The DIRK here might be one that maximizes the radius of absolute monotonicity (or SSP coefficient). Spijker’s condition is a measure of RK stability for non-linear problems. The method given (with coefficients in Wikipedia) is L-stable if certain coefficient chosen, but not sure if theirs is specifically L-stable or just A-stable. It’s included likely as an example of a stable DIRK.</p>\n<p>· <strong>Pareschi–Russo:</strong> They worked on IMEX (implicit-explicit) methods for hyperbolic and kinetic equations. The DIRK listed under them is likely the second order L-stable DIRK from one of their papers in early 2000s, which has a free parameter x (the condition for L-stability given was specific solutions for x).</p>\n<p>· <strong>Implicit vs Explicit in diffusion context:</strong> Explicit samplers are simpler and usually fine because diffusion ODEs aren’t extremely stiff (the “stiffness” comes from the fast decay of sigma at end, but Karras schedule mitigates that). Implicit ones shine if you have extremely large steps or combine with other processes (like controlling shift in video models – implicit ensures no blow-up when shifting prompts). They also can guarantee stability if someone cranks settings to extremes (high CFG, few steps – an implicit might handle it gracefully where an explicit would break).</p>\n<p>· <strong>Flux, HiDiff, WAN</strong> references: These are specific model names (Flux, HiDream, WAN) that the documentation references for features like regional conditioning and shift. Not directly part of samplers, but SHIFT parameters relate to them. Flux uses an “exponential” shift by default – meaning the influence of shift increases exponentially with time; SD3.5 uses linear. The UI giving both options allows using these samplers across different model architectures smoothly.</p>\n<p>Each of these terms and methods could warrant more explanation, but this glossary provides a starting point. For deeper theory, see references like Butcher’s <em>Numerical Methods for ODEs</em>, Hairer & Wanner’s <em>Solving ODEs</em> series (which cover stability, collocation, etc.), and the specific papers: - Karras et al., <em>Elucidating Design Spaces of Diffusion Models</em> (for sigma schedules). - Lu et al., <em>DPM-Solver</em> (for DPM++ family). - Zhang & Chen, <em>Fast Sampling of Diffusion Models with Exponential Integrator (DEIS)</em>. - Song et al., <em>DDIM</em> (2020) for DDIM. - Shu & Osher (1988) for SSPRK. - Lawson (1967) for integrating factor methods. - Cox & Matthews (2002), Krogstad (2005) for ETD. - etc., as cited.</p>\n<p>By understanding these, you can better appreciate why RES4LYF includes them – essentially to give you every possible tool to balance speed, quality, and stability in diffusion sampling.</p>\n<p><strong>References:</strong></p>\n<ol>\n<li>\n<p>ClownsharkBatwing, <em>“Superior Sampling with RES4LYF: The Power of BONGMATH”</em>, GitHub README – <em>RES samplers reach high quality in ~20 steps, far fewer than common samplers</em></p>\n</li>\n<li>\n<p>ClownsharkBatwing, <em>RES4LYF ComfyUI Node Documentation (Inputs)</em> – <em>SDE sampling identical to ODE except noise added each step (controlled continuous noise injection)</em></p>\n</li>\n<li>\n<p>ClownsharkBatwing, <em>RES4LYF Node Documentation (Noise Settings)</em> – <em>ETA ≥ 1 triggers internal scaling to avoid NaNs (except “exp” mode which allows >1)</em></p>\n</li>\n<li>\n<p>ClownsharkBatwing, <em>RES4LYF Node Documentation (Scheduler settings)</em> – <em>Extra scheduler “beta57”: beta schedule with α=0.5, β=0.7</em></p>\n</li>\n<li>\n<p>ComfyUI BasicScheduler Docs – <em>List of schedulers: normal (linear), karras (smooth non-linear), exponential (rapid decay), etc. – characteristics and use cases</em></p>\n</li>\n<li>\n<p>ComfyUI BasicScheduler Docs – <em>KL optimal: theoretical KL divergence optimized decay (from AlignYourSteps)</em></p>\n</li>\n<li>\n<p>ComfyUI BasicScheduler Docs – <em>SGM Uniform: scheduler optimized for score-based generative models; DDIM Uniform: for DDIM sampling</em></p>\n</li>\n<li>\n<p>ComfyUI Node listing (ClownsharKSampler) – <em>Extension description: enables SDE sampling with all models; 115 sampler types, 24 noise types, 11 noise scaling modes (bongmath)</em></p>\n</li>\n<li>\n<p>Reddit – <em>User guide on chaining samplers:</em> <em>Set both samplers same total steps; use steps_to_run = N on first, -1 on second to split steps</em></p>\n</li>\n<li>\n<p>Reddit Q&A – <em>Non-deterministic (SDE) samplers don’t converge at high steps – will keep adding noise each step</em></p>\n</li>\n<li>\n<p>Claudio Beck, <em>RES4LYF Sampler Comparison</em> – <em>Fastest: Exponential DDIM & Euler ~24.5s; Best quality: linear RK6_7s (6th order) ~167s; Very good: Radau IIA 7s ~227s; Very good: Exponential res_4s_minchev ~95s</em></p>\n</li>\n<li>\n<p>Wikipedia – <em>Gauss–Legendre methods:</em> <em>s-stage Gauss method has order 2s (highest possible), A-stable for all s (proved by Butcher)</em></p>\n</li>\n<li>\n<p>Arxiv (Persson 2008) – <em>Radau IIA schemes are L-stable with order 2s-1</em></p>\n</li>\n<li>\n<p>Arxiv (CWI 1991) – <em>Methods can be A-stable & order 2s (Gauss) or L-stable & order 2s-1 (Radau IIA)</em></p>\n</li>\n<li>\n<p>Wikipedia – <em>Diagonally Implicit RK:</em> <em>Example methods: Kraaijevanger–Spijker 2s (tableau given), Qin–Zhang 2s (2nd order symplectic), Pareschi–Russo 2s (2nd order, A-stable if x>=1/4, L-stable for specific x)</em></p>\n</li>\n<li>\n<p>Wikipedia – <em>Crouzeix DIRK:</em> <em>2-stage 3rd order, 3-stage 4th order (with parameter α)</em><a target=\"_blank\" rel=\"noopener ugc\" href=\"https://en.wikipedia.org/wiki/List_of_Runge%E2%80%93Kutta_methods#:~:text=Crouzeix%27s%20two,Implicit%20Runge%E2%80%93Kutta%20method\">[3]</a></p>\n</li>\n<li>\n<p>Butcher’s List (Wikipedia) – <em>Ralston’s methods:</em> <em>2-stage 2nd order (min error), 3-stage 3rd order (min error, used in Bogacki–Shampine embed), 4-stage 4th order (min error)</em></p>\n</li>\n<li>\n<p>Diffrax Docs – <em>Tsitouras’ 5/4 method:</em> <em>5th order, embedded 4th, 7 stages with FSAL (so effectively 6 function evals), used in ODE solvers</em></p>\n</li>\n<li>\n<p>Reddit (StableDiffusion) – <em>Explanation of DPM++ 2M vs 2S:</em> <em>“2M and 2S: variants of DPM++ using second-order; S means single-step, M means multi-step; slower but more accurate”</em></p>\n</li>\n<li>\n<p>Karras et al. (EDM 2022) – <em>Sigma schedule (ρ):</em> <em>Spacing sigmas with smooth non-linear decay yields improved sample quality; e.g. Karras schedule distributes steps in log space (here referred as “smooth transition”) – recommended for detail</em></p>\n</li>\n<li>\n<p>ClownsharkBatwing, <em>Node Docs (Sampler settings)</em> – <em>Samplers ending in “s” use multiple model calls (substeps) – e.g. 3s = 3 calls per step (3x slower than Euler), but much more accurate, especially with noise (SDE). “res” family are refined dpmpp with higher accuracy implemented here.</em></p>\n</li>\n<li>\n<p>ClownsharkBatwing, <em>Node Docs (Sampler settings cont.)</em> – <em>Samplers ending in “m” are multistep (recycle previous steps) – Euler-speed per step, less accurate but converge linearly (good for img2img, unsampling, etc. where linear convergence can be advantage).</em></p>\n</li>\n<li>\n<p>ClownsharkBatwing, <em>Node Docs (Implicit sampler usage)</em> – <em>Implicit_Sampler: using an implicit method as corrector improves coherence and reduces artifacts (notably on SD3.5 Medium). Use a fast explicit predictor (Euler, res_2m, etc.) to keep runtime sane; e.g. explicit res_5s + implicit gauss-legendre_5s for ultimate quality (but very slow, “commitment to climate change”).</em></p>\n</li>\n<li>\n<p>ClownsharkBatwing, <em>Node Docs (Implicit steps)</em> – <em>Implicit_steps: number of implicit refinement steps per explicit step. Increases runtime linearly (double, triple, etc.) – usually diminishing returns after 2-3 implicit steps.</em></p>\n</li>\n<li>\n<p>ClownsharkBatwing, <em>Node Docs (Shift settings)</em> – <em>SHIFT: same as model’s shift (max_shift in Flux) – set -1 to disable if not needed; BASE_SHIFT: only for Flux (set -1 to disable if not using Flux). SHIFT_SCALING: “exponential” (Flux default) vs “linear” (SD3.5 default) – exponential generally yields better results except niche uses.</em></p>\n</li>\n</ol>\n<p><a target=\"_blank\" rel=\"noopener ugc\" href=\"https://en.wikipedia.org/wiki/List_of_Runge%E2%80%93Kutta_methods#:~:text=,order%20method\">[1]</a> <a target=\"_blank\" rel=\"noopener ugc\" href=\"https://en.wikipedia.org/wiki/List_of_Runge%E2%80%93Kutta_methods#:~:text=Image%3A%20\">[2]</a> <a target=\"_blank\" rel=\"noopener ugc\" href=\"https://en.wikipedia.org/wiki/List_of_Runge%E2%80%93Kutta_methods#:~:text=Crouzeix%27s%20two,Implicit%20Runge%E2%80%93Kutta%20method\">[3]</a> List of Runge–Kutta methods - Wikipedia</p>\n<p><a target=\"_blank\" rel=\"noopener ugc\" href=\"https://en.wikipedia.org/wiki/List_of_Runge%E2%80%93Kutta_methods\">https://en.wikipedia.org/wiki/List_of_Runge%E2%80%93Kutta_methods</a></p>",
"author": {
"name": "admin"
},
"tags": [
"comfyui"
],
"date_published": "2026-07-21T15:58:41-04:00",
"date_modified": "2026-08-11T05:38:27-04:00"
},
{
"id": "https://graphicdesigngeek.com/krea2-expressions-with-muscle-prompting.html",
"url": "https://graphicdesigngeek.com/krea2-expressions-with-muscle-prompting.html",
"title": "Krea2 expressions with muscle prompting ",
"summary": "Krea2 face expressions with muscle prompt descriptors: Happiness Muscles involved: Zygomaticus major (pulls mouth corners up and out), orbicularis oculi (raises cheeks and creates \"crow's…",
"content_html": "<p>Krea2 face expressions with muscle prompt descriptors:</p>\n<ol>\n<li>\n<p>Happiness</p>\n</li>\n</ol>\n<p>Muscles involved: Zygomaticus major (pulls mouth corners up and out), orbicularis oculi (raises cheeks and creates \"crow's feet\" around the eyes).</p>\n<p>Description: A genuine (Duchenne) smile lifts both the lips and the outer corners of the eyes.</p>\n<p>2. Sadness</p>\n<p>Muscles involved: Corrugator supercilii (pulls brows inward and downward), depressor anguli oris (pulls lip corners down), mentalis (wrinkles the chin and protrudes the lower lip).</p>\n<p>Description: Characterized by the inner eyebrows lifting and drawing together, drooping eyelids, and the edges of the mouth turning downward.</p>\n<p>3. Anger</p>\n<p>Muscles involved: Corrugator supercilii & procerus (lower brows and pull them together), orbicularis oculi (tightens eyelids), orbicularis oris (tightens and thins the lips).</p>\n<p>Description: Eyebrows are pulled downward and together, the upper eyelids are raised, the eyes narrow, and lips are often pressed tightly together.</p>\n<p>4. Fear</p>\n<p>Muscles involved: Frontalis & corrugator (raise and pull brows together), levator palpebrae superioris (wide opening of upper eyelids), risorius (stretches the lips horizontally).</p>\n<p>Description: Eyebrows pull upwards and together, upper eyelids raise to expose the white of the eyes, and the lips stretch outward horizontally.</p>\n<p>5. Disgust</p>\n<p>Muscles involved: Levator labii superioris (raises the upper lip), nasalis (wrinkles the nose), depressor anguli oris (pulls lip corners down).</p>\n<p>Description: The nose wrinkles, the upper lip elevates, and the cheeks are raised.</p>\n<p>6. Surprise</p>\n<p>Muscles involved: Frontalis (raises the eyebrows), levator palpebrae superioris (widens upper lids), jaw drops (mandible depressor muscles).</p>\n<p>Description: Eyebrows curve upwards, eyes widen significantly, and the jaw drops open naturally.</p>\n<p>7. Contempt</p>\n<p>Muscles involved: Zygomaticus major & risorius (tightens the corner of the lip).</p>\n<p>Description: The only asymmetrical emotion; usually presents as a unilateral tightening and pulling back of a single corner of the mouth (an arrogant smirk).</p>\n<p>Sample Prompt: wide-angle lens distortion, forced perspective, face close to the lens</p>\n<p>soft diffused lighting, cinematic light halation</p>\n<p>subject: dynamic close-up shot of a french brunette woman with a diamonds ornate royal crown and red medieval dress</p>\n<p>Expression: blink scream with visible teeth</p>\n<p>Facial Muscle: blinking left eye with brow lowerer and nose wrinkler</p>\n<p>style:luminous photographic aesthetics defined by strong backlighting, radiant edge illumination, subtle translucency effects, atmospheric depth, graceful tonal transitions, and a heightened sense of visual separation, creating elegant and emotionally evocative imagery through carefully controlled exposure, naturalistic light diffusion, and refined portrait craftsmanship, reminiscent of Peter Lindbergh and Paolo Roversi, inspired by Vogue editorials and In the Mood for Love.</p>\n<p> </p>\n<p>https://www.reddit.com/r/StableDiffusion/comments/1v10uub/krea2_expressions_with_muscle_prompting/</p>",
"author": {
"name": "admin"
},
"tags": [
"krea2",
"comfyui"
],
"date_published": "2026-07-19T17:40:22-04:00",
"date_modified": "2026-08-11T05:38:36-04:00"
},
{
"id": "https://graphicdesigngeek.com/comfyu-tutorials.html",
"url": "https://graphicdesigngeek.com/comfyu-tutorials.html",
"title": "comfyui tutorials",
"summary": "https://docs.comfy.org/tutorials/",
"content_html": "<p>https://docs.comfy.org/tutorials/</p>",
"author": {
"name": "admin"
},
"tags": [
"comfyui workflow",
"comfyui"
],
"date_published": "2026-07-17T19:05:26-04:00",
"date_modified": "2026-08-11T05:38:14-04:00"
}
]
}