MiniMax H3 · ComfyUI · Stubelius

MiniMax H3 ComfyUI workflow: Speed, Hybrid and Quality modes timed on an RTX 5090

Published Updated 11 min readby
AI-generated frame: hot-air balloons over a rocky valley at sunrise Watch the video · I Built the Ultimate MiniMax H3 Workflow in ComfyUI Free, 4K, 60 FPS

Stubelius Ultimate H3 is my free MiniMax H3 workflow for ComfyUI with three presets: on an RTX 5090, one 8-second 16:9 clip finished at 1080p took 1:29 in Speed (8 steps) and 2:33 in Hybrid (10 steps). Quality (20 steps) took 5:55 at its native 768p, or about 50 minutes with the 2× latent polish to 1080p.

Short answer
  • Speed preset in Stubelius Ultimate H3: 8 steps, res_multistep / simple, 0.4 megapixels, turbo LoRA at 1.0. An 8 s MiniMax H3 clip took 1:07 to render plus 0:22 to finish at 1080p on an RTX 5090 (32 GB), peaking at 25.5 GB of VRAM.
  • Hybrid preset in Stubelius Ultimate H3: 10 steps, euler / beta, 0.5 megapixels, turbo LoRA at 0.75. An 8 s MiniMax H3 clip took 2:09 to render plus 0:24 to finish at 1080p on an RTX 5090 (32 GB), peaking at 28.2 GB of VRAM.
  • Quality preset in Stubelius Ultimate H3: 20 steps, euler / beta, 0.98 megapixels, no turbo LoRA. An 8 s MiniMax H3 clip took 5:48 to render on an RTX 5090 (32 GB), peaking at 30.8 GB of VRAM. The optional 2× latent polish to 1080p added 43:55 in one run.
  • Seed hunting in Stubelius Ultimate H3 renders 1 to 4 full takes with sound. WINNER 0 on the Finish node holds after the seeds, and setting WINNER to 1, 2, 3 or 4 re-runs only the finish, from cache.
  • The pack is MIT licensed and public at github.com/stuubszzz/Stubelius-Ultimate-H3. There is no release zip: install it with ComfyUI Manager (Install via Git URL) or git clone.

MiniMax H3 is MiniMax's video model with sound: from a text prompt, a first or last frame, or reference media, it generates the picture and the audio in the same pass. Its checkpoints are an open release on Hugging Face under the MiniMax H3 Community License (model card), and ComfyUI runs it with its stock nodes.

What are the Speed, Hybrid and Quality settings?

Speed, Hybrid and Quality are the three mode presets on the Setup node. Picking one fills in the steps, sampler, scheduler, megapixels and speed LoRA strength. The widgets hold the real values, so you can still change any of them by hand.

Mode presets in Stubelius Ultimate H3 v0.8
ModeStepsSampler / schedulerMegapixelsNative render at 16:9Speed LoRA strength
Speed8res_multistep / simple0.4864 × 4801.0
Hybrid10euler / beta0.5960 × 5440.75
Quality20euler / beta0.981344 × 7680 (off)

The speed LoRA is the 8-step turbo LoRA from the Comfy-Org MiniMax H3 files. FastVideo's FastH3 preview adapter is another speed LoRA for H3, but it does nothing in ComfyUI until you convert it with the FastH3 LoRA converter. The native sizes follow from the megapixel value, rounded to multiples of 32.

How long does an 8-second MiniMax H3 clip take on an RTX 5090?

An 8-second clip took 1:29 in Speed and 2:33 in Hybrid, both finished at 1080p, and 5:55 in Quality at its native 768p. Taking Quality to 1080p with the 2× latent polish made it 49:42. I ran the same prompt and seed in all three modes, split the way a seed hunt runs: the seed render with WINNER at 0, then the finish as a second queue with the seed coming from cache.

One 8-second 16:9 text-to-video clip (a glassblower), one seed, 24 fps. Times are min:sec.
Mode and outputSeed renderPeak VRAM, seedFinishPeak VRAM, finishTotal
Speed, 1080p1:0725.5 GB0:22 (RTX VSR)3.5 GB1:29
Hybrid, 1080p2:0928.2 GB0:24 (RTX VSR)3.6 GB2:33
Quality, native 768p5:4830.8 GB0:07 (no upscale)2.6 GB5:55
Quality, 1080p5:4830.8 GB43:55 (2× latent polish)31.4 GB49:42

Tested 2026-09-25 to 2026-09-26 on an RTX 5090 (32 GB), 96 GB RAM, Windows 11, ComfyUI 0.37.0, workflow v0.8. One seed per setting and one run each, apart from the Speed seed render, which ran cold and then warm.

Times are ComfyUI's own execution time, queue wait excluded, with totals added up before rounding. Peak VRAM is the card's memory in use as reported by nvidia-smi, polled every 5 seconds. The Speed seed render in the table is the warm run. The first one, with the models loading from disk, took 2:21 and peaked at 27.8 GB.

One difference from the file list further down: the timed runs loaded a community variant of the text encoder in the same int8 format (26.4 GB, against 27.1 GB for the official file). The checkpoints, VAEs and turbo LoRA were the files listed there.

The Quality polish is the expensive step. It doubles the chosen take in latent space, from 1344 × 768 to 2688 × 1536, and re-samples it at that size (here 12 steps at strength 0.3 with the learned 2× upscaler) before the result is scaled down to 1080p. It peaked at 31.4 GB on a 32 GB card. One longer data point: a 10-second Quality clip taken to 4K with the polish, DLSS5 + Color Lock and the low-VRAM options on took 7:07 for the seed and 62:10 for the finish, peaking at 31.3 and 31.4 GB.

AI-generated MiniMax H3 frame: a glassblower shaping glowing glass at a furnace
The glassblower is the pack's example prompt and the subject of the 8-second timing clip. AI-generated.

What each node does

You work top to bottom: Setup, Models, Output, Director, Finish.

The nodes in the Stubelius Ultimate H3 pack
NodeJob
Stubelius H3 SetupHow the video is made: the mode preset, the number of seeds (1 to 4), and the settings the mode filled in.
Stubelius H3 ModelsReference (ref2va) and First/Last Frame (fl2va) checkpoints as safetensors or GGUF, text encoder, video and audio VAE, speed LoRA, two global LoRAs, attention and low-VRAM options, live preview. A checkpoint loads only when the Director needs it.
Stubelius H3 OutputWhat comes out: final resolution, frame rate, upscaler and the Quality polish settings. It feeds only Finish.
Stubelius H3 Director V2What the video is: timeline or prompt, Reference (Omni) or First/Last Frame mode, reference media, aspect ratio, duration, seed, and a LoRA box per chunk.
Stubelius H3 FinishWINNER 1 to 4, or 0 to hold. Runs the polish, RIFE and the upscale on the chosen take.
Stubelius Live PreviewShows the Director, and the Quality polish, while it samples.
Stubelius RIFE to FPS, Stubelius Color LockFrame rate conversion that keeps hard cuts clean, and a step that restores the original colours after DLSS 5.
Stubelius ThemeA colour theme for this workflow only.

A single H3 call is capped at 15 seconds here. A longer duration is split into chunks, and each chunk can carry its own LoRAs while the turbo LoRA stays global. A 12-second Hybrid clip made of two 6-second chunks with different LoRAs rendered its seed in 2:52 and finished at 1080p in 0:14.

In this workflow I tested a style LoRA that I trained in AI Toolkit. If you train your own, the AI Toolkit post explains how the trainer deals with H3's guidance distillation.

How does seed hunting with full videos work?

Every candidate seed, 1 to 4 of them, renders as a complete video with sound, and only the one you pick is finished. My older Director workflow scouts seeds as low-resolution latents instead.

  1. Set seeds to 2, 3 or 4 on Setup and WINNER to 0 on Finish, then queue. WINNER 0 means hold: the seeds render and preview, and nothing is finished.
  2. Watch the seed previews and pick one.
  3. Set WINNER to that number and queue again. The seeds come from ComfyUI's cache and only the finish runs.

Four 8-second Hybrid takes of a different prompt (a dancer) took 7:48 on the RTX 5090 and peaked at 28.7 GB.

Because Output feeds only Finish, you can change the resolution, frame rate or upscaler after the hunt without rendering the seeds again. Leave Setup, Models and the Director alone in between, or they do render again. Running an unrelated workflow in between can also push the seeds out of ComfyUI's cache.

Output sizes per mode

Final sizes each mode offers on the Output node
ModeSizes offeredHow it gets there
Speednative (about 480p), 720p, 1080p, 2K (1440p)Upscale from the render. 4K is not offered.
Hybridnative (about 540p), 720p, 1080p, 2K, 4K (2160p)Upscale from the render.
Qualitynative (about 768p), 1080p, 2K, 4KA 2× latent polish from 1080p up. 1080p and 2K are scaled down from the polished frames, 4K goes through the upscaler.

The short side lands exactly on the number, in portrait too. The upscaler is RTX VSR, or DLSS 5 followed by Color Lock, which puts the clip's original colours back (both timed, with the colour measurements). The workflow renders at 24 fps, and any higher frame rate, up to 120, is interpolated by RIFE with length and audio sync kept.

Finish only, Hybrid mode, one 8-second take (a dancer) picked from a 4-seed hunt. One run each, RTX 5090.
OutputUpscalerFrame rateFinish timePeak VRAM
720pRTX VSR24 fps0:243.0 GB
1080pRTX VSR24 fps0:533.0 GB
1080pRTX VSR60 fps (RIFE)0:565.8 GB
2K (1440p)RTX VSR24 fps1:062.8 GB
2K (1440p)DLSS5 + Color Lock24 fps4:2716.7 GB
4K (2160p)RTX VSR24 fps1:263.0 GB

Same settings, different clip: the glassblower's Hybrid 1080p finish took 0:24 and the dancer's took 0:53. These are single runs, so read them as ballpark figures.

Install in 3 steps

  1. Install the pack. In ComfyUI Manager choose Install via Git URL, paste https://github.com/stuubszzz/Stubelius-Ultimate-H3 and restart. If your Manager security level blocks Git URLs, use the commands below.
  2. Open the workflow under Workflow, Browse Templates, Stubelius-Ultimate-H3, then run Manager, Install Missing Custom Nodes for the other packs it uses.
  3. Download the models listed below, check that the Models node shows your files, and press Queue. The example, a 12.5-second glassblower shot, ships with WINNER at 0, so the first queue stops after the seed. Set WINNER to 1 and queue again for the final video.
cd ComfyUI/custom_nodes
git clone https://github.com/stuubszzz/Stubelius-Ultimate-H3

On Windows portable, install its one requirement from the ComfyUI_windows_portable folder:

python_embeded\python.exe -m pip install -r ComfyUI\custom_nodes\Stubelius-Ultimate-H3\requirements.txt

To update, use Manager, Update All, or git pull in the pack folder, then restart. The repo has no tagged release, so the main branch is the current version.

Required node packs

The Director and Refine engines inside the pack are Muse Minimax Director V1.2 by Muse Collective, carried with changes under its MIT license. The workflow also uses these packs:

Other node packs the workflow uses
PackByNeeded for
ComfyUI-VideoHelperSuiteKosinkadinkVideo outputs (needed)
Nvidia RTX nodesComfy-OrgRTX VSR upscaling (needed)
ComfyUI-KJNodeskijaiLive preview and the low-VRAM options (needed)
ComfyUI-Frame-InterpolationFannovel16RIFE for any frame rate above 24 fps (needed)
ComfyUI-H3-Motion-Context-MultiRefseitanismRecommended: smoother continuity when a video runs past one chunk
ComfyUI-MiniMaxH3_LatentUpscaler and ComfyUI-H3-Latent-Upscaler-Mamad8Tr1dae, mamad8cThe Quality polish method "learned model (2x)". Without them, pick bislerp or bicubic.
ComfyUI-DLSS5-EnhancerBlueforcerThe DLSS 5 upscaler
ComfyUI-GGUFcity96GGUF models
ComfyUI-Spectrum-MiniMax-H3xmarreThe Spectrum option
Comfyui-PlagueKind-NodesPlagueKindThe H3 cache option

Where do the MiniMax H3 model files go in ComfyUI?

The two checkpoints go in models/diffusion_models, the text encoder in models/text_encoders, the video and audio VAEs in models/vae and the turbo LoRA in models/loras.

All six are in the Comfy-Org/MiniMax-H3 repository, about 77 GB in total. Direct links are in the pack's INSTALL.txt.

Model files the example workflow expects
FolderFileSize
models/diffusion_modelsminimax_h3_ref2va_pruned_int8_convrot.safetensors21 GB
models/diffusion_modelsminimax_h3_fl2va_pruned_int8_convrot.safetensors21 GB
models/text_encodersqwen3vl_32b_minimax_h3_int8_convrot.safetensors27 GB
models/vaeminimax_h3_video_vae_fp16.safetensors5.2 GB
models/vaeminimax_h3_audio_vae_fp32.safetensors0.6 GB
models/lorasminimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors2 GB

RTX 50-series cards can swap the text encoder for qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors (16 GB). The optional live preview uses Kijai's tiny H3 VAE, taeh3.safetensors (10 MB), in models/vae_approx. RIFE downloads rife47.pth by itself the first time you pick a frame rate above 24 fps.

Hardware notes

  • Every number here comes from one RTX 5090 (32 GB) with 96 GB of RAM. The VRAM peaks are what the workflow used on that card, not minimum requirements.
  • For cards with less VRAM, the Models node has low VRAM attention and chunk feed-forward switches, and it loads GGUF checkpoints and text encoders through ComfyUI-GGUF. My one run with both switches on, the 10-second Quality clip above, still peaked at 31.3 GB on the 32 GB card, so I have no measurement of how far they lower the requirement.
  • RTX VSR needs an NVIDIA RTX card. The DLSS 5 pack's README lists RTX 30, 40 and 50 cards and says it is Windows only.
  • A finished 4K frame holds about 100 MB of RAM, so keep 4K clips short, especially at 48 or 60 fps.
  • The note inside the workflow says the turbo modes can mumble dialogue. For talking shots, use Quality.
  • If VRAM budgeting is new to you, ComfyUI From Zero covers it in Part 2, with fit tables for 8, 12, 16 and 24 GB cards.

Sources and files

Common questions

Is Stubelius Ultimate H3 free?

Yes. The nodes and the example workflow are MIT licensed in a public GitHub repo. The MiniMax H3 model is a separate download under the MiniMax H3 Community License, which you need to read for yourself.

Can I keep the older Stubelius Director or the Muse Director installed?

Yes. Ultimate H3 registers only its own nodes and uses its own server routes, so it installs next to Stubelius-Director and the Muse Director without clashing.

Will it run on a card with less than 32 GB of VRAM?

I do not know, because I have only measured it on an RTX 5090 (32 GB), where the 8-second seed renders peaked between 25.5 and 30.8 GB. The Models node has two low VRAM switches and loads GGUF files for smaller cards, but my one run with both switches on still peaked at 31.3 GB, so I cannot say how far they lower the requirement.

StuubzzzBuilds self-hosted AI video pipelines and the Stubelius nodes for ComfyUI, and teaches them in the ComfyUI From Zero course. My own tests run on one RTX 5090. About
Want the why, not just the workflow?

Learn it, fix it live, or have it made.

The ComfyUI From Zero course explains the machine from the first node to training your own LoRA, and Part 1 is free. If something is fighting you right now, bring it to a 60-minute 1-on-1.