AI Toolkit MiniMax H3 LoRA training: what the Distillation Handling Method setting does
In AI Toolkit 0.13.23, the Distillation Handling Method box for MiniMax H3 writes two things into your job: an assistant LoRA path (the training adapter) and a contrastive guidance flag with a target of 3.5. The default since 2026-09-24 is the adapter alone, and the separate Do Differential Guidance checkbox only takes effect when contrastive guidance is on.
- In AI Toolkit 0.13.23 (commit
ecee894, 2026-09-27), the default Distillation Handling Method for MiniMax H3 LoRA training is Training Adapter alone, as it has been since commit60d0c28of 2026-09-24. - The MiniMax H3 training adapter in AI Toolkit is a frozen LoRA that is active during training steps, switched off for preview samples and not written into the LoRA you save. The Ref2VA adapter v2 is rank 32 and 310 MB.
- Contrastive Guidance in AI Toolkit adds one extra no-gradient forward pass with a blank prompt on every training step and moves the target to
uncond + 3.5 × (target − uncond). - In AI Toolkit at commit
ecee894, the Differential Guidance block sits inside the contrastive guidance branch ofSDTrainer.calculate_loss, so the checkbox has no effect with Training Adapter alone. - One MiniMax H3 Ref2VA style LoRA trained with the adapter alone on an RTX 5090 (32 GB) showed no artifacts at its step 4,250 checkpoint when rendered at 10 steps in ComfyUI. One training run, three prompts, one seed.

What does the Distillation Handling Method box set in the config?
It writes ordinary config keys and stores nothing of its own: model.assistant_lora_path for the training adapter, and train.do_guidance_loss with a train.guidance_loss_target of 3.5 for contrastive guidance. MiniMax H3 is guidance-distilled. The toolkit's help text says that training on it directly makes the distillation break down, and offers two ways around that: a training adapter and contrastive guidance.
AI Toolkit is Ostris's open-source trainer. Everything below is read from its code at commit ecee894 (version 0.13.23, 2026-09-27). The repo took 94 commits in September 2026 up to that one, so check the commit before you rely on a line number.
The box is defined in extensions_built_in/diffusion_models/ui.tsx. onChange writes the keys and getValue works the selection back out from them, so editing the keys by hand moves the box too.
| Option in the UI | train.do_guidance_loss | train.guidance_loss_target | model.assistant_lora_path |
|---|---|---|---|
| Training Adapter (default) | removed | removed | ostris/minimax_h3_training_adapter/minimax_h3_ref2va_training_adapter_v2.safetensors |
| Contrastive Guidance | true | 3.5, unless a target is already set | removed |
| Contrastive Guidance + Training Adapter | true | 3.5, unless a target is already set | same adapter path as the default |
| None | removed | removed | removed |
Ref2VA has a fifth option, D-OPSD, which clears all three keys and sets model.model_kwargs.dopsd instead. The help text describes it as self-distillation with a no-gradient teacher pass. I have not run it.
The text and image-to-video architecture minimax_h3 has the same four options with minimax_h3_training_adapter_v3.safetensors. The FastH3 8-Step V2 architecture only offers Training Adapter or None, with its own adapter file.
What is the training adapter?
It is a LoRA file that Ostris publishes in the Hugging Face repo ostris/minimax_h3_training_adapter. The loader's docstring names a de-distillation adapter as its example of an assistant LoRA. MinimaxH3Model.load_training_adapter loads it as a second LoRA network next to the one you are training.
- Frozen. Gradients are switched off for every parameter, so it shapes what your LoRA trains against and is never trained itself.
- Live, not merged. It stays a separate module at multiplier 1.0. The H3 transformer is loaded pre-quantized, and the docstring says a merge would resample every int8 scale.
- Off for previews.
BaseModel.generate_imagesdeactivates it before sampling and reactivates it afterwards. - Not in your file.
BaseSDTrainProcess.savewrites only the network being trained. My rank 16 checkpoints are 155 MB each. The Ref2VA adapter v2 on its own is rank 32 and 310 MB.
The adapter downloads itself on first use into loras/training_adapters under the toolkit's models folder. On 2026-09-30 the adapter repo held the weight files without a model card, so how the adapter was trained is not something I can read there. The repo also holds a Ref2VA v3 file that the toolkit at this commit does not reference.
What does contrastive guidance do on each training step?
It runs one extra forward pass with a blank prompt, without gradients, and pushes the training target 3.5 times as far from that prediction. In SDTrainer.calculate_loss, when do_guidance_loss is true:
- The trainer takes the blank-prompt embeddings it cached before the training loop started.
- It runs one extra forward pass on the same noisy latents with that blank prompt, without gradients.
- It replaces the target with an extrapolation away from that blank-prompt prediction.
- For a joint audio model like H3 it extrapolates the audio target the same way.
target = uncond + 3.5 * (target - uncond)
That is the classifier-free guidance formula applied to the training target. As I read it, the LoRA trains towards a guided prediction in a single pass, which is how a guidance-distilled model runs.
The cost is the second forward pass on every step. The help text calls the adapter faster but still able to break down over a long run, and contrastive guidance slower but less likely to. I have no controlled timing of the two, so there is no number here.
Differential Guidance needs contrastive guidance
Do Differential Guidance is its own checkbox in the Advanced card of the job form, with a default scale of 3. It amplifies the gap between the current prediction and the target:
target = pred + 3 * (target - pred)
At this commit the block is nested inside the if self.train_config.do_guidance_loss: branch of calculate_loss (lines 775 to 866), and it is the only place in the Python code that reads the flag. It has been in that branch since the commit that added it on 2025-11-10.
So with Training Adapter alone, or with None, ticking the checkbox changes nothing in the loss. With either contrastive option it is active, and at scale 3 it triples the gap. For the default MSE loss that makes the video loss 9 times larger for the same prediction. The audio target is not touched by this block.
How the default has changed
| Commit date | Commit | Default |
|---|---|---|
| 2026-08-04 | 183433a | Contrastive guidance |
| 2026-08-06 | 71625d1 | Training adapter (alpha version) |
| 2026-08-12 | 7eb65b8 | Contrastive guidance, target 3.5. The Distillation Handling Method box is added. |
| 2026-08-16 | 2042481, b982a03 | Contrastive guidance + training adapter, including the new Ref2VA adapter |
| 2026-09-24 | 60d0c28 | Training adapter alone, with new adapter files (v3, and v2 for Ref2VA). Version 0.13.23. |
The message of the 2026-09-24 commit says the new adapters "no longer need contrastive guidance". Five defaults in seven weeks is the reason this post carries a commit hash.
What I ran
One run on AI Toolkit 0.13.23: a 2.5D style LoRA for the course, in the look described in the 2.5D style post. Architecture minimax_h3_ref2va, Training Adapter alone with the v2 adapter, rank 16 and alpha 16, learning rate 1e-4, adamw8bit, resolution settings 256 and 512 px, clips bucketed at 73 frames with audio, a save every 500 steps at first and every 250 later in the run. It started with a 3,500-step target that I raised to 5,000, and I stopped it at step 4,452.
Do Differential Guidance was ticked at scale 3 in that job. Going by the code above, it did nothing.
I tested the step 3,500 and step 4,250 checkpoints in ComfyUI with my Stubelius Ultimate H3 workflow: 10 steps, euler and beta, an 8-step lightx2v turbo LoRA at 0.75, 736×1280, 124 frames at 24 fps (about 5 s), LoRA strength 1.0, no upscale. Each checkpoint rendered the same three prompts, one reference image per prompt, all on seed 42. None of the six clips showed artifacts, and every one performed the whole prompted action.
This run logged a median total loss of 0.43 over 4,451 logged steps. An earlier run with Contrastive Guidance + Training Adapter and Differential Guidance active logged a median of 29.5 over 3,999 steps. The two are not comparable. The targets differ, so 30 in one mode and 0.4 in the other says nothing about which LoRA is better.
Tested 2026-09-29 on an RTX 5090 (32 GB), 96 GB RAM, Windows 11. Trained in the last week of September 2026 with AI Toolkit 0.13.23. Code read on 2026-09-30. One training run, three prompts per checkpoint, one seed.
Which option should you pick?
- Start with the default. Training Adapter alone is what the toolkit ships at this commit, and it runs one forward pass per step instead of two. My run gave me no reason to change it up to step 4,250. I have no data beyond that.
- Judge it on the previews. They render at guidance scale 1 with the adapter switched off, which is how your LoRA will be used. Then test the checkpoints in your real ComfyUI workflow.
- Save often. The help text gives no step count for when a long run breaks down. Frequent saves let you step back to an earlier checkpoint.
- Check the Differential Guidance box before adding contrastive guidance. If it is ticked, switching to a contrastive option changes the target twice over.
None of these settings rescues a weak dataset. The rules I use for a character set are in the dataset post, and LoRA training is covered in Part 3 of ComfyUI From Zero.
Sources and files
- AI Toolkit by Ostris, the trainer this post reads (MIT licence).
ui.tsxat commitecee894: the Distillation Handling Method options and their help text.SDTrainer.pylines 775 to 866 at the same commit: contrastive guidance and the Differential Guidance block.- Commit
60d0c28: the switch to the adapter-only default. - MiniMax H3 model card on Hugging Face.
- Stubelius Ultimate H3, the ComfyUI workflow I tested the checkpoints in.
- On this site: the 2.5D style LoRA post, the character dataset post and the course.
Common questions
Does the training adapter end up inside my LoRA file?
No. In AI Toolkit 0.13.23 the adapter is a separate frozen LoRA network, and the save function writes only the network being trained. My rank 16 MiniMax H3 checkpoints are 155 MB each, while the Ref2VA adapter v2 alone is 310 MB.
Do I need to load the training adapter in ComfyUI?
No. The toolkit switches the adapter off for its own preview samples, so the previews already show your LoRA on the plain model. My ComfyUI tests loaded the trained LoRA and a turbo LoRA, and no adapter.
Why is the logged loss so much higher with contrastive guidance than without?
Because the target is different. Contrastive guidance extrapolates the target away from a blank-prompt prediction, and Differential Guidance at scale 3 multiplies the video part of an MSE loss by 9 on top of that. In my two runs the median logged loss was 29.5 with both active and 0.43 with the adapter alone. Loss values are only comparable between runs that use the same option.
STUUBZZZ


