LTX 2.5 modes in ComfyUI: what each one does and how to stop MSR duplicate characters
Watch the video · LTX 2.5 I Tested Every Mode in ComfyUI So You Don't Have To
Stubelius Ultimate LTX 2.5 makes an LTX 2.5 video eight ways besides text to video, and each mode wants its own prompt: image to video describes what happens next, IC-LoRA control the new look, custom audio what makes the sound. Against duplicate characters with Licon's MSR references, wording did more than any setting I changed: in 4-seed hunts (MSR V1, RTX 5090) a short label gave a second person in 2 of 4 seeds, a long label twins in 4 of 4, and @ref1 as the subject of every action 0 of 4.
- In 4-seed LTX 2.5 hunts with Licon's MSR LoRA V1 on an RTX 5090 (2026-10-03), a three-word slot label gave a second person in 2 of 4 seeds,
Ref strength1.5 or 25 fps gave 1 of 4, and a long description as the label gave twins in all 4. - In 4-seed LTX 2.5 MSR V1 hunts on an RTX 5090 (2026-10-03), naming
@ref1as the subject of every action, or adding "alone on the trail" and "Nobody else is around.", each gave 0 of 4 seeds with a second person. - With Lightricks' IC-LoRA Union Control on LTX 2.5,
Video Attn0.3 let go of the control video's walk, 0.65 and 1.0 looked the same, andVideo Strength0.5 barely changed the clip (one seed per setting, RTX 5090, 2026-10-03). - Retake in Stubelius Ultimate LTX 2.5 rendered the same walk for a spin prompt and a wave prompt in a region from 2.0 to 4.0 s of a 5 s LTX 2.5 clip. A region from 1.0 to 4.4 s made the wave happen (one seed per setting, RTX 5090, 2026-10-03).
Which LTX 2.5 mode does what in ComfyUI?
All eight are set up on the Director's timeline in my free Stubelius Ultimate LTX 2.5 workflow; that post has the install. The prompts follow Lightricks' prompting guide, which asks for a focused scene: a few clear characters read better than a crowded frame. That comes back with duplicates.
| Mode | Setup on the Director | The prompt describes | What my test showed |
|---|---|---|---|
| Image to video | An image block, Guide Strength 1.0 | What happens next, then the sound | First frame against the image: 29.8 dB PSNR at 1.0, 27.4 at 0.6, 21.6 at 0.3. Describing only the picture added a second woman |
| First and last frame, storyboards | Two or three image blocks, the last set to Pin to Last Frame | The change, one prompt per block | It moved from keyframe to keyframe. Mine came from one image with edits |
| Extend | Add Video with a clip's last 3 s and its sound, then a text block | Only what happens next | No seam at the join by frame-to-frame PSNR. She picked up the umbrella lying in the given frames |
| IC-LoRA control | Union Control on Models, a depth or canny video via Add IC Video | The new look, not the motion | An umbrella shape in the depth map became an umbrella I never asked for |
| Motion transfer | A DWPose video on the IC track, a start image of the new character | The new character, not the dance | One robot followed every move. Without the start image: two robots |
| MSR characters | MSR LoRA on Models, Licon MSR as reference option, a picture and short label in an @ref slot, 50 fps | @ref1 as the subject of every action | Her identity held; duplicates below |
| Custom audio | Add Audio on the audio track, Inpaint on or off | What we see making the sound | With Inpaint on, the model added a "Thank you." I never wrote. "On the drum hits" put a drum kit in the shot; naming a speaker as the source fixed it |
| Retake | Retake Mode (BETA), the clip, a region to redo | The new action | A 2 s region repeated the walk. 1.0 to 4.4 s gave a wave |
Tested 2026-10-03 on an RTX 5090 (32 GB), 96 GB RAM, Windows 11. Official LTX 2.5 distilled int8, 8 steps, cfg 1, euler_ancestral on linear_quadratic, seed 42, 1280×704, first-pass scale 1.0, 5 s at 25 fps (extend 8 s, audio 6 s, MSR 50 fps). One render per setting: PSNR where given, my eye otherwise.
How do you stop duplicate characters with MSR?
Make @ref1 the subject of every action, or say she is alone, and keep the slot label short. MSR is Licon's Multiple Subject Reference LoRA for LTX 2.5: a picture of the character goes in an @ref slot with a label, and @ref1 is swapped for that label before the prompt reaches the text encoder. My first MSR render had a second explorer behind a rock. So I ran 4-seed hunts of that shot, one change at a time, and had a person detector, YOLOX-L, count the people in each seed.
| Variant | Seeds with a second person | Which seeds |
|---|---|---|
Baseline: label "red-haired explorer woman", Ref strength 1.0, 50 fps | 2 of 4 | 1, 3 |
| A long description as the label | 4 of 4, twins | 1, 2, 3, 4 |
Ref strength 1.5 | 1 of 4 | 3 |
| 25 fps | 1 of 4 | 1 |
| "alone on the trail" and "Nobody else is around." added | 0 of 4 | none |
@ref1 as the subject of both actions | 0 of 4 | none |
Tested 2026-10-03 on an RTX 5090 (32 GB), 96 GB RAM, Windows 11. Licon MSR V1 at LoRA strength 1.0, one character slot, a text block, 5 s at 1280×704.
Give every action a subject. The baseline prompt was "@ref1 climbs over mossy rocks along a misty mountain trail, then stops on a ledge and looks out over the valley. Wind in the trees and distant birds." Its second action has no subject, and in seed 1 the double stood on the rocks. "Then @ref1 stops on a ledge" left one explorer in all four seeds. So did adding "alone on the trail" after @ref1 and "Nobody else is around." before the sound.


Keep the label short. The label is what the model reads in place of @ref1. Swapping "red-haired explorer woman" for a full description of her hair, freckles, jacket, scarf, satchel, trousers and boots gave twins in every seed. The picture already carries all that. My guess is that, written out again, it reads like a second person.

Ref strength and frame rate help less. Ref strength 1.5 and 25 fps each took the doubles from two seeds to one, a one-seed difference too small to call. At other rates the Director shows an "MSR 50 recommended" hint whose tooltip says Licon MSR is trained at 50 fps, so I keep 50.
Not every double is obvious: at Ref strength 1.0 and 1.5, seed 3's double is a small figure far down the valley. The long-label twins are in plain view.
Outside MSR. Other modes added extra figures too, one seed each: the picture-only image-to-video prompt, Video Attn 0.3, pose transfer without a start image and the drum-hits prompt. Check every mode for a second person.
Later note, 2026-10-07. These counts are Licon's original LTX-2.5-Licon-MSR-V1.safetensors, as in the video. Licon has since posted LTX-2.5-Licon-MSR-V2.safetensors in the same repo. I have not rerun this table on V2. My only V2 data is one later 4-seed hunt of a different shot, with a second person in one seed: a single observation, not a comparison with V1.
Does MSR work with a start image or a pose video?
Yes, since a fix on 2026-10-02: before it, MSR crashed when the timeline held an image. From an empty trail, @ref1 walked in. From a frame of her seen from behind, she turned to the camera and MSR kept her face. With a pose video and no start image, MSR gave two explorers next to a dancer in white. With her edited into the first frame and "alone" in the prompt, it gave one.
Which IC-LoRA setting matters: Video Attn or Video Strength?
Video Attn. Lightricks' IC-LoRA Union Control is one file for depth, canny and pose. The LTX 2.3 file, ltx-2.3-22b-ic-lora-union-control-ref0.5.safetensors, ran on 2.5 at strength 1.0. Select the clip on the IC-LoRA Video track for Video Strength (default 1.0) and Video Attn (default 0.65), how hard the model listens to the control video. On a depth video restyled to winter, 0.3 let go: a different walk, and someone far behind her. 0.65 and 1.0 looked the same, and Video Strength 0.5 barely changed anything.
Why does Retake render the same motion again?
In my test, the region was too short. Retake Mode (BETA) re-renders a region of a finished clip and keeps the frames outside it. On my 5-second rain clip I asked a region from 2.0 to 4.0 s for a spin, then a wave. Both times she walked again: inside the region the two takes were 32 to 42 dB PSNR apart, nearly identical. The kept frames show her walking, and two seconds left no room to do something new and get back. A region from 1.0 to 4.4 s gave a clear wave. Adding the wave to the global prompt helped little.
How long does a seed hunt take on an RTX 5090?
A full-size 4-seed hunt of the image-to-video shot, winner finished as rendered, took 147 s. Four at first-pass scale 0.5, held with WINNER 0, took 56.5 s, and finishing the winner from cache, a latent ×2 upscale and a 4-step refine at 0.35, took 26.5 s. What half size costs in detail is in the full-size first pass post. Single seeds took about 42 s plain, 51 to 59 s with IC-LoRA, 88 to 125 s with MSR at 50 fps and 155 to 165 s with MSR plus pose. Why the pack sizes its decode to the free VRAM is in the VAE decode post.
Tested 2026-10-03 on an RTX 5090 (32 GB), 96 GB RAM, Windows 11. 5 s at 1280×704, ComfyUI's execution time from its history. One run per hunt; the single seeds are approximate ranges.
Test setup and limits
- Duplicates: one shot, one reference image and four seeds per variant, 24 in six hunts, one change each from the baseline. Read the counts as directions, not rates.
- Seeds 1 and 3 ran
euler_ancestral, 2 and 4euler. Outside the long-label row every double came on seed 1 or 3. Seed and sampler change together, so I cannot tell which it follows. The sampler settings post explains the pair, and the free Part 1 of ComfyUI From Zero covers what a sampler does. - The detector (YOLOX-L) checks sampled frames, skips anyone under 4% of the frame height and counts a seed once if any checked frame holds two or more people.
- The MSR plus pose redo and the music redo each changed more than one thing: both also added "alone".
- Other modes: one seed (42) per setting, two for the music redo. Guide Strength, the extend seam and the retake takes are PSNR; the rest is my eye on playback.
- Not covered: plain text to video (in the Director post), Ghost Mask, chunk render, low-VRAM and GGUF setups. One RTX 5090 only.
Sources and files
- LTX 2.5 I Tested Every Mode in ComfyUI So You Don't Have To: the video this post belongs to, at the duplicate tests (16:11). Chapters: IC-LoRA control (11:01), characters with MSR (14:24) and Retake (20:02).
- stuubszzz/Stubelius-Ultimate-LTX2.5: the pack and workflow I tested.
- LiconStudio/LTX-2.5-Multiple-Subject-Reference: Licon's MSR LoRA for LTX 2.5 (V1 here, V2 posted since), and liconstudio/ComfyUI-LTX2.5-MSR, Licon's MSR nodes.
- Lightricks/LTX-2.3-22b-IC-LoRA-Union-Control, the Lightricks/LTX-2.5 model card and Lightricks' LTX prompting guide.
- WhatDreamsCost's LTX Director: the timeline my Director is built on.
- YOLOX-L, the person detector: the
yolox_l.onnxthat comfyui_controlnet_aux downloads for DWPose, which also made the pose video. - On this site: the Stubelius Ultimate LTX 2.5 guide, LTX 2.5 sampler settings, the full-size first pass and VAE decode flicker and VRAM on Windows.
Common questions
How do I stop LTX 2.5 MSR from duplicating my character?
Make @ref1 the subject of every action, or say she is alone, and keep the slot label short. In 4-seed hunts with Licon MSR V1 on an RTX 5090, each of the two prompt changes gave 0 of 4 seeds with a second person, against 2 of 4 for the baseline and 4 of 4 with a long label.
Does the LTX 2.3 IC-LoRA Union Control work on LTX 2.5?
It did in my tests: at strength 1.0 it drove depth, canny and pose videos on the LTX 2.5 distilled checkpoint (one seed each, RTX 5090). Start at the defaults, Video Attn 0.65 and Video Strength 1.0.
Why does Finish skip the refine steps at first-pass scale 1.0?
There is no second pass to run them in. Refine steps belong to the refine after the latent ×2 upscale, which only runs when seeds render below full size. At 1.0 the 8 steps were the whole render, so Finish takes the seed as you saw it.
STUUBZZZ


