LTX 2.5 · ComfyUI · Character consistency

LTX 2.5 modes in ComfyUI: what each one does and how to stop MSR duplicate characters

Published 11 min readby
2 of 4 → 0 of 4: seeds with a second person: baseline vs @ref1 in every action Watch the video · LTX 2.5 I Tested Every Mode in ComfyUI So You Don't Have To

Stubelius Ultimate LTX 2.5 makes an LTX 2.5 video eight ways besides text to video, and each mode wants its own prompt: image to video describes what happens next, IC-LoRA control the new look, custom audio what makes the sound. Against duplicate characters with Licon's MSR references, wording did more than any setting I changed: in 4-seed hunts (MSR V1, RTX 5090) a short label gave a second person in 2 of 4 seeds, a long label twins in 4 of 4, and @ref1 as the subject of every action 0 of 4.

Short answer
  • In 4-seed LTX 2.5 hunts with Licon's MSR LoRA V1 on an RTX 5090 (2026-10-03), a three-word slot label gave a second person in 2 of 4 seeds, Ref strength 1.5 or 25 fps gave 1 of 4, and a long description as the label gave twins in all 4.
  • In 4-seed LTX 2.5 MSR V1 hunts on an RTX 5090 (2026-10-03), naming @ref1 as the subject of every action, or adding "alone on the trail" and "Nobody else is around.", each gave 0 of 4 seeds with a second person.
  • With Lightricks' IC-LoRA Union Control on LTX 2.5, Video Attn 0.3 let go of the control video's walk, 0.65 and 1.0 looked the same, and Video Strength 0.5 barely changed the clip (one seed per setting, RTX 5090, 2026-10-03).
  • Retake in Stubelius Ultimate LTX 2.5 rendered the same walk for a spin prompt and a wave prompt in a region from 2.0 to 4.0 s of a 5 s LTX 2.5 clip. A region from 1.0 to 4.4 s made the wave happen (one seed per setting, RTX 5090, 2026-10-03).

Which LTX 2.5 mode does what in ComfyUI?

All eight are set up on the Director's timeline in my free Stubelius Ultimate LTX 2.5 workflow; that post has the install. The prompts follow Lightricks' prompting guide, which asks for a focused scene: a few clear characters read better than a crowded frame. That comes back with duplicates.

Eight LTX 2.5 modes: the setup, what the prompt describes and what one render per setting showed
ModeSetup on the DirectorThe prompt describesWhat my test showed
Image to videoAn image block, Guide Strength 1.0What happens next, then the soundFirst frame against the image: 29.8 dB PSNR at 1.0, 27.4 at 0.6, 21.6 at 0.3. Describing only the picture added a second woman
First and last frame, storyboardsTwo or three image blocks, the last set to Pin to Last FrameThe change, one prompt per blockIt moved from keyframe to keyframe. Mine came from one image with edits
ExtendAdd Video with a clip's last 3 s and its sound, then a text blockOnly what happens nextNo seam at the join by frame-to-frame PSNR. She picked up the umbrella lying in the given frames
IC-LoRA controlUnion Control on Models, a depth or canny video via Add IC VideoThe new look, not the motionAn umbrella shape in the depth map became an umbrella I never asked for
Motion transferA DWPose video on the IC track, a start image of the new characterThe new character, not the danceOne robot followed every move. Without the start image: two robots
MSR charactersMSR LoRA on Models, Licon MSR as reference option, a picture and short label in an @ref slot, 50 fps@ref1 as the subject of every actionHer identity held; duplicates below
Custom audioAdd Audio on the audio track, Inpaint on or offWhat we see making the soundWith Inpaint on, the model added a "Thank you." I never wrote. "On the drum hits" put a drum kit in the shot; naming a speaker as the source fixed it
RetakeRetake Mode (BETA), the clip, a region to redoThe new actionA 2 s region repeated the walk. 1.0 to 4.4 s gave a wave

Tested 2026-10-03 on an RTX 5090 (32 GB), 96 GB RAM, Windows 11. Official LTX 2.5 distilled int8, 8 steps, cfg 1, euler_ancestral on linear_quadratic, seed 42, 1280×704, first-pass scale 1.0, 5 s at 25 fps (extend 8 s, audio 6 s, MSR 50 fps). One render per setting: PSNR where given, my eye otherwise.

How do you stop duplicate characters with MSR?

Make @ref1 the subject of every action, or say she is alone, and keep the slot label short. MSR is Licon's Multiple Subject Reference LoRA for LTX 2.5: a picture of the character goes in an @ref slot with a label, and @ref1 is swapped for that label before the prompt reaches the text encoder. My first MSR render had a second explorer behind a rock. So I ran 4-seed hunts of that shot, one change at a time, and had a person detector, YOLOX-L, count the people in each seed.

Seeds with a second person in 4-seed LTX 2.5 MSR hunts, one change per row from the baseline, the same four seeds each time
VariantSeeds with a second personWhich seeds
Baseline: label "red-haired explorer woman", Ref strength 1.0, 50 fps2 of 41, 3
A long description as the label4 of 4, twins1, 2, 3, 4
Ref strength 1.51 of 43
25 fps1 of 41
"alone on the trail" and "Nobody else is around." added0 of 4none
@ref1 as the subject of both actions0 of 4none

Tested 2026-10-03 on an RTX 5090 (32 GB), 96 GB RAM, Windows 11. Licon MSR V1 at LoRA strength 1.0, one character slot, a text block, 5 s at 1280×704.

Give every action a subject. The baseline prompt was "@ref1 climbs over mossy rocks along a misty mountain trail, then stops on a ledge and looks out over the valley. Wind in the trees and distant birds." Its second action has no subject, and in seed 1 the double stood on the rocks. "Then @ref1 stops on a ledge" left one explorer in all four seeds. So did adding "alone on the trail" after @ref1 and "Nobody else is around." before the sound.

AI-generated LTX 2.5 frame in a misty forest: a red-haired explorer with a red backpack climbs mossy rocks while a smaller copy of her stands on top of the rocks, both marked by gold detection boxes
Baseline, seed 1. The prompt's second action, “then stops on a ledge”, has no subject, and a second explorer stands on the rocks. The gold boxes are the people the detector counted.
AI-generated LTX 2.5 frame in a misty pine forest: one red-haired explorer with a backpack walks up a stone path beside a mossy rock wall, marked by one gold detection box
Seed 1 again, with “then @ref1 stops on a ledge”: no checked frame held a second person. Not the same moment as the frame above.

Keep the label short. The label is what the model reads in place of @ref1. Swapping "red-haired explorer woman" for a full description of her hair, freckles, jacket, scarf, satchel, trousers and boots gave twins in every seed. The picture already carries all that. My guess is that, written out again, it reads like a second person.

AI-generated LTX 2.5 frame on a misty mountain trail: the same red-haired explorer in a green field jacket and cream scarf appears twice, side by side, each marked by a gold detection box
Seed 2 of the long-label hunt. With a full description as the slot label, all four seeds had twins. The gold boxes are the people the detector counted.

Ref strength and frame rate help less. Ref strength 1.5 and 25 fps each took the doubles from two seeds to one, a one-seed difference too small to call. At other rates the Director shows an "MSR 50 recommended" hint whose tooltip says Licon MSR is trained at 50 fps, so I keep 50.

Not every double is obvious: at Ref strength 1.0 and 1.5, seed 3's double is a small figure far down the valley. The long-label twins are in plain view.

Outside MSR. Other modes added extra figures too, one seed each: the picture-only image-to-video prompt, Video Attn 0.3, pose transfer without a start image and the drum-hits prompt. Check every mode for a second person.

Later note, 2026-10-07. These counts are Licon's original LTX-2.5-Licon-MSR-V1.safetensors, as in the video. Licon has since posted LTX-2.5-Licon-MSR-V2.safetensors in the same repo. I have not rerun this table on V2. My only V2 data is one later 4-seed hunt of a different shot, with a second person in one seed: a single observation, not a comparison with V1.

Does MSR work with a start image or a pose video?

Yes, since a fix on 2026-10-02: before it, MSR crashed when the timeline held an image. From an empty trail, @ref1 walked in. From a frame of her seen from behind, she turned to the camera and MSR kept her face. With a pose video and no start image, MSR gave two explorers next to a dancer in white. With her edited into the first frame and "alone" in the prompt, it gave one.

Which IC-LoRA setting matters: Video Attn or Video Strength?

Video Attn. Lightricks' IC-LoRA Union Control is one file for depth, canny and pose. The LTX 2.3 file, ltx-2.3-22b-ic-lora-union-control-ref0.5.safetensors, ran on 2.5 at strength 1.0. Select the clip on the IC-LoRA Video track for Video Strength (default 1.0) and Video Attn (default 0.65), how hard the model listens to the control video. On a depth video restyled to winter, 0.3 let go: a different walk, and someone far behind her. 0.65 and 1.0 looked the same, and Video Strength 0.5 barely changed anything.

Why does Retake render the same motion again?

In my test, the region was too short. Retake Mode (BETA) re-renders a region of a finished clip and keeps the frames outside it. On my 5-second rain clip I asked a region from 2.0 to 4.0 s for a spin, then a wave. Both times she walked again: inside the region the two takes were 32 to 42 dB PSNR apart, nearly identical. The kept frames show her walking, and two seconds left no room to do something new and get back. A region from 1.0 to 4.4 s gave a clear wave. Adding the wave to the global prompt helped little.

How long does a seed hunt take on an RTX 5090?

A full-size 4-seed hunt of the image-to-video shot, winner finished as rendered, took 147 s. Four at first-pass scale 0.5, held with WINNER 0, took 56.5 s, and finishing the winner from cache, a latent ×2 upscale and a 4-step refine at 0.35, took 26.5 s. What half size costs in detail is in the full-size first pass post. Single seeds took about 42 s plain, 51 to 59 s with IC-LoRA, 88 to 125 s with MSR at 50 fps and 155 to 165 s with MSR plus pose. Why the pack sizes its decode to the free VRAM is in the VAE decode post.

Tested 2026-10-03 on an RTX 5090 (32 GB), 96 GB RAM, Windows 11. 5 s at 1280×704, ComfyUI's execution time from its history. One run per hunt; the single seeds are approximate ranges.

Test setup and limits

  • Duplicates: one shot, one reference image and four seeds per variant, 24 in six hunts, one change each from the baseline. Read the counts as directions, not rates.
  • Seeds 1 and 3 ran euler_ancestral, 2 and 4 euler. Outside the long-label row every double came on seed 1 or 3. Seed and sampler change together, so I cannot tell which it follows. The sampler settings post explains the pair, and the free Part 1 of ComfyUI From Zero covers what a sampler does.
  • The detector (YOLOX-L) checks sampled frames, skips anyone under 4% of the frame height and counts a seed once if any checked frame holds two or more people.
  • The MSR plus pose redo and the music redo each changed more than one thing: both also added "alone".
  • Other modes: one seed (42) per setting, two for the music redo. Guide Strength, the extend seam and the retake takes are PSNR; the rest is my eye on playback.
  • Not covered: plain text to video (in the Director post), Ghost Mask, chunk render, low-VRAM and GGUF setups. One RTX 5090 only.

Sources and files

Common questions

How do I stop LTX 2.5 MSR from duplicating my character?

Make @ref1 the subject of every action, or say she is alone, and keep the slot label short. In 4-seed hunts with Licon MSR V1 on an RTX 5090, each of the two prompt changes gave 0 of 4 seeds with a second person, against 2 of 4 for the baseline and 4 of 4 with a long label.

Does the LTX 2.3 IC-LoRA Union Control work on LTX 2.5?

It did in my tests: at strength 1.0 it drove depth, canny and pose videos on the LTX 2.5 distilled checkpoint (one seed each, RTX 5090). Start at the defaults, Video Attn 0.65 and Video Strength 1.0.

Why does Finish skip the refine steps at first-pass scale 1.0?

There is no second pass to run them in. Refine steps belong to the refine after the latent ×2 upscale, which only runs when seeds render below full size. At 1.0 the 8 steps were the whole render, so Finish takes the seed as you saw it.

StuubzzzBuilds self-hosted AI video pipelines and the Stubelius nodes for ComfyUI, and teaches them in the ComfyUI From Zero course. My own tests run on one RTX 5090. About
Want the why, not just the workflow?

Learn it, fix it live, or have it made.

The ComfyUI From Zero course explains the machine from the first node to training your own LoRA, and Part 1 is free. If something is fighting you right now, bring it to a 60-minute 1-on-1.