Consistent character LoRA dataset: the five views and the rules
A character consistency LoRA needs 15 to 30 images of one subject, built around five views: front, side, three-quarter, back and a face detail, all in one outfit and one lighting setup. Most "my character keeps changing" problems are dataset problems. This is the sheet I build before any character LoRA, and why each column is there.
- A character consistency LoRA dataset is 15 to 30 images of one subject. Fewer tends not to generalise, and more starts averaging in mistakes.
- The core of the set is a five-view turnaround: front, side, three-quarter, back and a face detail, with the same outfit, hairstyle, lighting and neutral background in every image.
- Captions describe what changes between images (angle, expression, framing). The character's name token is written the same way in every caption.
- A LoRA for a video model can be trained on clips instead of stills: the MiniMax H3 2.5D style LoRA from Part 3 of my ComfyUI course used 50 clips of about 3 seconds (73 frames at 24 fps), each with its own caption file, trained with AI Toolkit on one RTX 5090 (32 GB).
- Rule-of-thumb set sizes from Part 3 of the course: 15 to 30 images for a face, 20 to 50 for a style, 10 to 20 for an object and 30 to 60 for a concept.
Which five views does a character LoRA dataset need?
Front, side, three-quarter, back and a face detail, all in one outfit and one lighting setup. What each view is for:
FrontIdentity anchor. Face, proportions, outfit silhouette.
SideNose and jaw profile, hair volume, posture.
3/4The angle most video frames actually land on.
BackHair, outfit construction, so turns do not invent new clothes.
Face detailSkin, eyes, texture of the fabric at close range.Rules for the set
- Lock everything except the angle on the sheet. One outfit, one hairstyle, one lighting setup, one neutral background. If two things change between images, the LoRA cannot tell which one is "the character". The cost: a LoRA learns whatever its images share, so a set made only of the sheet carries that light and backdrop too. Vary them in the extra images if the character has to work in other light and other places.
- Cover the turn. Front, both 3/4s, both sides, back. Video generation will ask for angles you never rendered; give it the neighbours.
- Add close-ups. Two or three face details and a fabric or prop detail. They are there to stop the face from melting when the subject is small in frame.
- Then vary expression and pose, not identity. Once the turn is covered, add a handful of expressions and a couple of full-body poses. Still the same outfit.
- 15 to 30 images total. Fewer tends not to generalise; more and you start averaging in mistakes.
- Caption for what changes. Captions describe the angle, the expression and the framing. The constant things (the character's name token, the outfit) stay constant in every caption.
Sheet first, then stills, then video
The order matters. A turnaround sheet like the one above is generated first, from a single reference or a written brief. The individual columns then seed the training set. Once the LoRA holds identity on stills, it goes into the video workflow, where its job is to keep the face from drifting through motion. A LoRA only works on the model family it was trained for, so it is trained on the model that workflow runs, MiniMax H3 or LTX 2.5. That is the whole "Brand Character System": sheet, LoRA, workflow, handover.
Need the sheet or the LoRA made rather than explained? Both are fixed-price commissions: a model sheet for €50, a trained character LoRA from €350.
Do I need a different dataset for image and video models?
For a character, no. The same sheet is the base for both. When the LoRA is headed for a video model I add a few extra 3/4 and side frames, because motion spends most of its time between the canonical angles.
What changes is that a video trainer also accepts clips. The one video LoRA I can give exact numbers for is a style LoRA, not a character: the 2.5D style LoRA for MiniMax H3 that is trained on screen in Part 3 of ComfyUI From Zero. Its dataset was 50 clips of about 3 seconds, not stills.
| Setting | Value |
|---|---|
| Dataset | 50 clips of about 3 seconds: 35 landscape, 15 portrait |
| Captions | One .txt file per clip, each starting with the trigger word 2.5D |
| Frames trained per clip | 73 at 24 fps, about 3 seconds (auto_frame_count: true) |
| Resolution settings | 256 and 512 px |
| Buckets in the training log | 320x192 and 672x384 px for the landscape clips, 192x320 and 384x672 px for the portrait clips |
| Audio | Trained together with the video (do_audio: true) |
| Trainer | AI Toolkit, MiniMax H3 reference-to-video architecture |
Dataset figures read from the job config and training logs of one AI Toolkit job that ended on 2026-09-29, on an RTX 5090 (32 GB), 96 GB RAM, Windows 11. One LoRA, one dataset. No stills-versus-clips comparison was run.
The clips give the trainer motion and an audio track to learn from, which a sheet of stills does not have. They did not need to be large: the 512 px setting put the landscape clips in a 672x384 px bucket.
Check the frame count in the log, not in the form. The job config also carried num_frames: 39, and that is not what trained. With auto frame count on, AI Toolkit took the count from each clip's length at 24 fps, and the bucket lines in the log show 73 frames for all 50 clips.
That is one style LoRA on one machine. I have no matching test for a character LoRA on H3 or LTX 2.5, so read the clip dataset as what I used for this style, not as a rule for every video model.
How many images for a face, a style, an object or a concept?
As rules of thumb: 15 to 30 for a face, 20 to 50 for a style, 10 to 20 for an object and 30 to 60 for a concept. Part 3 of the course gives these set sizes by goal. The 15 to 30 above is its row for a face, and the 50-clip set sits at the top of the style row, counted in clips instead of images.
| Goal | Set size | What stays constant |
|---|---|---|
| A face | 15 to 30 | The person |
| A style | 20 to 50 | The technique |
| An object | 10 to 20 | The object |
| A concept | 30 to 60 | The idea |
Everything else in the set should vary, because whatever the items share is what the LoRA learns. The caption rule from the same module is to caption what varies, not the concept: describe the things that change from one item to the next and leave the thing you are teaching to the trigger word. It is the same idea as rule 6 above, applied to clips.
The 2.5D look itself, and why MiniMax H3 does not reach it from a prompt alone, is covered in the MiniMax H3 2.5D style LoRA post. If you train on H3 yourself, there is a separate post on distillation handling in AI Toolkit for MiniMax H3 LoRAs.
Sources and files
- AI Toolkit by Ostris: the trainer used for the MiniMax H3 style LoRA described above.
- MiniMax H3 model card and LTX 2.5 model card: the two video models named in this post.
- Stubelius Ultimate H3 and Stubelius Ultimate LTX 2.5: my free ComfyUI workflows for those two models.
- ComfyUI From Zero: Part 3 covers datasets, captions and the training run the numbers above come from.
- Getting a 2.5D look out of MiniMax H3 with a style LoRA.
- AI Toolkit distillation handling for MiniMax H3 LoRAs.
- Commissions: turnaround sheets and trained LoRAs at a fixed price.
Common questions
Can I train from photos of a real person?
Technically yes, and the same rules apply. For client work I only train on subjects the client has rights to, and brand-safe projects only.
Should I lock the outfit or vary it?
Lock it when the outfit is part of the character, which is the case for a brand character and is what this sheet does. If you want a face that can wear anything, do the opposite: Part 3 of the course lists clothing and lighting among the things a face set should vary, with only the person held constant, because a LoRA learns whatever its images share.
Are more images better for a character LoRA?
Not by default. I stay at 15 to 30 images. Part 3 of the course makes the same point for every LoRA type: a small set of good images is worth more than a large mediocre one.
STUUBZZZ


