ComfyUI From Zero: The ComfyUI course that explains the machine
ComfyUI From Zero is a three-part video course by Stuubzzz Creative Studios that takes you from installing ComfyUI to training your own LoRA and writing a custom node. Read it. Change it. Build it. Part 1 is the whole beginner course, and it is free.
- Level
- Beginner to advanced
- Format
- On-demand video: slides plus live ComfyUI demos
- Parts
- 3. Part 1 free, Parts 2 and 3 sold as one bundle
- Modules
- 28 (10 + 9 + 9)
- Length
- Just over 10 hours: about 3 h 05 min, about 3 h 25 min, about 3 h 40 min
- Hardware
- 6 to 8 GB of VRAM for Part 1. A rented GPU works too
- Language
- English
- Price
- Part 1 free. Bundle: €29, one purchase
A peek inside · real slides from all three parts
278 slides in total · click one to enlarge
Six jobs, one photographer.Part 1
Four machines. Four different questions.Part 2
Four things, and you keep all of them.Part 3
The model is the photographer. ComfyUI is the studio.Part 1
A modern 12-billion image model, 1024×1024.Part 2
Everything your images have in common becomes the LoRA.Part 3
VRAM is the desk. RAM is the filing cabinet in the hallway.Part 1
The same seed, six strengths.Part 2
Load, change, send.Part 3
ComfyUI updates like a toddler.Part 1
Pick the lane, then stop worrying about it.Part 2
Same maths, different spelling.Part 3
How many times you wipe the foggy window.Part 1
Your VAE squashes time as well as space.Part 2
An average. That's genuinely all it is.Part 3
Read a model filename like a food label.Part 1
The console already tells you what it did.Part 2
Three prompts, chosen to answer three different questions.Part 3
The error shows where it failed. Not why.Part 1
What you pin, and what you leave free.Part 2
From “what is this” to “I built it”.Part 3
Ten things you can do right now.Part 1
Camera motion, written as a sentence.Part 2
A real node: brightness, with a slider.Part 3
Most tutorials hand you a workflow. This one hands you the reasons.
You download a 40-node workflow, drop the models in the right folders and press Run. It works, right up to the moment you need to change one thing. Then you are stuck, because nobody explained what the nodes are doing. This course is for that moment.
Why, not which button
Almost every module ends with a “You can now…” list: things you can do, not files you can download.
Your card, up front
Each module states the VRAM it assumes. Part 2 has fit tables for 8, 12, 16 and 24 GB cards.
One change at a time
Fix the seed, change exactly one thing, compare at 100%. Every claim in the course was tested that way.
Not dry
One running metaphor (the model is a photographer, ComfyUI is the studio) and slides with lines like “ComfyUI updates like a toddler.”
“Scary on day one. Free on day ten.”Part 1 · Module 1, on node graphs
Read it. Change it. Build it.
Part 1 is free and stands on its own. Parts 2 and 3 are one paid bundle and continue exactly where Part 1 stops.
Understand the machine
Install ComfyUI, build the seven-node text-to-image workflow from an empty canvas, and fix the eight errors beginners actually hit.
- What every node and every KSampler setting does
- Which model file fits your card, and where it goes
- The eight errors beginners hit, and the fix for each
Open the hood
Change a downloaded workflow and know what will happen before you press Run.
- Fit tables for 8, 12, 16 and 24 GB cards
- LoRAs, the four ways to edit a picture, and video explained from the VAE up
- Structured prompting, chaining clips, renting a GPU
Build the engine
Train your own LoRA, merge models, write a custom node and run your workflow from a script.
- A real LoRA trained on screen, from the first samples to the winning checkpoint
- LoRA formats, conversion and model merges
- A custom node in about twenty lines, and the API in twelve
28 modules, in the order you need them
Open any module to see what it covers and what you can do once you have watched it. The module titles are the real ones and the lists are taken from the slides.
Part 1: Understand the machine
1.01What ComfyUI actually is
The one metaphor the whole course runs on: the model is the photographer, ComfyUI is the studio. Why a node graph is scary on day one and free on day ten.
- Say what a model is and what the studio is
- Explain why ComfyUI shows you every cable
- Recognise a workflow and which way it flows
- Name the four ideas that carry the course
1.02Can my computer run this?
VRAM versus RAM, finding your real number in 30 seconds, five hardware tiers and one rule of thumb: model file size plus about 2 GB.
- Tell VRAM from RAM and why overflow is slow, not broken
- Find your real VRAM number in 30 seconds
- Name your tier and your install path
- Estimate if a model fits: file size + 2 GB
1.03Install and first launch
The two official installs, a live walkthrough, and the three failures that stop most people: the nervous bouncer, the overprotective antivirus and the path from hell.
- Install ComfyUI the official way
- Fix the three most common install failures
- Tell the engine room from the dashboard
- Find the canvas, Run, sidebar and bottom panel
1.04Your first image
A picture before any theory, then an error on purpose, so a red node becomes a warning light instead of a crash.
- Load a workflow by dragging a .json onto the canvas
- Read a red node and fix a missing model
- Make an image and find it in the output folder
- Know the five folders that matter
1.05The words
41 terms in five groups. The 28 core ones each get what it is, the photographer version and what to remember. Ends on the two-page cheat sheet.
- Files, nodes and settings in plain English
- The photographer version of the core terms
- A two-page cheat sheet to keep
1.06Meet the photographer
The seven-node workflow built from an empty canvas, then what steps, seed, CFG, sampler and scheduler actually do, tested one change at a time.
- Build text-to-image from an empty canvas
- Explain positive vs negative and the cable rules
- Say what steps, seed, CFG, sampler and scheduler do
- Test anything with the compare ritual
1.07Reading model files
FP16, FP8 and GGUF, checkpoints versus models in pieces, and how to read a filename like a food label so the file lands in the right folder.
- Choose FP16, FP8 or GGUF for your card
- Put checkpoints and pieces in the right folders
- Read a filename and pick the right loader
- Find a new model's nodes in its official template
1.08Saving, sharing, not breaking things
Save versus export, the workflow hidden in every PNG, custom nodes as code on your machine, and update rules for software that updates like a toddler.
- Choose Save, Save As or Export on purpose
- Get a workflow back from any ComfyUI PNG
- Install, fix and remove custom nodes safely
- Update without breaking things, and use bypass for tests
1.09The eight errors
The eight errors beginners actually see, each with what you see, what it means and the fix. Plus the golden rule: walk upstream.
- Read any error: kind, where, clue
- Fix the eight errors beginners actually see
- Walk upstream to the real cause
- Ask for help with everything a helper needs
1.10What you can do now
Ten concrete skills, three homework challenges and one honest pitch for Parts 2 and 3.
- Rebuild text-to-image from memory
- Swap to a GGUF model and compare on one seed
- Fix a workflow that is broken in three places
Part 2: Open the hood
2.01What changes now
Distilled models break Part 1's rules. The three questions every workflow answers, how to read a 40-node graph, and the one-change, fixed-seed method behind every claim in the course.
- Spot when a model breaks Part 1's rules
- Ask the three questions of any workflow
- Find the loaders, sampler and output in forty nodes
- Test a change one at a time, with a fixed seed
2.02What your card can run
The four things that fill your VRAM, the quantisation ladder, and fit tables for images and video on 8, 12, 16 and 24 GB cards.
- Name the four things holding your VRAM
- Choose a quantisation for your card on purpose
- Predict whether a model fits before downloading it
- Work the out-of-memory checklist in the right order
2.03Where the memory goes
Dynamic VRAM, launch flags, attention backends, and the console lines that tell you what ComfyUI actually did.
- Explain dynamic VRAM and when to turn it off
- Choose launch flags one at a time, with evidence
- Say what attention is and why resolution costs so much
- Read the console lines that reveal offloading
2.04LoRAs, properly
What a LoRA does to a model, matching it to its base, finding a strength in one sweep, stacking, and the steps and CFG a speed LoRA demands.
- Explain what a LoRA does to the model
- Match a LoRA to its base before downloading
- Find a working strength in one sweep
- Set steps and CFG correctly for a speed LoRA
2.05Four ways to change a picture
img2img, inpaint, edit models and reference as four machines that answer four different questions. Crop-and-stitch, outpainting and composite-then-harmonise.
- Pick between four mechanisms on purpose
- Kill an inpaint seam with feather and context
- Write an instruction an edit model can follow
- Composite a real object and harmonise it
2.06Video, under the hood
Why frame counts are 8n+1, why sizes are multiples of 32, why the frame rate belongs to the model, and how high-noise and low-noise samplers split the work.
- Explain 8n+1 and multiples of 32 from the VAE
- Say why fps isn't yours to choose
- Use a high / low noise pair correctly
- Plan a shoot: draft small, finish big
2.07Reference and structured prompting
The three-field prompt that audio-video models were trained on: shots, camera moves, dialogue tags and word budgets, so you can write it without someone else's chatbot.
- Fill in the three fields without a chatbot
- Write shots, cuts and camera moves the model follows
- Get dialogue timed and tagged correctly
- Align a reference so it actually applies
2.08Chaining clips
Longer pieces from short clips: a clean handover frame, stopping colour drift, audio across a cut, and planning a minute as a shot list.
- Chain clips on a clean last frame
- Stop colour drift before it stacks up
- Keep audio continuous across cuts
- Plan a minute as a shot list, not a lottery
2.09When your card says no
Renting a GPU properly: network volumes, a six-step setup, ballpark cost per finished thing, and what the licences let you sell.
- Pick between renting, hosted and API deliberately
- Keep models on a network volume and download once
- Work out the cost per finished thing
- Check the licence before you sell the work
Part 3: Build the engine
3.01What you're about to own
Training demystified in five steps, and the four things you build: a LoRA, a merged model, a custom node and a script that runs your workflow.
- Explain training in five sentences
- Say what a LoRA file actually contains
- Know what you'll have built by the end
3.02Datasets that work
Everything your images have in common becomes the LoRA. Set sizes, buckets and captions, then the real 50-clip dataset for the course LoRA: the first images generated on screen, then the finished clips and captions walked through.
- Choose images by what they don't share
- Pick a set size for the job
- Crop and size so the trainer doesn't decide for you
- Caption the variation, not the concept
3.03Training, for real
Twenty settings, six decisions. AI Toolkit configured field by field and run live, with what to check when a run stalls without an error.
- Explain rank, alpha, learning rate and steps
- Fill in a config without copying someone's screenshot
- Recognise overtraining before it finishes
- Train on a small card, deliberately
3.04Reading a training run
Ignore the loss, read the samples. Three diagnostic prompts, the three states of a LoRA, and a routine for picking the winning checkpoint.
- Ignore loss and read the samples instead
- Write sample prompts that diagnose
- Spot overbaked before the run ends
- Pick a checkpoint by comparison, not by hope
3.05LoRA formats and conversion
Why a LoRA can load perfectly and do nothing: what is inside a safetensors file, why key names are the contract, and when conversion is possible.
- Read a safetensors header without loading it
- Explain why a LoRA silently does nothing
- Tell convertible from impossible
- Verify a fix with an image, not an absent error
3.06Model merges
A merge is a weighted average, and that is all it is. Baking a LoRA in, block weights, and an honest test routine.
- Explain a merge as weighted averaging
- Bake a LoRA in, knowing what you gave up
- Use block weights deliberately
- Judge a merge against a fixed test set
3.07Writing a custom node
A working ComfyUI node from four declarations in about twenty lines of Python, the tensor shapes that trip everyone up, and how to use AI on code you can read.
- Write a working node from four declarations
- Handle IMAGE and MASK tensors correctly
- Read an import error and fix it
- Use AI to extend code you can read
3.08ComfyUI without the UI
The HTTP API: two file formats, three endpoints and a twelve-line script that queues your workflow.
- Export a workflow in API format
- Queue jobs from a script and collect the results
- Automate the testing you were avoiding
- Keep the port where it belongs
3.09Where this leaves you
The whole course on one page, three things to build next, and what stays true when the models change.
- Train your own LoRA
- Merge and bake models
- Write and read custom nodes
- Drive ComfyUI from code
Modules 1 to 4 are the long hands-on ones (about 2 h 50 min of the part). Modules 5 to 9 are shorter, about 50 minutes in total.
Go deeper on the blog: the MiniMax H3 three-field prompt (module 2.07) · a consistent character LoRA dataset (3.02) · AI Toolkit's distillation setting for MiniMax H3 (3.03) · why a FastH3 LoRA loads and does nothing (3.05).
Slides explain it. Then ComfyUI proves it.
Close to 150 short ComfyUI screen demos are cut into the three parts, each one showing the thing the slide just claimed. Three of them, as they appear in the course:
Every image starts as TV static
The same render stopped after 1, 3, 6, 10, 16 and 25 steps, side by side. Early steps decide the big shapes, late steps add detail.
10 steps or 30? Look at 100%
A distilled model at 10 steps and at 30 steps on the same seed. The 30-step picture comes out different, not better, after three times the wait.
A node you wrote yourself
The Brighten Image node from Module 7: four declarations, about twenty lines of Python, one slider.
One dial. Watch what it does.
These are the images from two Part 2 slides. Drag the sliders: this is how the course makes a setting click.
Denoise decides how far back in time you go.
A fresh coat of paint. The picture is noised only a little, so everything big survives.
Image-to-image noises your picture to a point on the schedule, then denoises from there. Low denoise keeps the large shapes, because those were decided in the early steps you never undid.
The same seed, six strengths.
No LoRA at all. This is your control. Always keep one.
Sweep it once for each LoRA you care about, with a fixed seed. Ten minutes that saves you a hundred bad renders.
How much VRAM do you need for ComfyUI?
For a 12-billion-parameter image model at 1024×1024, 8 GB is enough: the 6.5 GB Q4 GGUF fits in VRAM. A 12 GB card fits the 9 GB Q6, 16 GB the 12 GB fp8 or int8, and 24 GB the 24 GB bf16. Whether a model fits is arithmetic: weights, plus text encoder, plus working space, against your VRAM.
Part 2 turns that into tables. Below is its table for a modern 12-billion-parameter image model.
8 GB
Q4 and Q6 image models. Short 480p video. Offload on, patience on.
12 GB
Q6 image models in VRAM, fp8 and int8 with offloading. Real 480p video, 720p if you wait.
16 GB
Almost everything, most of the time. The best value tier right now.
24 GB+
Stop thinking about memory. Start thinking about time.
| Card | bf16 · 24 GB file | fp8 / int8 · 12 GB | GGUF Q6 · 9 GB | GGUF Q4 · 6.5 GB |
|---|---|---|---|---|
| 8 GB | no | no | offload | fits |
| 12 GB | no | offload | fits | fits |
| 16 GB | offload | fits | fits | fits |
| 24 GB | fits | fits | fits | fits |
Why a 12 GB model does not fit a 12 GB card: the text encoder is another 4 to 5 GB and the working space 2 to 4 GB more. Part 2 has the same table for video, plus the out-of-memory checklist in the order that costs you least.
Four things, and you keep all of them.
Part 2 is about using other people's models. Part 3 is about making your own. The LoRA trained on screen is a real one: a 2.5D style LoRA for MiniMax H3, built from a 50-clip dataset and followed from the first samples to the checkpoint that won.
A LoRA of your own
Trained on your own images, on your own machine, that nobody else has.
A merged model
Two models combined, or a LoRA baked in, tested properly rather than hopefully.
A custom node
Written by hand first, then extended with AI once you can read it.
A script that runs it
Your workflow, called from code, a hundred times, while you sleep.
Hardware, honestly: training an image LoRA is comfortable from 12 GB of VRAM at 512 px, given enough system RAM to offload into. Video LoRAs are heavier: 24 GB is the floor, and the one on screen was trained on a 32 GB RTX 5090 with 96 GB of RAM, still offloading. Below 24 GB, rent a GPU; Part 2 explains how. The custom node, the merges and the API script run on small cards.
Hi, I'm Stuubzzz.
About a year ago I started ComfyUI on a PC with 8 GB of VRAM and 16 GB of RAM. I was stuck with the smallest GGUF files and could not find one place that said what would actually run. I learned it because I wanted to make animation for myself. The teaching came later.
Today I build the Stubelius node packs for ComfyUI, train and release LoRAs, and run almost everything on one RTX 5090. This is the beginner course I wish had existed when I started, followed by two parts that go where I wanted to go next.
Start free. Pay only if you want to go further.
Part 1 stays free. Understanding should be.
Understand the machine
- 10 modules, about 3 h 05 min of video
- The two-page cheat sheet
- No account, no email, no catch
Play Part 1 here
Plays on this page, or opens on YouTube. The cheat sheet is further down.
Open the hood, then build the engine
- 18 modules, just over 7 hours of video in two parts
- VRAM fit tables, LoRAs, the four ways to edit a picture, video from the VAE up, structured prompting, chaining clips, renting a GPU
- Training your own LoRA, LoRA formats, model merges, a custom node, the API
- One purchase on this site. No subscription needed
Payment by Stripe. By buying you accept the course terms.
Already a patron? The bundle is also included in the Studio tier on Patreon.
Take the cheat sheet.
Two pages: the files you download, the nodes in a basic workflow, what each KSampler setting does, the folders you touch and the first-week errors. Plain meaning, the photographer version, and what to remember.
ComfyUI cheat sheet
2 pages · print it ZIPPart 1 course files
18 MB · slides, workflows, help templateFree, no strings.

Straight answers
What is ComfyUI?
ComfyUI is a free, open-source program that runs image and video models on your own computer, or on a rented GPU. Instead of one Generate button, you build the process as a graph of nodes: boxes that each do one job, joined by cables, running left to right. Read the full definition in the glossary.
Is Part 1 of the ComfyUI course really free?
Yes. Part 1 is the complete beginner course, about 3 h 05 min in 10 modules, and it stays free. The cheat sheet is free as well. Parts 2 and 3 are sold together as one paid bundle.
What do I need to follow along?
For Part 1, a Windows PC with an Nvidia graphics card is the main path shown on screen. Apple Silicon Macs, AMD cards and rented cloud GPUs each get a slide with the route to take, not a full walkthrough. You need about 30 GB of free disk space to start. Several modules are watch-only and need no GPU at all.
Do I need an expensive graphics card?
No. Part 1 works on 6 to 8 GB cards with FP8 or GGUF files. Part 2 gives fit tables for 8, 12, 16 and 24 GB cards and says which tier each module assumes. In Part 3, training an image LoRA is comfortable from 12 GB with enough system RAM; video LoRAs want 24 GB or more (the one in the course was trained on a 32 GB card, with offloading), or a rented GPU, which Part 2 explains how to set up.
Which models does the course use?
The lessons are about how ComfyUI works, so they carry over between models. On screen you will see an SDXL-class checkpoint in Part 1; Krea 2 Turbo, Qwen-Image 2.1, MiniMax H3 and the Wan 2.2 template in Part 2; and AI Toolkit training a MiniMax H3 LoRA in Part 3.
Will it be out of date when the next model comes out?
Models change every month. The machine underneath does not. The last module lists what stays true: model, interface and workflow are separate things; whether a model fits is arithmetic; change one thing and fix the seed; a model learns what your images have in common; and the console tells you what happened.
Do I need to know Python for Part 3?
The course does not assume Python experience. You read and type a little: about twenty lines for the custom node, twelve for the API script and six for the LoRA inspector. Each one is walked through on a slide first.
Is the course narrated by AI?
No. I recorded every module in my own voice, ad-libs and dad jokes included.
How long is the whole course?
Just over ten hours in 28 modules: about 3 h 05 min for Part 1, about 3 h 25 min for Part 2 and about 3 h 40 min for Part 3.
How much does the ComfyUI course cost?
Part 1 is free. Parts 2 and 3 are sold together as one bundle, as a single purchase on this page. The bundle costs €29, as one purchase.
Can I follow the course without a GPU?
Yes, with a rented cloud GPU. Part 1 names that route in three steps and leaves the setup to Part 2, whose last module covers renting properly: a network volume so models download once, a six-step setup, and ballpark costs per finished image, clip and training run.
Is this course for complete beginners?
Part 1 assumes you have never opened ComfyUI. It starts with what ComfyUI is and whether your computer can run it, and you make your first image in Module 4. Parts 2 and 3 assume you have done Part 1 or already know that material.
What if I get stuck?
Part 1 has a whole module on the eight errors beginners hit and how to ask for help so you get an answer. If something on your own machine is fighting you, you can book a 60-minute 1-on-1 session and fix it live on a screen share.
Understand the machine before you copy a single workflow.
Part 1 costs nothing. Watch the first module and decide for yourself.
STUUBZZZ