ComfyUI course · 3 parts · just over 10 hours

ComfyUI From Zero: The ComfyUI course that explains the machine

ComfyUI From Zero is a three-part video course by Stuubzzz Creative Studios that takes you from installing ComfyUI to training your own LoRA and writing a custom node. Read it. Change it. Build it. Part 1 is the whole beginner course, and it is free.

Start Part 1, free Buy the full course The full course is €29, paid once: Parts 1 to 3 and all the course files.
Part 1 free 28 modules just over 10 hours slides + live ComfyUI demos
workflow-01-first-image.json · seven nodesPart 1
Load Checkpoint MODELCLIPVAE CLIP Text Encodea vintage cameraon a wooden desk CLIP Text Encode(negative) Empty Latent Image1024 × 1024 KSampler modelpositivenegativelatent seed 42 · steps 8cfg 1.0 · euler VAE Decodesamplesvae Save Image ComfyUI_00001_.png
The seven-node workflow you build from an empty canvas in Part 1. By the end you can read it at a glance.
Level
Beginner to advanced
Format
On-demand video: slides plus live ComfyUI demos
Parts
3. Part 1 free, Parts 2 and 3 sold as one bundle
Modules
28 (10 + 9 + 9)
Length
Just over 10 hours: about 3 h 05 min, about 3 h 25 min, about 3 h 40 min
Hardware
6 to 8 GB of VRAM for Part 1. A rented GPU works too
Language
English
Price
Part 1 free. Bundle: €29, one purchase

A peek inside · real slides from all three parts

278 slides in total · click one to enlarge
Why this course exists

Most tutorials hand you a workflow. This one hands you the reasons.

You download a 40-node workflow, drop the models in the right folders and press Run. It works, right up to the moment you need to change one thing. Then you are stuck, because nobody explained what the nodes are doing. This course is for that moment.

Why, not which button

Almost every module ends with a “You can now…” list: things you can do, not files you can download.

Your card, up front

Each module states the VRAM it assumes. Part 2 has fit tables for 8, 12, 16 and 24 GB cards.

One change at a time

Fix the seed, change exactly one thing, compare at 100%. Every claim in the course was tested that way.

Not dry

One running metaphor (the model is a photographer, ComfyUI is the studio) and slides with lines like “ComfyUI updates like a toddler.”

“Scary on day one. Free on day ten.”Part 1 · Module 1, on node graphs

Curriculum

28 modules, in the order you need them

Open any module to see what it covers and what you can do once you have watched it. The module titles are the real ones and the lists are taken from the slides.

Part 1 · Beginner

Part 1: Understand the machine

Free10 modules · about 3 h 05 min
1.01

What ComfyUI actually is

The one metaphor the whole course runs on: the model is the photographer, ComfyUI is the studio. Why a node graph is scary on day one and free on day ten.

You can now
  • Say what a model is and what the studio is
  • Explain why ComfyUI shows you every cable
  • Recognise a workflow and which way it flows
  • Name the four ideas that carry the course
1.02

Can my computer run this?

VRAM versus RAM, finding your real number in 30 seconds, five hardware tiers and one rule of thumb: model file size plus about 2 GB.

You can now
  • Tell VRAM from RAM and why overflow is slow, not broken
  • Find your real VRAM number in 30 seconds
  • Name your tier and your install path
  • Estimate if a model fits: file size + 2 GB
1.03

Install and first launch

The two official installs, a live walkthrough, and the three failures that stop most people: the nervous bouncer, the overprotective antivirus and the path from hell.

You can now
  • Install ComfyUI the official way
  • Fix the three most common install failures
  • Tell the engine room from the dashboard
  • Find the canvas, Run, sidebar and bottom panel
1.04

Your first image

A picture before any theory, then an error on purpose, so a red node becomes a warning light instead of a crash.

You can now
  • Load a workflow by dragging a .json onto the canvas
  • Read a red node and fix a missing model
  • Make an image and find it in the output folder
  • Know the five folders that matter
1.05

The words

41 terms in five groups. The 28 core ones each get what it is, the photographer version and what to remember. Ends on the two-page cheat sheet.

You can now
  • Files, nodes and settings in plain English
  • The photographer version of the core terms
  • A two-page cheat sheet to keep
1.06

Meet the photographer

The seven-node workflow built from an empty canvas, then what steps, seed, CFG, sampler and scheduler actually do, tested one change at a time.

You can now
  • Build text-to-image from an empty canvas
  • Explain positive vs negative and the cable rules
  • Say what steps, seed, CFG, sampler and scheduler do
  • Test anything with the compare ritual
1.07

Reading model files

FP16, FP8 and GGUF, checkpoints versus models in pieces, and how to read a filename like a food label so the file lands in the right folder.

You can now
  • Choose FP16, FP8 or GGUF for your card
  • Put checkpoints and pieces in the right folders
  • Read a filename and pick the right loader
  • Find a new model's nodes in its official template
1.08

Saving, sharing, not breaking things

Save versus export, the workflow hidden in every PNG, custom nodes as code on your machine, and update rules for software that updates like a toddler.

You can now
  • Choose Save, Save As or Export on purpose
  • Get a workflow back from any ComfyUI PNG
  • Install, fix and remove custom nodes safely
  • Update without breaking things, and use bypass for tests
1.09

The eight errors

The eight errors beginners actually see, each with what you see, what it means and the fix. Plus the golden rule: walk upstream.

You can now
  • Read any error: kind, where, clue
  • Fix the eight errors beginners actually see
  • Walk upstream to the real cause
  • Ask for help with everything a helper needs
1.10

What you can do now

Ten concrete skills, three homework challenges and one honest pitch for Parts 2 and 3.

You can now
  • Rebuild text-to-image from memory
  • Swap to a GGUF model and compare on one seed
  • Fix a workflow that is broken in three places
Part 2 · Intermediate

Part 2: Open the hood

Paid bundle9 modules · about 3 h 25 min
2.01

What changes now

Distilled models break Part 1's rules. The three questions every workflow answers, how to read a 40-node graph, and the one-change, fixed-seed method behind every claim in the course.

You can now
  • Spot when a model breaks Part 1's rules
  • Ask the three questions of any workflow
  • Find the loaders, sampler and output in forty nodes
  • Test a change one at a time, with a fixed seed
2.02

What your card can run

The four things that fill your VRAM, the quantisation ladder, and fit tables for images and video on 8, 12, 16 and 24 GB cards.

You can now
  • Name the four things holding your VRAM
  • Choose a quantisation for your card on purpose
  • Predict whether a model fits before downloading it
  • Work the out-of-memory checklist in the right order
2.03

Where the memory goes

Dynamic VRAM, launch flags, attention backends, and the console lines that tell you what ComfyUI actually did.

You can now
  • Explain dynamic VRAM and when to turn it off
  • Choose launch flags one at a time, with evidence
  • Say what attention is and why resolution costs so much
  • Read the console lines that reveal offloading
2.04

LoRAs, properly

What a LoRA does to a model, matching it to its base, finding a strength in one sweep, stacking, and the steps and CFG a speed LoRA demands.

You can now
  • Explain what a LoRA does to the model
  • Match a LoRA to its base before downloading
  • Find a working strength in one sweep
  • Set steps and CFG correctly for a speed LoRA
2.05

Four ways to change a picture

img2img, inpaint, edit models and reference as four machines that answer four different questions. Crop-and-stitch, outpainting and composite-then-harmonise.

You can now
  • Pick between four mechanisms on purpose
  • Kill an inpaint seam with feather and context
  • Write an instruction an edit model can follow
  • Composite a real object and harmonise it
2.06

Video, under the hood

Why frame counts are 8n+1, why sizes are multiples of 32, why the frame rate belongs to the model, and how high-noise and low-noise samplers split the work.

You can now
  • Explain 8n+1 and multiples of 32 from the VAE
  • Say why fps isn't yours to choose
  • Use a high / low noise pair correctly
  • Plan a shoot: draft small, finish big
2.07

Reference and structured prompting

The three-field prompt that audio-video models were trained on: shots, camera moves, dialogue tags and word budgets, so you can write it without someone else's chatbot.

You can now
  • Fill in the three fields without a chatbot
  • Write shots, cuts and camera moves the model follows
  • Get dialogue timed and tagged correctly
  • Align a reference so it actually applies
2.08

Chaining clips

Longer pieces from short clips: a clean handover frame, stopping colour drift, audio across a cut, and planning a minute as a shot list.

You can now
  • Chain clips on a clean last frame
  • Stop colour drift before it stacks up
  • Keep audio continuous across cuts
  • Plan a minute as a shot list, not a lottery
2.09

When your card says no

Renting a GPU properly: network volumes, a six-step setup, ballpark cost per finished thing, and what the licences let you sell.

You can now
  • Pick between renting, hosted and API deliberately
  • Keep models on a network volume and download once
  • Work out the cost per finished thing
  • Check the licence before you sell the work
Part 3 · Advanced

Part 3: Build the engine

Paid bundle9 modules · about 3 h 40 min
3.01

What you're about to own

Training demystified in five steps, and the four things you build: a LoRA, a merged model, a custom node and a script that runs your workflow.

You can now
  • Explain training in five sentences
  • Say what a LoRA file actually contains
  • Know what you'll have built by the end
3.02

Datasets that work

Everything your images have in common becomes the LoRA. Set sizes, buckets and captions, then the real 50-clip dataset for the course LoRA: the first images generated on screen, then the finished clips and captions walked through.

You can now
  • Choose images by what they don't share
  • Pick a set size for the job
  • Crop and size so the trainer doesn't decide for you
  • Caption the variation, not the concept
3.03

Training, for real

Twenty settings, six decisions. AI Toolkit configured field by field and run live, with what to check when a run stalls without an error.

You can now
  • Explain rank, alpha, learning rate and steps
  • Fill in a config without copying someone's screenshot
  • Recognise overtraining before it finishes
  • Train on a small card, deliberately
3.04

Reading a training run

Ignore the loss, read the samples. Three diagnostic prompts, the three states of a LoRA, and a routine for picking the winning checkpoint.

You can now
  • Ignore loss and read the samples instead
  • Write sample prompts that diagnose
  • Spot overbaked before the run ends
  • Pick a checkpoint by comparison, not by hope
3.05

LoRA formats and conversion

Why a LoRA can load perfectly and do nothing: what is inside a safetensors file, why key names are the contract, and when conversion is possible.

You can now
  • Read a safetensors header without loading it
  • Explain why a LoRA silently does nothing
  • Tell convertible from impossible
  • Verify a fix with an image, not an absent error
3.06

Model merges

A merge is a weighted average, and that is all it is. Baking a LoRA in, block weights, and an honest test routine.

You can now
  • Explain a merge as weighted averaging
  • Bake a LoRA in, knowing what you gave up
  • Use block weights deliberately
  • Judge a merge against a fixed test set
3.07

Writing a custom node

A working ComfyUI node from four declarations in about twenty lines of Python, the tensor shapes that trip everyone up, and how to use AI on code you can read.

You can now
  • Write a working node from four declarations
  • Handle IMAGE and MASK tensors correctly
  • Read an import error and fix it
  • Use AI to extend code you can read
3.08

ComfyUI without the UI

The HTTP API: two file formats, three endpoints and a twelve-line script that queues your workflow.

You can now
  • Export a workflow in API format
  • Queue jobs from a script and collect the results
  • Automate the testing you were avoiding
  • Keep the port where it belongs
3.09

Where this leaves you

The whole course on one page, three things to build next, and what stays true when the models change.

You can now
  • Train your own LoRA
  • Merge and bake models
  • Write and read custom nodes
  • Drive ComfyUI from code

Modules 1 to 4 are the long hands-on ones (about 2 h 50 min of the part). Modules 5 to 9 are shorter, about 50 minutes in total.

Go deeper on the blog: the MiniMax H3 three-field prompt (module 2.07) · a consistent character LoRA dataset (3.02) · AI Toolkit's distillation setting for MiniMax H3 (3.03) · why a FastH3 LoRA loads and does nothing (3.05).

How it is taught

Slides explain it. Then ComfyUI proves it.

Close to 150 short ComfyUI screen demos are cut into the three parts, each one showing the thing the slide just claimed. Three of them, as they appear in the course:

Part 1 · Meet the photographer

Every image starts as TV static

The same render stopped after 1, 3, 6, 10, 16 and 25 steps, side by side. Early steps decide the big shapes, late steps add detail.

Part 2 · What changes now

10 steps or 30? Look at 100%

A distilled model at 10 steps and at 30 steps on the same seed. The 30-step picture comes out different, not better, after three times the wait.

Part 3 · Writing a custom node

A node you wrote yourself

The Brighten Image node from Module 7: four declarations, about twenty lines of Python, one slider.

Try a lesson right here

One dial. Watch what it does.

These are the images from two Part 2 slides. Drag the sliders: this is how the course makes a setting click.

A red bicycle leaning on a pale blue wall, image-to-image at denoise 0.2: almost identical to the source photo The same bicycle scene at denoise 0.4: same composition, cleaner surfaces The same bicycle scene at denoise 0.6: recognisably the same idea, different details The bicycle scene at denoise 0.8: a different bicycle and angle, loosely based on the source The bicycle scene at denoise 1.0: a completely new picture
Part 2 · Four ways to change a picture

Denoise decides how far back in time you go.

denoise0.2
0.20.40.60.81.0

A fresh coat of paint. The picture is noised only a little, so everything big survives.

Image-to-image noises your picture to a point on the schedule, then denoises from there. Low denoise keeps the large shapes, because those were decided in the early steps you never undid.

Portrait of an older bearded man in front of a rainy window, generated with no LoRA applied (strength 0.0) The same prompt and seed with the LoRA at strength 0.3: a different pose and background from the no-LoRA picture The same portrait and seed with the LoRA at strength 0.6 The same portrait and seed with the LoRA at strength 0.8 The same portrait and seed with the LoRA at strength 1.0 The same portrait and seed with the LoRA at strength 1.2
Part 2 · LoRAs, properly

The same seed, six strengths.

LoRA strength0.0
0.00.30.60.81.01.2

No LoRA at all. This is your control. Always keep one.

Sweep it once for each LoRA you care about, with a fixed seed. Ten minutes that saves you a hundred bad renders.

Pick your VRAM

How much VRAM do you need for ComfyUI?

For a 12-billion-parameter image model at 1024×1024, 8 GB is enough: the 6.5 GB Q4 GGUF fits in VRAM. A 12 GB card fits the 9 GB Q6, 16 GB the 12 GB fp8 or int8, and 24 GB the 24 GB bf16. Whether a model fits is arithmetic: weights, plus text encoder, plus working space, against your VRAM.

Part 2 turns that into tables. Below is its table for a modern 12-billion-parameter image model.

12 GB

Q6 image models in VRAM, fp8 and int8 with offloading. Real 480p video, 720p if you wait.

Images any
Video usable
Training small LoRAs
Fit table from Part 2, Module 2. It states whether a model fits, not how many seconds it takes: speed depends on your exact card.
Cardbf16 · 24 GB filefp8 / int8 · 12 GBGGUF Q6 · 9 GBGGUF Q4 · 6.5 GB
8 GBnonooffloadfits
12 GBnooffloadfitsfits
16 GBoffloadfitsfitsfits
24 GBfitsfitsfitsfits
fits runs in VRAM offload runs while offloading, expect 3 to 10 times slower no not worth attempting

Why a 12 GB model does not fit a 12 GB card: the text encoder is another 4 to 5 GB and the working space 2 to 4 GB more. Part 2 has the same table for video, plus the out-of-memory checklist in the order that costs you least.

Part 3 · Build the engine

Four things, and you keep all of them.

Part 2 is about using other people's models. Part 3 is about making your own. The LoRA trained on screen is a real one: a 2.5D style LoRA for MiniMax H3, built from a 50-clip dataset and followed from the first samples to the checkpoint that won.

A LoRA of your own

Trained on your own images, on your own machine, that nobody else has.

A merged model

Two models combined, or a LoRA baked in, tested properly rather than hopefully.

A custom node

Written by hand first, then extended with AI once you can read it.

A script that runs it

Your workflow, called from code, a hundred times, while you sleep.

Output of the LoRA trained in Part 3: step 4,250, strength 1.0, native 720p.

Hardware, honestly: training an image LoRA is comfortable from 12 GB of VRAM at 512 px, given enough system RAM to offload into. Video LoRAs are heavier: 24 GB is the floor, and the one on screen was trained on a 32 GB RTX 5090 with 96 GB of RAM, still offloading. Below 24 GB, rent a GPU; Part 2 explains how. The custom node, the merges and the API script run on small cards.

Who is teaching

Hi, I'm Stuubzzz.

About a year ago I started ComfyUI on a PC with 8 GB of VRAM and 16 GB of RAM. I was stuck with the smallest GGUF files and could not find one place that said what would actually run. I learned it because I wanted to make animation for myself. The teaching came later.

Today I build the Stubelius node packs for ComfyUI, train and release LoRAs, and run almost everything on one RTX 5090. This is the beginner course I wish had existed when I started, followed by two parts that go where I wanted to go next.

28modules across three parts
278slides, every one drawn for this course
~10 hof video, Part 1 free
3free ComfyUI node packs on GitHub
Get the course

Start free. Pay only if you want to go further.

Part 1 stays free. Understanding should be.

Part 1Beginner

Understand the machine

Free
  • 10 modules, about 3 h 05 min of video
  • The two-page cheat sheet
  • No account, no email, no catch
Slide from Part 1: Six jobs, one photographer.Play Part 1 here

Plays on this page, or opens on YouTube. The cheat sheet is further down.

Parts 2 + 3Intermediate + advanced

Open the hood, then build the engine

€29one purchase
  • 18 modules, just over 7 hours of video in two parts
  • VRAM fit tables, LoRAs, the four ways to edit a picture, video from the VAE up, structured prompting, chaining clips, renting a GPU
  • Training your own LoRA, LoRA formats, model merges, a custom node, the API
  • One purchase on this site. No subscription needed

Payment by Stripe. By buying you accept the course terms.

Already a patron? The bundle is also included in the Studio tier on Patreon.

Free with Part 1

Take the cheat sheet.

Two pages: the files you download, the nodes in a basic workflow, what each KSampler setting does, the folders you touch and the first-week errors. Plain meaning, the photographer version, and what to remember.

Free, no strings.

comfyui-cheatsheet.pdf · page 1 of 2
Page 1 of the ComfyUI cheat sheet: files you download, formats, nodes in a basic workflow, the seven-node pipeline, the folders you touch and first-week errors
Questions

Straight answers

What is ComfyUI?

ComfyUI is a free, open-source program that runs image and video models on your own computer, or on a rented GPU. Instead of one Generate button, you build the process as a graph of nodes: boxes that each do one job, joined by cables, running left to right. Read the full definition in the glossary.

Is Part 1 of the ComfyUI course really free?

Yes. Part 1 is the complete beginner course, about 3 h 05 min in 10 modules, and it stays free. The cheat sheet is free as well. Parts 2 and 3 are sold together as one paid bundle.

What do I need to follow along?

For Part 1, a Windows PC with an Nvidia graphics card is the main path shown on screen. Apple Silicon Macs, AMD cards and rented cloud GPUs each get a slide with the route to take, not a full walkthrough. You need about 30 GB of free disk space to start. Several modules are watch-only and need no GPU at all.

Do I need an expensive graphics card?

No. Part 1 works on 6 to 8 GB cards with FP8 or GGUF files. Part 2 gives fit tables for 8, 12, 16 and 24 GB cards and says which tier each module assumes. In Part 3, training an image LoRA is comfortable from 12 GB with enough system RAM; video LoRAs want 24 GB or more (the one in the course was trained on a 32 GB card, with offloading), or a rented GPU, which Part 2 explains how to set up.

Which models does the course use?

The lessons are about how ComfyUI works, so they carry over between models. On screen you will see an SDXL-class checkpoint in Part 1; Krea 2 Turbo, Qwen-Image 2.1, MiniMax H3 and the Wan 2.2 template in Part 2; and AI Toolkit training a MiniMax H3 LoRA in Part 3.

Will it be out of date when the next model comes out?

Models change every month. The machine underneath does not. The last module lists what stays true: model, interface and workflow are separate things; whether a model fits is arithmetic; change one thing and fix the seed; a model learns what your images have in common; and the console tells you what happened.

Do I need to know Python for Part 3?

The course does not assume Python experience. You read and type a little: about twenty lines for the custom node, twelve for the API script and six for the LoRA inspector. Each one is walked through on a slide first.

Is the course narrated by AI?

No. I recorded every module in my own voice, ad-libs and dad jokes included.

How long is the whole course?

Just over ten hours in 28 modules: about 3 h 05 min for Part 1, about 3 h 25 min for Part 2 and about 3 h 40 min for Part 3.

How much does the ComfyUI course cost?

Part 1 is free. Parts 2 and 3 are sold together as one bundle, as a single purchase on this page. The bundle costs €29, as one purchase.

Can I follow the course without a GPU?

Yes, with a rented cloud GPU. Part 1 names that route in three steps and leaves the setup to Part 2, whose last module covers renting properly: a network volume so models download once, a six-step setup, and ballpark costs per finished image, clip and training run.

Is this course for complete beginners?

Part 1 assumes you have never opened ComfyUI. It starts with what ComfyUI is and whether your computer can run it, and you make your first image in Module 4. Parts 2 and 3 assume you have done Part 1 or already know that material.

What if I get stuck?

Part 1 has a whole module on the eight errors beginners hit and how to ask for help so you get an answer. If something on your own machine is fighting you, you can book a 60-minute 1-on-1 session and fix it live on a screen share.

See you in the studio

Understand the machine before you copy a single workflow.

Part 1 costs nothing. Watch the first module and decide for yourself.