ComfyUI glossary · 41 terms

ComfyUI terms, in plain English

Every word you meet in your first week of ComfyUI: the files you download, the nodes in a basic workflow, the settings inside the KSampler, and the canvas and hardware words. Each one has a plain definition, and most add the “photographer” version from the course and the one thing to remember. It is the terminology module of ComfyUI From Zero, Part 1, which is free.

What is ComfyUI?

ComfyUI is a free, open-source program that runs image and video models on your own computer or on a rented GPU. Instead of pressing one Generate button, you build the process as a node graph: boxes that each do one job, wired into a workflow. You can load, change and share workflows as JSON files.

The one metaphor
  • The model is the photographer. ComfyUI is the studio. You can swap the photographer and keep the studio.
  • The text encoder is the translator, the VAE is the darkroom, the latent is the negative, and the KSampler is the shoot itself.
  • You do not need to memorise any of this. You need to know where to look: the same content fits on the free two-page cheat sheet (PDF).
10 terms

The files you download

Model (checkpoint)

A model (checkpoint) is a big file of numbers that learned what images look like. It is the thing that actually "draws".

The photographer version: the photographer you hire. Different models are different photographers with different styles and different day rates (in VRAM).

Remember: a workflow without a model is a studio with nobody in it. Goes in models/checkpoints.

Diffusion model (the "bare" model)

A diffusion model (the "bare" model) is the same model that "draws", but shipped alone, without the two helpers below: the text encoder and the VAE. Newer models mostly come this way.

The photographer version: she shows up with just her camera. You have to hire the translator and the darkroom separately.

Remember: goes in models/diffusion_models, not checkpoints. Wrong folder = "value not in list".

Text encoder (CLIP)

A text encoder (CLIP) is the part that turns your words into numbers the model understands. "CLIP" is the old name for one kind of text encoder; ComfyUI still calls the slot CLIP even when the file inside is something else (T5, Qwen, Gemma).

The photographer version: the translator. You speak human, the photographer speaks maths. The translator sits between you.

Remember: a checkpoint has the translator built in. A diffusion model needs one from models/text_encoders.

VAE

A VAE (Variational AutoEncoder) is the part that converts between the small, compressed "latent" the model works in and the actual pixels you see. Two directions: encode (pixels → latent) and decode (latent → pixels).

The photographer version: the darkroom. The photographer works on the negative; the darkroom develops it into a photo you can look at.

Remember: built into checkpoints; separate file for diffusion models (models/vae). Wrong VAE family = weird colours or a shape error.

Go deeper: LTX 2.5 VAE decode: flicker and the Windows stall

Latent

A latent is the compressed working version of an image, about 8× smaller in each direction. All the generation happens here; you only see pixels at the end.

The photographer version: the negative. Not viewable by itself, but everything happens on it.

Remember: "Empty Latent Image" = blank negative of a given size. Sizes should be multiples of 8 (some models want 16, 32 or 64).

LoRA

A LoRA is a small add-on file that nudges a model toward a style, a character, or a skill. It doesn't replace the model; it rides on top of it. It must match the model family.

The photographer version: the photographer took a weekend class. Same camera, new trick.

Remember: models/loras; needs its trigger word in the prompt; strength 0.6 to 1.0 to start. (Using them: Part 2. Training them: Part 3.)

Go deeper: MiniMax H3 2.5D style LoRA: which strength to use and Character LoRA dataset: the five views and the rules

ControlNet

A ControlNet is a helper network that lets you steer with an image (a pose, an edge map, a depth map) instead of only words.

The photographer version: you hand her a reference photo: "stand like this."

Remember: prompt says WHAT, ControlNet says WHERE. models/controlnet. (Part 2.)

Upscaler model

An upscaler model is a small model that enlarges an image and invents plausible detail.

The photographer version: the lab that prints your photo bigger without it going blurry.

Remember: models/upscale_models. Different from "upscale latent", which is just resizing the negative.

FP16 / FP8 / GGUF (Q8, Q6, Q4)

FP16, FP8 and GGUF (Q8, Q6, Q4) are ways of saving the same model with more or fewer decimal places. Smaller file, less VRAM, slightly rougher output. Q8 is nearly perfect, Q4 is visibly rougher.

The photographer version: the same photo saved at 4K, 1080p, 720p. The photographer isn't dumber in the small file, just a bit blurrier.

Remember: file size + 2 GB ≈ VRAM needed. Step down until it fits. GGUF files need the GGUF loader nodes.

safetensors

Safetensors is the modern file format for models. Its point is safety: it cannot contain hidden code. Older .ckpt files could.

Remember: prefer .safetensors always. It is not "faster", it is safer.

12 terms

The nodes in a basic workflow

Node

A node is a box that does one job. Inputs on the left, outputs on the right, settings in the middle.

The photographer version: one person on the crew, with one task.

Remember: double-click the canvas to search for one.

Workflow

A workflow is all the nodes and links together, saved as a JSON file. It is also hidden inside every PNG ComfyUI saves.

The photographer version: the shot list for the whole day.

Remember: workflows don't contain models. Drag a PNG onto the canvas to get its workflow back.

Load Checkpoint

Load Checkpoint is the node that loads a checkpoint and gives you three outputs: MODEL, CLIP, VAE.

The photographer version: hire the photographer; she brings the translator and darkroom.

Load Diffusion Model / Load CLIP / Load VAE

Load Diffusion Model, Load CLIP and Load VAE together are the three-node version of Load Checkpoint, for models shipped in pieces. Load CLIP has a "type" dropdown that must match the model family.

The photographer version: hire the photographer, the translator, and book the darkroom, separately.

Remember: GGUF files use "Unet Loader (GGUF)" and "CLIP Loader (GGUF)" instead.

CLIP Text Encode (positive and negative)

CLIP Text Encode is the node that turns a prompt into "conditioning". You usually have two: positive for what you want, negative for what you don't.

The photographer version: the brief. "Sunset, warm light" in one hand, "no people, no text" in the other.

Remember: the two nodes are identical. The KSampler socket you plug into decides which is which.

Conditioning

Conditioning is the encoded prompt (the orange cable in the default theme): what the sampler steers toward or away from.

Remember: some modern models ignore the negative; the workflow then uses a "Conditioning Zero Out" node as a polite empty plug.

Empty Latent Image

Empty Latent Image is the node that makes a blank latent of the size you choose (width, height, batch).

The photographer version: choosing the blank photo paper. Bigger paper costs more time.

Remember: replaced by Load Image → VAE Encode when you start from an existing picture.

KSampler

The KSampler is the node that actually generates. It takes model + positive + negative + latent, runs the denoising steps and outputs a finished latent.

The photographer version: the shoot itself.

Remember: every setting you'll ever argue about lives here (see the KSampler settings below).

VAE Decode / VAE Encode

VAE Decode is the node that turns a latent into pixels (always at the end); VAE Encode turns pixels into a latent (when starting from an image).

The photographer version: the darkroom, in both directions.

Save Image / Preview Image

Save Image and Preview Image are output nodes: Save writes the picture to output/, Preview shows it and writes to temp/, which is cleared.

Remember: "prompt has no output" error = you forgot one of these.

Load Image

Load Image is the node that brings a picture from input/ onto the canvas. It outputs IMAGE and MASK.

Remember: the mask editor lives here (right-click).

6 terms

The settings inside the KSampler

Seed

The seed is the random starting noise, as a number. Same seed + same settings = same image.

The photographer version: which roll of the dice you started with.

Remember: "control after generate": fixed / increment / randomize. Fix it when comparing.

Steps

Steps is the number of denoising passes the KSampler runs. More = cleaner, slower, then no better.

The photographer version: how many wipes of the foggy window.

Remember: classic models 20 to 30; turbo models 4 to 8. Use the model page's number.

CFG

CFG (classifier-free guidance) is how hard the model is pushed to obey the prompt (and away from the negative).

The photographer version: the tension on the rope. Positive pulls, negative pushes.

Remember: turbo/distilled models want CFG 1, which is why their workflows have no negative box. Too high = burnt colours, six fingers.

Sampler

The sampler is the KSampler setting that decides HOW the noise gets removed each step (Euler, DPM++ 2M, …).

The photographer version: the photographer's walking style down the stairs.

Remember: use what the model page recommends. Euler is the safe default.

Go deeper: MiniMax H3 sampler and scheduler settings and LTX 2.5 sampler and scheduler settings

Scheduler

The scheduler is the KSampler setting that decides WHEN the noise gets removed: how much per step (simple, karras, beta, …).

The photographer version: the height of each stair.

Go deeper: MiniMax H3 sampler and scheduler settings and LTX 2.5 sampler and scheduler settings

Denoise

Denoise is how much of the input is allowed to change, 0 to 1. It only matters when you start from an image.

The photographer version: "redraw this sketch, but keep 70% of it."

Remember: 1.0 = ignore the input entirely (the usual mistake). 0.3 to 0.7 for image-to-image.

8 terms

Canvas words

Group

A group is a coloured rectangle that moves nodes together. Cosmetic.

Bypass (purple, Ctrl+B)

Bypass is skipping a node while the wire runs through it. Use it for A/B tests.

Mute (grey, Ctrl+M)

Mute is scissors on the wire: the nodes downstream error. Rarely useful.

Subgraph

A subgraph is a set of nodes folded into one box you can open. Double-click to look inside.

Template

A template is a saved copy-paste of nodes.

Custom node

A custom node is a folder of Python someone published on GitHub, living in custom_nodes/. Delete the folder = uninstall.

Manager

The Manager is the built-in app store for custom nodes. "Install Missing Custom Nodes" fixes red nodes.

Queue / Run

Run is the button that adds the workflow to a queue; ComfyUI only re-runs nodes whose inputs changed.

5 terms

Hardware words

VRAM

VRAM is the memory on the graphics card: the photographer's desk. Everything fast happens here.

RAM

RAM is system memory: the filing cabinet in the hallway. Overflow goes here and slows down.

Offloading / dynamic VRAM

Offloading (dynamic VRAM) is ComfyUI moving pieces between the desk and the hallway (VRAM and RAM) automatically. Usually leave it alone; it's a Part 2 topic.

Attention (sage / flash / pytorch)

Attention backends (sage, flash, pytorch) are different maths engines for the same job. Faster ones can break some models (black image, buzzing audio). If something looks broken, try the default.

CUDA

CUDA is Nvidia's language for talking to the GPU. "CUDA out of memory" = desk is full.

8 folders

Where do model files go in ComfyUI?

Each kind of model file goes in its own subfolder of ComfyUI/models; a file in the wrong subfolder, or one added without pressing R to refresh, gives the error "value not in list".

Paths are inside your ComfyUI folder, so models/checkpoints is ComfyUI/models/checkpoints.
Kind of fileFolderExamples
Checkpointsmodels/checkpointsAll-in-one models with the text encoder and VAE built in, such as z_image_turbo_aio.safetensors
Diffusion modelsmodels/diffusion_modelsBare models shipped without a text encoder or VAE, such as minimax_h3_ref2va_pruned_int8_convrot.safetensors (MiniMax H3)
Text encodersmodels/text_encodersThe text encoder a diffusion model needs, such as gemma4-12b-with-proj-ltx-2.5-comfy-int8-convrot.safetensors (LTX 2.5)
VAEmodels/vaeSeparate VAE files, such as ltx-2.5-video-vae-bf16.safetensors and ltx-2.5-audio-vae-bf16.safetensors (LTX 2.5)
LoRAsmodels/lorasSmall add-on files, such as minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16.safetensors (MiniMax H3 turbo LoRA)
ControlNetsmodels/controlnetHelper networks that steer with a pose, an edge map or a depth map
Upscalersmodels/upscale_modelsSmall models that enlarge the decoded image and invent plausible detail
Latent upscalersmodels/latent_upscale_modelsLearned models that upscale in latent space, such as ltx-2.5-latent-spatial-upscaler-x2-bf16-1.0.safetensors (LTX 2.5)
Want the why, not just the workflow?

Learn it, fix it live, or have it made.

The ComfyUI From Zero course explains the machine from the first node to training your own LoRA, and Part 1 is free. If something is fighting you right now, bring it to a 60-minute 1-on-1.