DiffHub Logo
HomeGlossaryGuidesClassesComparison
DiffHub Logo
DiffHub Logo

Master ComfyUI and diffusion models with comprehensive education and community support.

Resources

  • Glossary
  • Classes
  • Creator Plan
  • Service Comparison
  • Contact
  • About

Stay Updated

Get the latest tutorials, tips, and news delivered to your inbox.

No spam, unsubscribe at any time.

This site contains affiliate links. We may earn a commission from purchases made through these links at no additional cost to you.

© 2026 DiffHub. All rights reserved.
Privacy PolicyTerms of ServiceImprintSupport

Glossary

Comprehensive definitions of terms related to ComfyUI, diffusion models, LLMs, image, video, audio, and 3D generation models, and upscalers

Diffusion Model
AI/ML

A generative model that learns to reverse a noise process to generate data. It works by gradually removing noise from random data to create meaningful outputs.

Example:

Stable Diffusion uses a diffusion model to generate images from text prompts.

Related Terms:

Stable Diffusion
VAE
UNet
Noise Schedule
View Source
ComfyUI
Interface

A powerful and modular stable diffusion GUI and backend. It uses a node-based interface for creating complex workflows.

Example:

ComfyUI allows users to create custom image generation pipelines using a visual node editor.

Related Terms:

Node
Workflow
Stable Diffusion
Interface
View Source
VAE
Architecture

Variational Autoencoder - a neural network that compresses images into a latent space and can decode them back to pixel space.

Example:

The VAE in Stable Diffusion converts between latent representations and actual images.

Related Terms:

Latent Space
Autoencoder
Compression
Decoding
View Source
UNet
Architecture

A U-shaped neural network architecture commonly used in diffusion models for denoising. It processes data through downsampling and upsampling layers.

Example:

The UNet in Stable Diffusion is responsible for the actual denoising process that generates images.

Related Terms:

Diffusion Model
Denoising
Neural Network
Architecture
View Source
Latent Space
AI/ML

A compressed representation of data in a lower-dimensional space. In diffusion models, images are processed in latent space for efficiency.

Example:

Stable Diffusion works in a 64x64 latent space instead of the full 512x512 pixel space.

Related Terms:

VAE
Compression
Representation
Dimensionality
View Source
Prompt
Interface

Text input that describes what you want to generate. The model uses this to guide the image generation process.

Example:

A prompt like 'a beautiful sunset over mountains' tells the model what kind of image to create.

Related Terms:

Text Encoder
CLIP
Conditioning
Input
View Source
CFG Scale
Parameters

Classifier-Free Guidance scale - controls how closely the model follows the prompt. Higher values make the model follow the prompt more strictly.

Example:

A CFG scale of 7.5 is often a good starting point for most image generation tasks.

Related Terms:

Guidance
Prompt
Conditioning
Parameters
View Source
Sampling Steps
Parameters

The number of denoising steps the model takes to generate an image. More steps generally mean higher quality but longer generation time.

Example:

20 sampling steps is a common setting that balances quality and speed.

Related Terms:

Denoising
Quality
Speed
Iterations
View Source
Checkpoint
Models

A saved model file containing the trained weights of a neural network. Different checkpoints can produce different styles and capabilities.

Example:

The Stable Diffusion 1.5 checkpoint is widely used for general-purpose image generation.

Related Terms:

Model
Weights
Training
File
View Source
LoRA
Models

Low-Rank Adaptation - a technique for fine-tuning large models efficiently. LoRA files can add specific styles or concepts to a base model.

Example:

A LoRA trained on anime characters can be applied to make any model generate anime-style images.

Related Terms:

Fine-tuning
Adaptation
Style
Training
View Source
Stable Diffusion
Models

A latent diffusion model for generating high-quality images from text descriptions. It combines diffusion models with latent space processing for efficiency.

Example:

Stable Diffusion can generate photorealistic images from prompts like 'a cat sitting on a windowsill'.

Related Terms:

Diffusion Model
Latent Space
Text-to-Image
OpenAI
View Source
CLIP
Architecture

Contrastive Language-Image Pre-training - a neural network that learns to associate images with text descriptions.

Example:

CLIP is used in Stable Diffusion to encode text prompts into embeddings that guide image generation.

Related Terms:

Text Encoder
Embedding
Multimodal
OpenAI
View Source
Text Encoder
Architecture

A neural network component that converts text prompts into numerical embeddings that can guide the image generation process.

Example:

The text encoder in Stable Diffusion uses CLIP to convert 'a red car' into a vector representation.

Related Terms:

CLIP
Embedding
Prompt
Encoding
View Source
Noise Schedule
Parameters

A predefined sequence that determines how much noise is added at each step of the diffusion process.

Example:

Different noise schedules can affect the quality and style of generated images.

Related Terms:

Diffusion Model
Denoising
Steps
Process
View Source
Denoising
Process

The process of removing noise from data. In diffusion models, this is the core mechanism for generating images.

Example:

The UNet performs denoising by predicting and removing noise at each sampling step.

Related Terms:

UNet
Diffusion Model
Noise
Generation
View Source
Sampler
Parameters

An algorithm that determines how the denoising process is performed. Different samplers can produce different results.

Example:

DPM++ 2M Karras is a popular sampler that balances quality and speed.

Related Terms:

Sampling Steps
Algorithm
Denoising
Quality
View Source
Seed
Parameters

A random number that initializes the generation process. The same seed with the same prompt will produce the same image.

Example:

Setting seed to 42 will always generate the same image for a given prompt and settings.

Related Terms:

Random
Reproducibility
Generation
Initialization
View Source
Negative Prompt
Interface

Text that describes what you don't want in the generated image. It helps guide the model away from unwanted elements.

Example:

Using 'blurry, low quality' as a negative prompt helps avoid generating poor quality images.

Related Terms:

Prompt
Guidance
Quality
Control
View Source
Embedding
AI/ML

A numerical representation of data in a high-dimensional space. Text and images are converted to embeddings for processing.

Example:

The text 'sunset' is converted to a 768-dimensional embedding vector by CLIP.

Related Terms:

CLIP
Text Encoder
Vector
Representation
View Source
Workflow
ComfyUI

A sequence of connected nodes in ComfyUI that defines how images are processed and generated.

Example:

A workflow might include nodes for loading models, encoding prompts, sampling, and saving images.

Related Terms:

Node
ComfyUI
Process
Pipeline
View Source
Node
ComfyUI

A visual component in ComfyUI that performs a specific function, such as loading models or processing images.

Example:

The 'Load Checkpoint' node loads a Stable Diffusion model, while the 'KSampler' node generates images.

Related Terms:

Workflow
ComfyUI
Function
Component
View Source
ControlNet
Models

A neural network that allows precise control over image generation by using additional input conditions like poses or edges.

Example:

ControlNet can generate images that follow specific poses or architectural layouts.

Related Terms:

Conditioning
Control
Pose
Structure
View Source
Inpainting
Process

The process of filling in or modifying specific parts of an existing image while keeping the rest unchanged.

Example:

Inpainting can be used to remove objects from photos or add new elements to specific areas.

Related Terms:

Mask
Editing
Modification
Partial
View Source
Outpainting
Process

The process of extending an image beyond its original boundaries by generating new content.

Example:

Outpainting can extend a landscape photo to show more of the surrounding area.

Related Terms:

Extension
Boundary
Expansion
Generation
View Source
Upscaling
Process

The process of increasing the resolution of an image using AI models to add detail and improve quality.

Example:

Upscaling can convert a 512x512 image to 1024x1024 or higher resolution.

Related Terms:

Resolution
Quality
Enhancement
Super-resolution
View Source
Face Restoration
Process

The process of improving the quality and detail of faces in generated or low-quality images.

Example:

Face restoration can fix blurry faces or add missing facial details in generated images.

Related Terms:

Quality
Enhancement
Face
Detail
View Source
Style Transfer
Process

The process of applying the artistic style of one image to another while preserving the content.

Example:

Style transfer can make a photo look like a Van Gogh painting.

Related Terms:

Style
Artistic
Transfer
Aesthetic
View Source
Hypernetwork
Models

A small neural network that modifies the behavior of a larger model to achieve specific styles or effects.

Example:

A hypernetwork can be trained to make any model generate images in a specific artistic style.

Related Terms:

Modification
Style
Training
Adaptation
View Source
Textual Inversion
Models

A technique that learns to represent specific concepts or styles as text embeddings that can be used in prompts.

Example:

Textual inversion can learn to represent a specific person's face as a new word that can be used in prompts.

Related Terms:

Embedding
Concept
Learning
Personalization
View Source
DreamBooth
Models

A technique for fine-tuning diffusion models to generate images of specific subjects using just a few example images.

Example:

DreamBooth can teach a model to generate images of your pet using just 3-5 photos.

Related Terms:

Fine-tuning
Personalization
Subject
Training
View Source
IP-Adapter
Models

A model that allows image prompts to guide text-to-image generation, enabling style and content transfer.

Example:

IP-Adapter can use a reference image to guide the style of a generated image while following a text prompt.

Related Terms:

Image Prompt
Style Transfer
Conditioning
Reference
View Source
Latent Diffusion
Architecture

A diffusion model that operates in latent space rather than pixel space, making it more efficient for high-resolution image generation.

Example:

Stable Diffusion is a latent diffusion model that works in 64x64 latent space for 512x512 images.

Related Terms:

Diffusion Model
Latent Space
Efficiency
Resolution
View Source
Cross-Attention
Architecture

A mechanism in neural networks that allows different modalities (like text and images) to interact and influence each other.

Example:

Cross-attention in Stable Diffusion allows text prompts to guide the image generation process.

Related Terms:

Attention
Multimodal
Interaction
Guidance
View Source
Self-Attention
Architecture

A mechanism that allows neural networks to focus on different parts of the input data when making predictions.

Example:

Self-attention helps the model understand relationships between different parts of an image or text.

Related Terms:

Attention
Relationships
Focus
Mechanism
View Source
Transformer
Architecture

A neural network architecture based on attention mechanisms that has revolutionized natural language processing and computer vision.

Example:

CLIP uses a transformer architecture to understand the relationship between text and images.

Related Terms:

Attention
Architecture
Neural Network
CLIP
View Source
Residual Connection
Architecture

A connection that allows information to flow directly from one layer to another, helping with training deep networks.

Example:

Residual connections in UNet help preserve important features during the denoising process.

Related Terms:

Connection
Training
Deep Network
Information Flow
View Source
Batch Normalization
Architecture

A technique that normalizes the inputs to each layer, helping with training stability and convergence.

Example:

Batch normalization is used throughout the UNet to ensure stable training.

Related Terms:

Normalization
Training
Stability
Convergence
View Source
Dropout
Architecture

A regularization technique that randomly sets some neurons to zero during training to prevent overfitting.

Example:

Dropout is used in various parts of the diffusion model to improve generalization.

Related Terms:

Regularization
Overfitting
Training
Generalization
View Source
Learning Rate
Training

A hyperparameter that controls how much the model weights are updated during training.

Example:

A learning rate of 0.0001 is commonly used for fine-tuning diffusion models.

Related Terms:

Training
Hyperparameter
Update
Optimization
View Source
Gradient Descent
Training

An optimization algorithm that iteratively adjusts model parameters to minimize the loss function.

Example:

Gradient descent is used to train diffusion models by minimizing the difference between predicted and actual noise.

Related Terms:

Optimization
Training
Loss Function
Parameters
View Source
Loss Function
Training

A function that measures how well the model's predictions match the actual data, used to guide training.

Example:

Diffusion models use a loss function that measures the difference between predicted and actual noise.

Related Terms:

Training
Prediction
Measurement
Optimization
View Source
Overfitting
Training

When a model learns the training data too well and performs poorly on new, unseen data.

Example:

An overfitted diffusion model might generate images that look exactly like the training data but fail on new prompts.

Related Terms:

Training
Generalization
Performance
Data
View Source
Underfitting
Training

When a model is too simple to capture the underlying patterns in the data.

Example:

An underfitted diffusion model might generate blurry or low-quality images regardless of the prompt.

Related Terms:

Training
Complexity
Patterns
Quality
View Source
Data Augmentation
Training

Techniques used to artificially increase the size of the training dataset by creating variations of existing data.

Example:

Data augmentation for images might include rotation, scaling, or color adjustments.

Related Terms:

Training
Dataset
Variation
Techniques
View Source
Transfer Learning
Training

A technique where a model trained on one task is adapted for use on a different but related task.

Example:

Fine-tuning a pre-trained Stable Diffusion model for a specific art style is an example of transfer learning.

Related Terms:

Training
Adaptation
Pre-trained
Task
View Source
Fine-tuning
Training

The process of adapting a pre-trained model to a specific task or dataset by training it further.

Example:

Fine-tuning Stable Diffusion on a dataset of anime images to make it generate anime-style artwork.

Related Terms:

Training
Adaptation
Pre-trained
Specific
View Source
Pre-training
Training

The initial training phase where a model learns general features from a large, diverse dataset.

Example:

Stable Diffusion was pre-trained on millions of image-text pairs from the internet.

Related Terms:

Training
General
Large Dataset
Features
View Source
Inference
Process

The process of using a trained model to make predictions or generate new data.

Example:

Running Stable Diffusion to generate an image from a text prompt is an inference process.

Related Terms:

Prediction
Generation
Trained Model
Output
View Source
GPU
Hardware

Graphics Processing Unit - specialized hardware that can perform many calculations in parallel, essential for AI model training and inference.

Example:

Training and running Stable Diffusion requires a powerful GPU with sufficient VRAM.

Related Terms:

Hardware
Parallel
Training
Inference
View Source
VRAM
Hardware

Video Random Access Memory - the memory on a graphics card that stores data for GPU processing.

Example:

Running Stable Diffusion typically requires at least 4GB of VRAM, with 8GB+ recommended for optimal performance.

Related Terms:

GPU
Memory
Graphics Card
Performance
View Source
SAM (Segment Anything Model)
Models

A promptable segmentation model by Meta AI that can segment any object in an image. It was trained on 11M images and 1B masks, providing zero-shot segmentation capabilities.

Example:

SAM can be used in ComfyUI workflows to automatically detect and segment faces or objects for targeted processing.

Related Terms:

Segmentation
Face Detailer
Mask
Object Detection
View Source
YOLO (You Only Look Once)
Models

A family of real-time object detection models that can identify and locate multiple objects in images. YOLO processes the entire image in a single pass, making it extremely fast.

Example:

YOLOv8 can detect faces, people, or other objects in generated images for post-processing refinement.

Related Terms:

Object Detection
Bounding Box
Face Detection
Real-time
View Source
YOLOv8
Models

The latest version of YOLO by Ultralytics featuring an anchor-free approach, CSPNet backbone for enhanced feature extraction, and FPN+PAN neck for superior multi-scale object detection.

Example:

YOLOv8 is used in face detailing workflows to accurately detect faces before applying enhancement.

Related Terms:

YOLO
Object Detection
Face Detection
Ultralytics
View Source
Bounding Box
Process

A rectangular annotation that defines the location and size of an object within an image. Bounding boxes are defined by coordinates (x, y, width, height) or corner coordinates.

Example:

Face detailer nodes use bounding boxes to identify face regions before applying enhancement algorithms.

Related Terms:

Object Detection
YOLO
Detection
Coordinates
View Source
Segmentation
Process

The process of partitioning an image into multiple segments or regions, typically to identify objects and boundaries. Can be semantic (classifying pixels) or instance-based (identifying individual objects).

Example:

Image segmentation is used to create precise masks for inpainting or selective image editing.

Related Terms:

SAM
Mask
Object Detection
Instance Segmentation
View Source
DPM++ Sampler
Parameters

DPM-Solver++ is a high-order solver for diffusion models that can generate high-quality samples in 15-20 steps. It solves the diffusion ODE with improved efficiency and quality.

Example:

DPM++ 2M Karras is a popular sampler choice in ComfyUI for balancing speed and image quality.

Related Terms:

Sampler
Sampling Steps
Karras Scheduler
Quality
View Source
Karras Scheduler
Parameters

A noise schedule based on the paper 'Elucidating the Design Space of Diffusion-Based Generative Models' by Karras et al. It applies a smaller amount of noise per step near the end of sampling for improved quality.

Example:

Using the Karras scheduler with DPM++ sampler often produces higher quality results than other noise schedules.

Related Terms:

Noise Schedule
Sampler
DPM++
Quality
View Source
Wildcards
Interface

A templating system for prompts that allows random selection from predefined lists of terms. Wildcards use the syntax __filename__ to insert random values from text files.

Example:

Using __hairstyle__ in a prompt might randomly select from 'long hair', 'short hair', 'braided', etc., creating prompt variety.

Related Terms:

Prompt
Dynamic Prompts
Randomization
Template
View Source
Tiled Diffusion
Process

A technique that divides large images into smaller tiles, processes each independently, and seamlessly stitches them together. Based on MultiDiffusion and Mixture of Diffusers algorithms.

Example:

Tiled diffusion allows upscaling images to 4K or 8K resolution without running out of VRAM.

Related Terms:

Upscaling
VRAM
MultiDiffusion
Large Images
View Source
MultiDiffusion
Architecture

A method for fusing diffusion paths to enable controlled image generation at high resolutions. It allows for panorama generation and region-based text control by processing overlapping tiles.

Example:

MultiDiffusion enables generating ultra-high resolution images by processing them in overlapping tiles.

Related Terms:

Tiled Diffusion
High Resolution
Panorama
Controlled Generation
View Source
Face Detailer
Process

A workflow component that detects faces in generated images and applies targeted enhancement, including upscaling, denoising, and feature refinement to improve facial quality.

Example:

Face detailer can fix blurry or distorted faces in group shots or distant subjects.

Related Terms:

Face Restoration
YOLO
SAM
Enhancement
View Source
Euler Ancestral
Parameters

A stochastic sampler that uses the Euler method with added noise at each step. The 'ancestral' variant adds randomness, producing more varied results but less deterministic outputs.

Example:

Euler Ancestral sampler is often used for creative exploration due to its stochastic nature.

Related Terms:

Sampler
Stochastic
Euler
Sampling Steps
View Source
Feathering
Process

A mask processing technique that softens the edges of a mask by gradually transitioning from opaque to transparent. This creates smoother blending between masked and unmasked regions.

Example:

Feathering a face mask by 20 pixels prevents harsh edges when applying face detailing.

Related Terms:

Mask
Inpainting
Blending
Edge Softening
View Source
Detection Threshold
Parameters

A confidence value (0.0 to 1.0) that determines the minimum certainty required for an object detector to report a detection. Higher thresholds reduce false positives but may miss objects.

Example:

Setting face detection threshold to 0.75 means only faces detected with 75% confidence or higher will be processed.

Related Terms:

Object Detection
YOLO
Confidence
False Positives
View Source
Dilation
Process

A morphological operation that expands the boundaries of regions in a binary image or mask. In object detection, it's used to expand bounding boxes or masks to include surrounding context.

Example:

Dilating a face bounding box by 10 pixels ensures hair and neck are included in face detailing.

Related Terms:

Mask
Morphological Operations
Bounding Box
Expansion
View Source
KSampler
ComfyUI

A core ComfyUI node that performs the denoising process to generate images. It controls sampling steps, CFG scale, sampler type, scheduler, and seed for the generation process.

Example:

The KSampler node is typically connected between the model loader and VAE decoder in a workflow.

Related Terms:

Sampling
Denoising
Workflow
Node
View Source
Ultimate SD Upscale
Process

An advanced upscaling technique that uses tiled diffusion to upscale images to very high resolutions. It includes seam fixing and supports multiple upscale models.

Example:

Ultimate SD Upscale can upscale a 512x512 image to 2048x2048 or higher while adding new details.

Related Terms:

Upscaling
Tiled Diffusion
Resolution
Enhancement
View Source
Mixture of Diffusers
Architecture

A technique for high-resolution image generation that combines multiple diffusion models or regions. Each region can have different prompts or settings, enabling complex scene composition.

Example:

Mixture of Diffusers allows creating a landscape where the sky and ground are generated with different prompts.

Related Terms:

MultiDiffusion
Tiled Diffusion
Composition
High Resolution
View Source
Instance Segmentation
Process

A computer vision task that identifies each distinct object in an image and creates a separate segmentation mask for each instance, even for objects of the same class.

Example:

Instance segmentation can distinguish between multiple people in a photo, creating separate masks for each person.

Related Terms:

Segmentation
Object Detection
SAM
Mask
View Source
Denoising Strength
Parameters

A parameter (0.0 to 1.0) that controls how much the AI modifies an input image. Lower values preserve more of the original, while higher values allow more creative freedom.

Example:

A denoising strength of 0.3 in face detailer makes subtle improvements while 0.7 allows more dramatic changes.

Related Terms:

Denoising
img2img
Strength
Modification
View Source
Mesh
3D Generation

A collection of vertices, edges, and faces that defines the surface geometry of a 3D object. Meshes are the most common representation for real-time rendering, animation, and 3D printing.

Example:

An AI image-to-3D pipeline outputs a triangle mesh that can be imported into Blender or Unity.

Related Terms:

Vertex
Polygon
Topology
Retopology
View Source
Vertex
3D Generation

A point in 3D space defined by coordinates (x, y, z). Vertices are connected by edges to form faces, which together make up a mesh.

Example:

A cube mesh has 8 vertices at its corners, each with a fixed position in 3D space.

Related Terms:

Mesh
Polygon
Topology
UV Mapping
View Source
Polygon
3D Generation

A flat face on a 3D mesh, typically made of three or more vertices. Triangles and quads are the most common polygon types in 3D graphics.

Example:

Game engines often require meshes to be triangulated, converting all quads into pairs of triangle polygons.

Related Terms:

Mesh
Vertex
Topology
Triangle
View Source
UV Mapping
3D Generation

The process of projecting a 2D texture image onto a 3D mesh surface by assigning 2D coordinates (U, V) to each vertex.

Example:

After generating a 3D model, UV unwrapping maps the AI-generated texture onto the mesh without stretching.

Related Terms:

Texture Map
Mesh
Vertex
Material
View Source
PBR (Physically Based Rendering)
3D Generation

A shading approach that simulates how light interacts with real-world materials using physically accurate properties like roughness, metallic, and normal maps.

Example:

Exporting an AI-generated asset with PBR materials ensures it looks correct under different lighting in game engines.

Related Terms:

Normal Map
Material
Texture Map
Rendering
View Source
Normal Map
3D Generation

A texture that stores surface normal direction per pixel, creating the illusion of fine geometric detail without adding polygons.

Example:

A low-poly AI mesh can appear highly detailed when paired with a baked normal map from a high-resolution sculpt.

Related Terms:

PBR (Physically Based Rendering)
Texture Map
Mesh
Baking
View Source
Topology
3D Generation

The structure and flow of polygons across a 3D mesh, including edge loops and face distribution. Good topology is essential for deformation and animation.

Example:

AI-generated meshes often have messy topology that requires manual cleanup before rigging.

Related Terms:

Mesh
Retopology
Polygon
Rigging
View Source
Retopology
3D Generation

The process of rebuilding a mesh with cleaner polygon flow, usually to reduce polygon count or fix topology issues while preserving shape.

Example:

After generating a detailed sculpt with AI, artists retopologize it into a game-ready quad mesh.

Related Terms:

Topology
Mesh
Polygon
LOD
View Source
Rigging
3D Generation

The process of creating a skeleton (armature) and assigning skin weights to a mesh so it can be posed and animated.

Example:

An AI-generated character mesh must be rigged with bones before it can walk or perform actions in animation.

Related Terms:

Mesh
Topology
Animation
Skeleton
View Source
Level of Detail (LOD)
3D Generation

A technique that uses multiple versions of a model at different polygon counts, swapping to simpler meshes at greater distances to save performance.

Example:

A high-poly AI-generated asset is simplified into LOD0, LOD1, and LOD2 variants for real-time use.

Related Terms:

Mesh
Retopology
Polygon
Rendering
View Source
Point Cloud
3D Generation

A set of data points in 3D space, each with position and optionally color or normal information. Point clouds are common intermediate outputs in 3D scanning and reconstruction.

Example:

LiDAR sensors produce point clouds that can be converted into meshes or used directly for visualization.

Related Terms:

Voxel
Photogrammetry
Mesh
Gaussian Splatting
View Source
Voxel
3D Generation

A volumetric pixel — a value on a regular 3D grid. Voxel representations are used in medical imaging, Minecraft-style worlds, and some 3D generation models.

Example:

Some early AI 3D generators represented shapes as occupancy grids of voxels before extracting surfaces.

Related Terms:

Point Cloud
Signed Distance Field (SDF)
Mesh
Marching Cubes
View Source
glTF
3D Generation

GL Transmission Format — an open standard 3D file format optimized for efficient transmission and loading on the web and in real-time engines.

Example:

Exporting an AI-generated model as .glb (binary glTF) allows it to be viewed directly in browsers and AR viewers.

Related Terms:

Mesh
PBR (Physically Based Rendering)
UV Mapping
Material
View Source
Depth Map
3D Generation

A 2D image where each pixel value represents the distance from the camera to the scene surface. Depth maps are key inputs for single-image 3D reconstruction.

Example:

A monocular depth estimation model predicts a depth map from a photo, which is then used to lift the image into 3D.

Related Terms:

Image-to-3D
Novel View Synthesis
Photogrammetry
Point Cloud
View Source
Photogrammetry
3D Generation

The science of extracting 3D measurements and geometry from photographs by analyzing overlapping images of a subject from multiple angles.

Example:

Taking 50 photos around a statue and running photogrammetry software produces a detailed textured mesh.

Related Terms:

Point Cloud
Mesh
Multi-view Reconstruction
Novel View Synthesis
View Source
NeRF (Neural Radiance Field)
3D Generation

A neural network that represents a 3D scene as a continuous volumetric function, encoding color and density at any point in space. NeRFs excel at photorealistic novel view synthesis.

Example:

A NeRF trained on photos of a room can render new camera angles with realistic lighting and reflections.

Related Terms:

Novel View Synthesis
Radiance Field
Gaussian Splatting
Instant-NGP
View Source
Gaussian Splatting
3D Generation

3D Gaussian Splatting (3DGS) represents a scene as millions of colored 3D Gaussians that are rasterized in real time. It achieves NeRF-quality visuals with much faster rendering.

Example:

Converting a NeRF or set of photos into Gaussian splats enables interactive 60fps exploration of a 3D scene in a browser.

Related Terms:

NeRF (Neural Radiance Field)
Point Cloud
Novel View Synthesis
Radiance Field
View Source
Radiance Field
3D Generation

A function that maps any 3D point and viewing direction to emitted color and opacity. NeRFs and Gaussian splats are both radiance field representations.

Example:

Instead of storing explicit geometry, a radiance field encodes how light appears from every viewpoint in a scene.

Related Terms:

NeRF (Neural Radiance Field)
Gaussian Splatting
Novel View Synthesis
Rendering
View Source
Instant-NGP
3D Generation

Instant Neural Graphics Primitives — a method by NVIDIA that uses multi-resolution hash encodings to train NeRFs and other neural fields in seconds instead of hours.

Example:

Instant-NGP can reconstruct a photorealistic NeRF from a few dozen images in under a minute on a consumer GPU.

Related Terms:

NeRF (Neural Radiance Field)
Gaussian Splatting
Multi-view Reconstruction
Radiance Field
View Source
Signed Distance Field (SDF)
3D Generation

A mathematical function that returns the shortest distance from any point in space to the nearest surface, with sign indicating inside or outside. SDFs define smooth implicit 3D shapes.

Example:

AI text-to-3D models like DreamFusion optimize an SDF that is converted to a mesh via marching cubes.

Related Terms:

Marching Cubes
Voxel
Mesh
Text-to-3D
View Source
Marching Cubes
3D Generation

An algorithm that extracts a polygonal mesh from a 3D scalar field (such as an SDF or voxel grid) by identifying surface crossings between grid cells.

Example:

After an AI model predicts a voxel occupancy grid, marching cubes converts it into a triangle mesh for export.

Related Terms:

Signed Distance Field (SDF)
Voxel
Mesh
Polygon
View Source
Text-to-3D
3D Generation

AI generation of 3D assets directly from text descriptions, typically using diffusion models, NeRF optimization, or large 3D datasets.

Example:

Typing 'a wooden chair with carved legs' into a text-to-3D tool produces a downloadable 3D model.

Related Terms:

Image-to-3D
Signed Distance Field (SDF)
DreamFusion
Mesh
View Source
Image-to-3D
3D Generation

AI reconstruction of a 3D model from one or more 2D images, using depth estimation, multi-view diffusion, or neural field optimization.

Example:

Uploading a product photo to an image-to-3D service returns a textured mesh suitable for e-commerce AR.

Related Terms:

Text-to-3D
Depth Map
TripoSR
Mesh
View Source
Novel View Synthesis
3D Generation

Generating images of a scene from camera viewpoints that were not in the original input, requiring the model to understand 3D structure and appearance.

Example:

Given 10 photos of a building, novel view synthesis renders what it looks like from angles between the captured shots.

Related Terms:

NeRF (Neural Radiance Field)
Gaussian Splatting
Photogrammetry
Multi-view Reconstruction
View Source
Multi-view Reconstruction
3D Generation

Building a 3D representation of a scene or object by combining information from multiple images taken from different viewpoints.

Example:

AI multi-view reconstruction takes 4 input images of an object and outputs a consistent 360° 3D model.

Related Terms:

Photogrammetry
Novel View Synthesis
NeRF (Neural Radiance Field)
Image-to-3D
View Source
TripoSR
3D Generation

An open-source image-to-3D model by Tripo AI that generates textured 3D meshes from a single image in seconds using a transformer-based architecture.

Example:

Feeding a photo of a shoe into TripoSR produces an .obj mesh with albedo texture in under 10 seconds.

Related Terms:

Image-to-3D
Mesh
Text-to-3D
Shap-E
View Source
Shap-E
3D Generation

An OpenAI model that generates 3D assets from text or images by decoding implicit functions into textured meshes or NeRFs, trained on a large corpus of 3D data.

Example:

Shap-E can generate a 3D model of 'an ice cream cone' from a text prompt without needing multi-view photos.

Related Terms:

Text-to-3D
Signed Distance Field (SDF)
Mesh
TripoSR
View Source
DreamFusion
3D Generation

A pioneering text-to-3D method that optimizes a NeRF using a pretrained 2D diffusion model as a score distillation signal, generating 3D without 3D training data.

Example:

DreamFusion demonstrated that a text prompt alone could produce a coherent 3D object by distilling knowledge from Stable Diffusion.

Related Terms:

Text-to-3D
NeRF (Neural Radiance Field)
Signed Distance Field (SDF)
Score Distillation
View Source
Score Distillation Sampling (SDS)
3D Generation

A technique that uses gradients from a pretrained 2D diffusion model to optimize a 3D representation, enabling text-to-3D without paired text-3D datasets.

Example:

SDS drives DreamFusion and many follow-up text-to-3D methods by nudging a NeRF toward images the diffusion model would approve of.

Related Terms:

DreamFusion
Text-to-3D
Diffusion Model
NeRF (Neural Radiance Field)
View Source
Mesh Extraction
3D Generation

The post-processing step of converting an implicit or volumetric 3D representation (NeRF, SDF, voxels) into an explicit triangle mesh for editing and export.

Example:

After NeRF training, mesh extraction via marching cubes produces an .obj file that can be edited in Blender.

Related Terms:

Marching Cubes
Mesh
NeRF (Neural Radiance Field)
Signed Distance Field (SDF)
View Source
Large Language Model (LLM)
LLMs

A transformer-based neural network trained on huge amounts of text to predict the next token. At scale this yields models that can write, reason, code, and follow instructions.

Example:

ChatGPT, Claude, and Gemini are LLM-powered assistants; in image workflows LLMs are often used to expand short ideas into detailed prompts.

Related Terms:

Transformer
Token
Context Window
Prompt
View Source
Token
LLMs

The basic unit of text a language model reads and writes, usually a word fragment produced by a tokenizer. Pricing, speed, and context limits are all measured in tokens.

Example:

The word 'diffusion' may be split into two or three tokens; roughly 750 English words fit into 1,000 tokens.

Related Terms:

Large Language Model (LLM)
Context Window
Text Encoder
View Source
Context Window
LLMs

The maximum number of tokens a model can consider at once, covering the prompt, conversation history, attached documents, and its own reply.

Example:

A model with a 200K-token context window can read an entire codebase or a long PDF in a single request.

Related Terms:

Token
Large Language Model (LLM)
View Source
Reasoning Model
LLMs

An LLM trained, usually with reinforcement learning, to think step by step in a chain of thought before answering. It trades extra time and tokens for better results on math, code, and planning.

Example:

OpenAI o-series models, DeepSeek-R1, and Claude with extended thinking all spend extra tokens reasoning before they reply.

Related Terms:

Large Language Model (LLM)
DeepSeek
GPT (OpenAI)
View Source
Mixture of Experts (MoE)
Architecture

An architecture that splits a layer into many 'expert' sub-networks and routes each token to only a few of them. The model gets a large total parameter count while keeping per-token compute low.

Example:

DeepSeek-V3 has 671B total parameters but activates only about 37B per token; Wan 2.2 uses a two-expert MoE for high-noise and low-noise denoising.

Related Terms:

Transformer
DeepSeek
Wan (Wan-Video)
View Source
Open Weights
LLMs

A model whose trained weights are published for download so anyone can run, fine-tune, or quantize it locally. The license may still restrict commercial use, so open weights is not always the same as open source.

Example:

FLUX.1 [dev], Llama, Qwen, and Wan are open-weight models, while Midjourney and GPT Image are only available through an app or API.

Related Terms:

Checkpoint
Quantization
Fine-tuning
View Source
Quantization
Hardware

Storing model weights at lower numerical precision, such as 8-bit or 4-bit instead of 16-bit, to cut VRAM use and speed up inference at a small quality cost.

Example:

An FP8 or Q4 GGUF version of FLUX lets the model run on a 12GB GPU that could not load the full BF16 weights.

Related Terms:

VRAM
GGUF
Inference
View Source
GGUF
Hardware

A single-file model format from the llama.cpp / ggml project designed for quantized weights. It is the standard way to run LLMs locally and is also used for quantized diffusion models in ComfyUI.

Example:

Loading a Q5_K_M GGUF of Wan 2.2 through the ComfyUI-GGUF node pack fits video generation into consumer VRAM.

Related Terms:

Quantization
VRAM
Checkpoint
View Source
Vision-Language Model (VLM)
LLMs

A multimodal model that takes images (and sometimes video) alongside text as input, so it can describe, caption, or answer questions about visual content.

Example:

A VLM like Qwen2.5-VL or Florence-2 can auto-caption a folder of images to build a LoRA training dataset.

Related Terms:

Large Language Model (LLM)
Florence-2
LoRA
View Source
GPT (OpenAI)
LLMs

OpenAI's family of Generative Pre-trained Transformer models behind ChatGPT. It spans GPT-4o, the reasoning-focused o-series, GPT-5 (August 2025), which merged chat and reasoning into one system, and later releases.

Example:

Many prompt-enhancer nodes call a GPT model to turn 'cozy cabin' into a detailed, lighting-aware image prompt.

Related Terms:

Large Language Model (LLM)
Reasoning Model
GPT Image
View Source
gpt-oss
LLMs

OpenAI's open-weight reasoning models released in August 2025 under the Apache 2.0 license, in 20B and 120B MoE sizes.

Example:

gpt-oss-20b runs locally on a 16GB GPU, making it a private option for prompt generation inside a workflow.

Related Terms:

GPT (OpenAI)
Open Weights
Mixture of Experts (MoE)
View Source
Claude (Anthropic)
LLMs

Anthropic's family of LLMs, known for strong coding, long-context work, and agentic tool use. The Claude 5 family includes Claude Fable 5.1, Claude Opus 5.5, and Claude Sonnet 5, alongside the faster Claude Haiku 4.5.

Example:

Claude can write custom ComfyUI nodes or debug a broken workflow JSON file.

Related Terms:

Large Language Model (LLM)
Reasoning Model
Context Window
View Source
Gemini (Google)
LLMs

Google DeepMind's natively multimodal model family that understands text, images, audio, and video. It spans Gemini 2.5 through the Gemini 3 generation, in Pro, Flash, and Flash-Lite tiers.

Example:

Gemini can watch a reference video and describe its camera motion so you can reuse it in a video prompt.

Related Terms:

Large Language Model (LLM)
Vision-Language Model (VLM)
Nano Banana (Gemini Image)
View Source
Gemma
LLMs

Google's family of lightweight open-weight models built from Gemini research. Gemma 3 adds image input and runs on a single GPU or even a laptop.

Example:

Gemma 3 12B is a popular local model for captioning and prompt rewriting in offline pipelines.

Related Terms:

Gemini (Google)
Open Weights
Vision-Language Model (VLM)
View Source
Llama (Meta)
LLMs

Meta's open-weight LLM family that popularized running LLMs locally. Llama 4 (April 2025) introduced natively multimodal MoE models such as Scout and Maverick.

Example:

Many community fine-tunes and early local-LLM tools were built on Llama 2 and Llama 3.

Related Terms:

Open Weights
Mixture of Experts (MoE)
Fine-tuning
View Source
DeepSeek
LLMs

A Chinese lab known for highly efficient open-weight models. DeepSeek-V3 (MoE) and the reasoning model DeepSeek-R1 (January 2025) matched frontier models at a fraction of the training cost, followed by later V3.x and V4 updates.

Example:

DeepSeek-R1 showed that reinforcement learning can teach a model to reason step by step.

Related Terms:

Reasoning Model
Mixture of Experts (MoE)
Open Weights
View Source
Qwen (Alibaba)
LLMs

Alibaba's family of open-weight LLMs and multimodal models, from small edge models to large MoE flagships. It includes Qwen3, the Qwen-VL vision models, and the Qwen-Image generators.

Example:

Qwen2.5-VL is a common choice for captioning training data, and a Qwen language model serves as the text encoder of Qwen-Image.

Related Terms:

Open Weights
Vision-Language Model (VLM)
Qwen-Image
View Source
Mistral
LLMs

A French AI lab releasing efficient open-weight and commercial models, including Mistral Small, Mistral Large, Mixtral (MoE), Codestral for code, and Pixtral for vision.

Example:

Mistral Small is a popular local LLM for prompt expansion because it balances quality and VRAM use.

Related Terms:

Open Weights
Mixture of Experts (MoE)
Large Language Model (LLM)
View Source
Kimi K2
LLMs

Moonshot AI's open-weight MoE model with about 1 trillion total parameters (32B active), released in July 2025 and tuned for agentic coding and tool use.

Example:

Kimi K2 is often used as an open alternative to proprietary models in coding agents.

Related Terms:

Mixture of Experts (MoE)
Open Weights
Large Language Model (LLM)
View Source
GLM (Z.ai)
LLMs

The General Language Model family from Zhipu AI (Z.ai), including the open-weight GLM-4.5 MoE series with strong agentic and coding performance, and GLM-V vision models.

Example:

GLM-4.5-Air is a mid-sized open model that runs on a multi-GPU workstation.

Related Terms:

Open Weights
Mixture of Experts (MoE)
CogVideoX
View Source
Grok (xAI)
LLMs

xAI's LLM family integrated into X. It includes reasoning-focused Grok 4 models and the Grok Imagine image and video generator.

Example:

Grok can be accessed through the xAI API as an OpenAI-compatible endpoint.

Related Terms:

Large Language Model (LLM)
Reasoning Model
View Source
Diffusion Transformer (DiT)
Architecture

A diffusion model that replaces the UNet with a transformer operating on patches of the latent image. It scales better with parameters and data and is the backbone of most modern image and video models.

Example:

FLUX, Stable Diffusion 3, Qwen-Image, Wan, and Sora are all built on DiT-style architectures.

Related Terms:

Transformer
UNet
Latent Diffusion
MMDiT
View Source
MMDiT
Architecture

Multimodal Diffusion Transformer: a DiT variant where text and image tokens are processed in separate streams that share joint attention, improving prompt adherence and text rendering.

Example:

Stable Diffusion 3 introduced MMDiT, and FLUX.1 and Qwen-Image use closely related designs.

Related Terms:

Diffusion Transformer (DiT)
Stable Diffusion 3.5
FLUX
View Source
Flow Matching
Training

A training objective that teaches a model to follow a straight path from noise to data instead of predicting noise at each step. Rectified flow is a popular variant that gives good results in fewer steps.

Example:

FLUX and Stable Diffusion 3 are flow-matching models, which is why they use different samplers and schedulers than SD 1.5.

Related Terms:

Diffusion Model
Noise Schedule
Sampler
View Source
T5 Text Encoder
Architecture

Google's T5 language model, whose encoder half is used by many modern diffusion models to understand long, detailed prompts far better than CLIP alone.

Example:

FLUX.1 and SD3 load a T5-XXL encoder next to CLIP; an FP8 version of T5 saves several GB of VRAM.

Related Terms:

Text Encoder
CLIP
Prompt
View Source
Step Distillation
Training

Training a fast 'student' model to reproduce a slow model's output in far fewer sampling steps. It is behind 'Turbo', 'Lightning', 'Schnell', and 'Hyper' model variants.

Example:

SDXL Turbo generates an image in 1 to 4 steps, and Lightning LoRAs let Wan 2.2 render video in 4 to 8 steps.

Related Terms:

Sampling Steps
LoRA
FLUX
View Source
Stable Diffusion XL (SDXL)
Image Models

Stability AI's 2023 successor to SD 1.5 with a larger UNet, two text encoders, and native 1024x1024 output. It remains one of the most fine-tuned model families thanks to its huge LoRA and checkpoint ecosystem.

Example:

Popular community checkpoints such as Juggernaut XL and anime models like Illustrious and Pony are SDXL-based.

Related Terms:

Stable Diffusion
Checkpoint
LoRA
View Source
Stable Diffusion 3.5
Image Models

Stability AI's MMDiT-based image models released in October 2024 in Large, Large Turbo, and Medium sizes, with improved typography and prompt adherence over SDXL.

Example:

SD 3.5 Medium runs on around 10GB of VRAM and handles multi-subject prompts better than SDXL.

Related Terms:

MMDiT
Stable Diffusion XL (SDXL)
T5 Text Encoder
View Source
FLUX
Image Models

Black Forest Labs' image model family, created by former Stable Diffusion researchers. FLUX.1 (2024) came as [pro], [dev], and [schnell]; FLUX.1 Kontext (2025) added instruction-based editing; FLUX.2 (November 2025) raised quality and added multi-reference editing, followed by the compact FLUX.2 [klein] models.

Example:

FLUX.1 [dev] is the most common base for LoRAs in ComfyUI, and Kontext can 'change the jacket to red' while keeping the person identical.

Related Terms:

MMDiT
Flow Matching
Image Editing Model
View Source
Qwen-Image
Image Models

Alibaba's open-weight image generation and editing models. The original 20B MMDiT release (August 2025) was known for excellent text rendering, including Chinese. Qwen-Image-2.0 (February 2026) unified generation and editing in one model, and Qwen-Image 3.0 (July 2026) improved long prompts and multilingual layouts.

Example:

Qwen-Image-Edit can replace the text on a sign in a photo while matching the original font and lighting.

Related Terms:

MMDiT
Image Editing Model
Qwen (Alibaba)
View Source
HiDream-I1
Image Models

An open-weight 17B sparse diffusion transformer from HiDream.ai released in 2025 under the MIT license, with Full, Dev, and Fast variants.

Example:

HiDream-I1 Full is used when a permissive license and high prompt adherence matter more than speed.

Related Terms:

Diffusion Transformer (DiT)
Open Weights
Mixture of Experts (MoE)
View Source
Chroma
Image Models

A community-trained open model built on FLUX.1 [schnell] with fewer parameters, Apache 2.0 licensing, and restored support for negative prompts and CFG.

Example:

Chroma is popular for fine-tuning because it avoids the FLUX.1 [dev] non-commercial license.

Related Terms:

FLUX
Negative Prompt
CFG Scale
View Source
PixArt-Σ
Image Models

A lightweight, efficiently trained DiT text-to-image model that generates up to 4K images and uses a T5 encoder for detailed prompts.

Example:

PixArt-Σ is sometimes used as a fast first pass before refining details with an SDXL model.

Related Terms:

Diffusion Transformer (DiT)
T5 Text Encoder
View Source
HunyuanImage
Image Models

Tencent Hunyuan's open-weight text-to-image models, including HunyuanImage 2.1 (2025) with native 2K output, and the larger multimodal HunyuanImage 3.0.

Example:

HunyuanImage 2.1 produces 2048x2048 images directly without an upscaling pass.

Related Terms:

Diffusion Transformer (DiT)
Open Weights
HunyuanVideo
View Source
Image Editing Model
Image Models

A model that takes an existing image plus a text instruction and returns a modified image, replacing many masking and inpainting workflows with a single prompt.

Example:

FLUX.1 Kontext, Qwen-Image-Edit, Nano Banana, and GPT Image can all follow edits like 'remove the people in the background'.

Related Terms:

Inpainting
FLUX
Qwen-Image
View Source
GPT Image
Image Models

OpenAI's natively multimodal image generator built into GPT models (GPT-4o image generation, gpt-image-1 in the API). It excels at following complex instructions, rendering text, and conversational editing.

Example:

Asking ChatGPT to 'turn this sketch into a product photo with the logo on the box' uses GPT Image.

Related Terms:

GPT (OpenAI)
Image Editing Model
View Source
Nano Banana (Gemini Image)
Image Models

The nickname for Google's Gemini image generation and editing models, starting with Gemini 2.5 Flash Image. They are known for keeping characters consistent across edits and blending multiple reference images.

Example:

You can upload a portrait and a jacket photo and ask Nano Banana to dress the person in that jacket.

Related Terms:

Gemini (Google)
Image Editing Model
Imagen
View Source
Imagen
Image Models

Google DeepMind's dedicated text-to-image diffusion models. Imagen 4 (2025) improved fine detail and typography and is available in Gemini and Vertex AI.

Example:

Imagen 4 is used in Google tools for photorealistic marketing images with legible text.

Related Terms:

Diffusion Model
Nano Banana (Gemini Image)
Veo
View Source
Midjourney
Image Models

A closed image generator known for its strong default aesthetic. V7 (2025) improved coherence and added draft mode and personalization, and Midjourney has since added video generation.

Example:

The --sref parameter in Midjourney copies the style of a reference image into new generations.

Related Terms:

Prompt
Style Transfer
View Source
Ideogram
Image Models

A closed image model focused on accurate text and typography in images, well suited to posters, logos, and graphic design.

Example:

Ideogram can render a poster headline like 'SUMMER SALE 50% OFF' with correct spelling.

Related Terms:

Prompt
Recraft
View Source
Seedream
Image Models

ByteDance Seed's image generation and editing model family. Seedream 4.0 (2025) unified generation and editing with fast 2K to 4K output and multi-image references.

Example:

Seedream can create a consistent series of product shots from several reference photos.

Related Terms:

Image Editing Model
Seedance
View Source
Recraft
Image Models

A closed image model aimed at designers that can output vector (SVG) graphics, icons, and brand-consistent styles in addition to raster images.

Example:

Recraft can generate an editable SVG icon set in one consistent style.

Related Terms:

Ideogram
Style Transfer
View Source
Text-to-Video
Video Models

Generating a video clip directly from a text prompt. Modern models denoise a latent that spans width, height, and time with a video diffusion transformer.

Example:

'A drone shot flying over a foggy pine forest at sunrise' can produce a 5-second 720p clip.

Related Terms:

Image-to-Video
Diffusion Transformer (DiT)
Temporal Consistency
View Source
Image-to-Video
Video Models

Animating a still image into a video, optionally guided by a text prompt. It gives much more control over composition and characters than pure text-to-video.

Example:

Generate a character with FLUX, then animate it with Wan 2.2 I2V so the character's look stays exact.

Related Terms:

Text-to-Video
Wan (Wan-Video)
First-Last Frame
View Source
Temporal Consistency
Video Models

How stable objects, faces, colors, and lighting stay from frame to frame. Poor temporal consistency shows up as flicker, morphing, or identities that drift.

Example:

Video models attend across frames to keep a character's face identical as they turn their head.

Related Terms:

Text-to-Video
Frame Interpolation
View Source
First-Last Frame
Video Models

A video generation mode where you supply both the starting and ending frame, and the model generates the motion in between.

Example:

Supplying a closed flower and an open flower as first and last frames produces a blooming time-lapse.

Related Terms:

Image-to-Video
Frame Interpolation
Wan (Wan-Video)
View Source
Frame Interpolation
Video Models

Generating in-between frames to increase frame rate or smooth motion. Models such as RIFE and FILM are commonly used to turn 16 fps AI video into smooth 32 or 60 fps.

Example:

Running RIFE on a 16 fps Wan clip doubles its frame rate for smoother playback.

Related Terms:

Temporal Consistency
Video Upscaling
View Source
Wan (Wan-Video)
Video Models

Alibaba's open-weight video generation family. Wan 2.1 (February 2025) offered 1.3B and 14B text- and image-to-video models, and Wan 2.2 (July 2025, Apache 2.0) introduced a two-expert MoE for higher quality. Later 2.x releases added audio and were offered mainly through APIs.

Example:

Wan 2.2 I2V 14B with a Lightning LoRA is one of the most popular local video workflows in ComfyUI.

Related Terms:

Mixture of Experts (MoE)
Image-to-Video
VACE
View Source
VACE
Video Models

Video All-in-one Creation and Editing: a Wan-based framework that handles reference-to-video, video inpainting, pose or depth control, and motion transfer in one model.

Example:

VACE can make a character from a reference image follow the dance moves of a driving video.

Related Terms:

Wan (Wan-Video)
ControlNet
Inpainting
View Source
HunyuanVideo
Video Models

Tencent Hunyuan's open-weight video diffusion transformer. The original 13B model (December 2024) was followed by an image-to-video variant and the lighter HunyuanVideo 1.5 (2025), which runs on consumer GPUs.

Example:

HunyuanVideo is known for cinematic camera motion and realistic human movement.

Related Terms:

Diffusion Transformer (DiT)
Text-to-Video
HunyuanImage
View Source
LTX-Video
Video Models

Lightricks' fast open-weight video model, able to generate video close to real time on high-end GPUs. LTX-2 added synchronized audio and higher-resolution output in a single model.

Example:

LTX-Video is popular for rapid iteration because a short clip renders in seconds rather than minutes.

Related Terms:

Text-to-Video
Image-to-Video
Native Audio
View Source
Mochi 1
Video Models

Genmo's open-weight 10B video diffusion model released in 2024 under Apache 2.0, notable for strong motion quality and prompt adherence at the time.

Example:

Mochi 1 was one of the first open models to rival closed video generators on motion quality.

Related Terms:

Text-to-Video
Open Weights
View Source
CogVideoX
Video Models

Zhipu AI's open video generation models (2B and 5B) with a 3D VAE and expert transformer, and one of the early open text-to-video options.

Example:

CogVideoX-5B I2V animates product photos into short 6-second clips.

Related Terms:

Text-to-Video
VAE
View Source
Stable Video Diffusion (SVD)
Video Models

Stability AI's 2023 image-to-video model that turns a single image into a short 14 or 25 frame clip. It is largely superseded but was an important early open model.

Example:

SVD-XT was a common way to add subtle motion to still images in 2024 workflows.

Related Terms:

Image-to-Video
Stable Diffusion
View Source
AnimateDiff
Video Models

A plug-in motion module that adds temporal layers to existing Stable Diffusion 1.5 or SDXL checkpoints, so any fine-tuned image model can create animations.

Example:

AnimateDiff with an anime checkpoint and ControlNet was the go-to method for stylized animation before dedicated video models.

Related Terms:

Stable Diffusion
Temporal Consistency
ControlNet
View Source
Sora
Video Models

OpenAI's video generation model. Sora (2024) demonstrated long, coherent video with a diffusion transformer, and Sora 2 (September 2025) added synchronized dialogue and sound effects plus a social app.

Example:

Sora 2 can generate a clip where a character speaks a scripted line with matching lip movement.

Related Terms:

Diffusion Transformer (DiT)
Text-to-Video
Native Audio
View Source
Veo
Video Models

Google DeepMind's video generation family. Veo 3 (May 2025) was one of the first major models to generate native audio with the video, and Veo 3.1 added reference-image guidance and scene extension.

Example:

Veo 3.1 can take three reference images of a product and generate an ad where it stays consistent across shots.

Related Terms:

Native Audio
Text-to-Video
Imagen
View Source
Kling
Video Models

Kuaishou's closed video generation models, known for realistic human motion and physics. Newer versions such as Kling 2.x and 3.0 add higher resolution, longer clips, and audio.

Example:

Kling is a common choice for image-to-video of people dancing or doing sports.

Related Terms:

Image-to-Video
Text-to-Video
View Source
Seedance
Video Models

ByteDance Seed's video generation family. Seedance 1.0 (2025) supported multi-shot storytelling, and Seedance 2.0 (February 2026) accepts many reference images, clips, and audio files per generation for precise control.

Example:

Seedance can generate a multi-shot sequence where a product label stays readable while the camera orbits.

Related Terms:

Seedream
Image-to-Video
Native Audio
View Source
Runway Gen-4
Video Models

Runway's video models for filmmakers, focused on consistent characters, locations, and objects across shots using reference images. Runway also offers Aleph for editing existing video.

Example:

Gen-4 References keeps the same character's face across several separately generated shots.

Related Terms:

Image-to-Video
Temporal Consistency
View Source
Hailuo (MiniMax)
Video Models

MiniMax's video generation models, such as Hailuo 02, known for strong physics and complex motion at competitive prices.

Example:

Hailuo handles fast-action prompts like acrobatics with fewer limb artifacts than many models.

Related Terms:

Text-to-Video
Image-to-Video
View Source
Native Audio
Video Models

The ability of a video model to generate synchronized sound, including dialogue, effects, and ambience, in the same pass as the video rather than adding it afterwards.

Example:

Veo 3, Sora 2, and LTX-2 can generate footsteps and speech that line up with the on-screen action.

Related Terms:

Veo
Sora
Text-to-Speech (TTS)
View Source
Upscale Model
Upscalers

A neural network that increases image resolution by a fixed factor (commonly 2x or 4x) while reconstructing detail. In ComfyUI it is loaded with the Load Upscale Model node and applied in pixel space.

Example:

Running a 1024px render through a 4x model produces a sharp 4096px image in seconds.

Related Terms:

Upscaling
Real-ESRGAN
OpenModelDB
View Source
ESRGAN
Upscalers

Enhanced Super-Resolution GAN (2018), the architecture behind a huge number of community upscale models trained for specific content like photos, anime, or textures.

Example:

Most '4x-…' .pth files shared by the community use the ESRGAN architecture.

Related Terms:

Real-ESRGAN
Upscale Model
4x-UltraSharp
View Source
Real-ESRGAN
Upscalers

An ESRGAN variant trained on synthetically degraded images so it handles real-world blur, noise, and JPEG artifacts. It remains a fast, reliable default upscaler.

Example:

RealESRGAN_x4plus cleans up a noisy, compressed photo while enlarging it 4x.

Related Terms:

ESRGAN
Upscale Model
Upscaling
View Source
4x-UltraSharp
Upscalers

A popular community ESRGAN model tuned for crisp detail on AI-generated and photographic images, widely used in Stable Diffusion upscaling workflows.

Example:

4x-UltraSharp is a common choice inside Ultimate SD Upscale for adding fine texture.

Related Terms:

ESRGAN
Ultimate SD Upscale
Upscale Model
View Source
SwinIR
Upscalers

A transformer-based image restoration model (Swin Transformer) for super-resolution, denoising, and JPEG artifact removal, often more faithful than GAN upscalers.

Example:

SwinIR is used when you want clean enlargement without the over-sharpened look of some GAN models.

Related Terms:

Transformer
Upscale Model
View Source
SUPIR
Upscalers

A diffusion-based restoration and upscaling method built on SDXL that can reconstruct realistic detail in heavily degraded images, guided by a text prompt.

Example:

SUPIR can turn a blurry low-resolution old photo into a detailed high-resolution image, at the cost of speed and VRAM.

Related Terms:

Stable Diffusion XL (SDXL)
Diffusion Upscaling
Face Restoration
View Source
SeedVR2
Upscalers

ByteDance Seed's one-step diffusion transformer for video and image restoration (2025). It upscales video with strong temporal consistency and is available as ComfyUI nodes.

Example:

SeedVR2 upscales a 480p Wan clip to 1080p or higher without the flicker of frame-by-frame upscaling.

Related Terms:

Video Upscaling
Diffusion Upscaling
Temporal Consistency
View Source
Diffusion Upscaling
Upscalers

Upscaling that re-runs a diffusion model over the enlarged image at low denoising strength, inventing new plausible detail instead of only sharpening existing pixels.

Example:

Upscaling 2x with a GAN model and then running img2img at 0.3 denoise adds skin texture and fabric detail.

Related Terms:

Denoising Strength
Ultimate SD Upscale
Tiled Diffusion
View Source
Latent Upscale
Upscalers

Enlarging the latent representation directly before a second sampling pass (a 'hires fix'), instead of decoding to pixels first. It is fast but needs enough denoising to avoid artifacts.

Example:

A common SDXL workflow samples at 1024px, latent-upscales 1.5x, then resamples at 0.5 denoise.

Related Terms:

Latent Space
Denoising Strength
Upscaling
View Source
Video Upscaling
Upscalers

Increasing the resolution of video while keeping detail consistent across frames. Dedicated video upscalers avoid the flicker that happens when each frame is upscaled independently.

Example:

SeedVR2 and Topaz Video are commonly used to bring 720p AI video up to 4K.

Related Terms:

SeedVR2
Temporal Consistency
Frame Interpolation
View Source
Topaz Labs
Upscalers

Commercial desktop and cloud software (Topaz Photo and Topaz Video) for AI upscaling, denoising, sharpening, and frame interpolation, widely used for finishing AI-generated media.

Example:

Topaz Video can upscale and frame-interpolate a 16 fps AI clip to smooth 4K at 60 fps.

Related Terms:

Video Upscaling
Frame Interpolation
View Source
OpenModelDB
Upscalers

A community database of upscaling and restoration models, searchable by architecture, scale, and content type, with sample comparisons.

Example:

OpenModelDB is where you find a 2x anime upscaler or a JPEG-artifact removal model for ComfyUI.

Related Terms:

Upscale Model
ESRGAN
View Source
GFPGAN
Upscalers

A face restoration model that uses a pretrained face GAN as a prior to repair blurry, low-quality faces in images.

Example:

GFPGAN fixes distorted faces in small background figures after upscaling.

Related Terms:

Face Restoration
CodeFormer
Face Detailer
View Source
CodeFormer
Upscalers

A transformer-based face restoration model with an adjustable fidelity weight that lets you trade identity preservation against restoration strength.

Example:

Setting CodeFormer's fidelity to 0.7 restores a face while keeping it recognizable.

Related Terms:

Face Restoration
GFPGAN
View Source
Whisper
Audio Models

OpenAI's open-weight speech recognition model that transcribes and translates speech in about 100 languages.

Example:

Whisper can produce subtitles for a generated talking-head video.

Related Terms:

Text-to-Speech (TTS)
Open Weights
View Source
Text-to-Speech (TTS)
Audio Models

Models that turn text into natural speech, often with voice cloning from a short sample. Examples include ElevenLabs as well as open models such as Kokoro, F5-TTS, and Chatterbox.

Example:

A TTS model voices a script, and a lip-sync model then animates a generated character to match.

Related Terms:

Native Audio
Whisper
View Source
MusicGen
Audio Models

Meta's open music generation model from the AudioCraft library that creates music from text descriptions, optionally conditioned on a melody.

Example:

'Lo-fi hip hop beat with warm piano' generates a short background track for a video.

Related Terms:

Stable Audio
ACE-Step
View Source
Stable Audio
Audio Models

Stability AI's latent diffusion models for music and sound effects. Stable Audio Open is an open-weight version suited to samples, loops, and sound design.

Example:

Stable Audio Open can generate a 'rain on a tin roof' ambience loop for a scene.

Related Terms:

Latent Diffusion
MusicGen
View Source
ACE-Step
Audio Models

An open-source music generation foundation model that generates full songs with vocals and lyrics quickly, and runs locally including in ComfyUI.

Example:

ACE-Step can turn lyrics plus a genre tag into a complete song with vocals in under a minute.

Related Terms:

MusicGen
Stable Audio
View Source
Depth Anything
Vision Models

A monocular depth estimation model family trained on large-scale data. It produces detailed depth maps from a single image and is widely used for depth ControlNets.

Example:

Depth Anything V2 creates the depth map that guides a ControlNet to keep a room's layout while restyling it.

Related Terms:

Depth Map
ControlNet
View Source
SAM 2
Vision Models

Meta's Segment Anything Model 2, which extends promptable segmentation from images to video with memory for tracking objects across frames.

Example:

Click a person once and SAM 2 produces a mask for them through the whole clip, ready for video inpainting.

Related Terms:

SAM (Segment Anything Model)
Segmentation
Inpainting
View Source
Florence-2
Vision Models

Microsoft's compact vision foundation model that performs captioning, OCR, object detection, and grounding from task prompts.

Example:

Florence-2 is a popular ComfyUI node for auto-captioning images or detecting a face region to mask.

Related Terms:

Vision-Language Model (VLM)
Grounding DINO
Bounding Box
View Source
Grounding DINO
Vision Models

An open-set object detector that finds objects described by free text, not just a fixed list of classes. It is often paired with SAM to create masks from a text prompt.

Example:

Grounding DINO plus SAM turns the text 'the red car' into a precise mask for inpainting.

Related Terms:

SAM (Segment Anything Model)
Bounding Box
Segmentation
View Source