MiniMax H3 is the rare AI video release where the interesting part is not just the demo reel. The open weights arrived with native ComfyUI support for text-to-video, image-to-video and reference-to-video, including stereo sound generated with the picture.
This guide shows you how to install MiniMax H3 in ComfyUI on Windows with the official native templates. We will start with the stable four-file T2V/I2V path, explain the separate Ref2VA checkpoint, and only then touch optional speed tricks. That order matters: a clean five-second clip is far more useful than an elaborate workflow full of red nodes.
Verified scope: The filenames, folders, modes, limits and version notes below were checked against first-party documentation on 10 August 2026. I have not run H3 on the site’s test machine, so any hands-on speed or minimum-VRAM claim would be fiction. Community reports show that heavily optimised low-VRAM runs are possible, but MiniMax and ComfyUI do not publish one universal consumer-GPU minimum.

The 30-second version
- Use ComfyUI v0.31.0 or newer. H3 first appeared in v0.30.0, while v0.31.0 added important VAE, audio, sampling and offload fixes.
- Choose the native MiniMax H3 T2V or I2V template for your first run.
- T2V and I2V use the
fl2vadiffusion checkpoint. R2V uses a differentref2vacheckpoint. - Allow about 42.5 GB for the T2V/I2V model stack, or about 63.5 GB if you also install Ref2VA—plus room for caches and outputs.
- Local H3-Base works on a 768-pixel-short-edge canvas. MiniMax’s complete 2K workflow also uses hosted modules that are not part of the open release.
- Read the community licence before downloading. It has meaningful territorial and use restrictions.
What MiniMax H3 can do locally
MiniMax describes H3 as an omni-modal generation system. It accepts text and, depending on the checkpoint, image, video and audio references. It produces 24fps video with 32kHz stereo audio. The published duration range is 4 to 15 seconds.
| ComfyUI template | Use it for | Diffusion checkpoint |
|---|---|---|
| T2V | A prompt with no starting image | fl2va |
| I2V | A first frame, last frame, or both | fl2va |
| R2V | Character, style, motion, camera or voice references | ref2va |
The open H3-Base checkpoint generates on a native canvas with a 768-pixel short edge, capped at 768 × 1344 and rounded to a multiple of 32. MiniMax’s headline “up to 2K” result comes from H3-Regenerate-2K; its full workflow also uses the hosted H3-Context-IR service. Neither module is included in the initial open-weight release. In plain English: local 768p is real, but fully local native 2K is not what this guide is promising.
Before you download anything
Licence warning: H3 uses the MiniMax H3 Community License Agreement, not Apache 2.0 or an OSI-approved open-source licence. As published on 2 August 2026, its applicable territory excludes the United States, European Union, United Kingdom and Republic of Korea unless separate permission is obtained. It also contains commercial, distribution, disclosure and acceptable-use conditions. Check the current agreement for your location and use case; this paragraph is not legal advice.
- A 64-bit Windows 10 or Windows 11 computer.
- Comfy Desktop or ComfyUI Portable. The portable build is the easiest folder layout to follow in this guide.
- A supported dedicated GPU. NVIDIA is the most established Windows route, although current ComfyUI Portable packages also exist for AMD and Intel hardware.
- A fast SSD with at least 50 GB free for the T2V/I2V setup. Keep more space available for Ref2VA, temporary downloads and generated video.
- A current graphics driver and enough system memory to support offloading.
Do not buy a graphics card because one post says “H3 needs exactly X GB.” Duration, resolution, checkpoint format, offloading, system RAM and optional custom nodes can all change the result. Start with the official pruned and quantised components below, then measure your own machine.
Step 1: Install or update ComfyUI
If you would rather let Claude Code or Cursor validate and run a known workflow, the Comfy MCP setup guide connects an AI client to this local ComfyUI installation without exposing it publicly.
H3 native support began in ComfyUI v0.30.0. Use v0.31.0 or later because that release added H3 fixes for raw VAE parameters, the int8_convrot VAE path, noise-mask sampling, audio samplers, full audio-VAE offload and EasyCache audio corruption.
Comfy Desktop
Install the official Windows app from the Comfy Desktop guide. If it is already installed, use Desktop Update Ready or open Desktop Settings → Updates → Check for updates, then restart.
ComfyUI Portable
Close ComfyUI, open the portable installation’s update folder and run update_comfyui_stable.bat. The official documentation reserves update_comfyui_and_python_dependencies.bat for runtime problems; it is not the first button to press because dependency changes can break custom nodes.
Step 2: Choose the right H3 workflow
Open ComfyUI, go to Template Library → Video, and search for MiniMax H3. Current ComfyUI builds ship three native templates:
- MiniMax H3 T2V: the simplest first test and the best place to prove the installation works.
- MiniMax H3 I2V: accepts an optional first frame, last frame, or both.
- MiniMax H3 R2V: accepts reference images, videos and audio, but needs the separate Ref2VA diffusion checkpoint.
Choose T2V for now. When the template opens, ComfyUI may offer to fetch missing models. You can use that prompt or download the files manually from the official Comfy-Org/MiniMax-H3 repository.
Step 3: Download the exact model files
For T2V or I2V, you need four files totalling about 42.5 GB: a 20.97 GB diffusion model, 15.69 GB text encoder, 5.21 GB video VAE and 0.61 GB audio VAE. R2V shares the text encoder and both VAEs but replaces the FL2VA model with another roughly 20.97 GB checkpoint. Installing both diffusion models takes the set to about 63.5 GB before caches and outputs. Keeping that distinction clear saves a very large wrong download.
| Official file | Folder | Used by |
|---|---|---|
| minimax_h3_fl2va_pruned_int8_convrot.safetensors | ComfyUI/models/diffusion_models/ | T2V and I2V |
| minimax_h3_ref2va_pruned_int8_convrot.safetensors | ComfyUI/models/diffusion_models/ | R2V only; download later if needed |
| qwen3vl_32b_minimax_h3_nvfp4_awq.safetensors | ComfyUI/models/text_encoders/ | All three |
| minimax_h3_video_vae_fp16.safetensors | ComfyUI/models/vae/ | All three |
| minimax_h3_audio_vae_fp32.safetensors | ComfyUI/models/vae/ | All three |
For the first T2V run, download every row except Ref2VA. Restart ComfyUI after moving the files. In File Explorer, enable View → Show → File name extensions so a failed browser download does not masquerade as model.safetensors.safetensors.
Step 4: Run a boring first test
- Open the native MiniMax H3 T2V template.
- Confirm that the FL2VA checkpoint, Qwen text encoder, video VAE and audio VAE appear in their loader nodes.
- Use the template’s fast preview size for the first queue. For full native quality later, about 1.0 megapixel at 16:9 produces roughly 1344 × 768.
- Keep the resolution multiple at
32. - Choose a short duration. H3 supports 4–15 seconds at 24fps; ComfyUI snaps duration to the model’s valid 17-frame-block grid.
- Leave the supplied sampler and model settings alone until one output succeeds.
- Use a fixed seed while troubleshooting.
Start boring. Make one short T2V clip before adding Sage Attention, Turbo LoRAs or half the internet’s custom nodes. If the clean template fails, every optional component makes the error harder to isolate.
Step 5: Prompt the picture and the sound
H3 generates the audio and video together, so a silent visual prompt leaves part of the model under-directed. Describe the scene, subject, movement, camera, lighting and sound in one block. For multi-shot clips, add rough timestamps.
A close shot of a small red robot repairing a radio on a rain-soaked workbench at night. The camera slowly pushes forward while the robot turns a brass dial. Warm workshop light, realistic reflections and shallow depth of field. Stereo audio: rain tapping on the metal roof, a soft electrical hum, tiny tool clicks from the centre, and distant thunder moving from left to right. No speech, captions or logos.
Queue the prompt and watch the ComfyUI console during the first load. Large components may spend a while moving from SSD to system memory and GPU memory. When the queue finishes, play the MP4 with sound enabled before changing anything.
Step 6: Move to I2V or R2V
Image-to-video and first/last frames
Open the I2V template and connect a first frame, last frame, or both to the MiniMaxH3ImageToVideo node. The FL2VA checkpoint remains correct. Describe the movement between the frames rather than repeating a static description of the image.
Reference-to-video
Open the R2V template and select minimax_h3_ref2va_pruned_int8_convrot.safetensors. The official model limits are up to nine images, three video clips and three audio clips, with no more than 12 files in total. A video or audio clip must be 2–15 seconds; audio cannot be the only reference type.
Refer to inputs in connection order—such as <Picture 1>, <Video 1> and <Audio 1>—and give each one a job: identity, art style, motion, camera or voice. “Use Picture 1 for the character and Video 1 for the camera movement” is much clearer than dumping references into the graph and hoping the model reads your mind.
Optional: Sage Attention after the native workflow works
ComfyUI’s official H3 page says Sage Attention can roughly double generation speed with minimal quality loss. That is an upstream estimate, not a benchmark from this site. It is also optional and adds a Python wheel plus KJNodes, so it belongs after the stable baseline—not before it.
If you try it, use the SageAttention wheel matching your installed PyTorch and CUDA versions. The official route adds KJNodes’ Patch Sage Attention KJ between UNETLoader and BasicGuider, or launches ComfyUI with --use-sage-attention. Audit third-party nodes before installing them. If the workflow fails, remove the patch and reproduce the problem with the untouched native template first.
How to confirm H3 is really local
After the model downloads, H3-Base generation runs through your local ComfyUI instance. You can see the local process using GPU, system memory and disk in Task Manager, and the result appears in the local output folder. That does not make MiniMax’s entire production pipeline local: H3-Context-IR and H3-Regenerate-2K are hosted services in the current release.
MiniMax H3 troubleshooting
| Problem | What to check first |
|---|---|
| Missing or red H3 nodes | Update to ComfyUI v0.31.0+ and restart. MiniMaxH3ReferenceToVideo is a native core node, not something to hunt for in Manager. |
| Model absent from a dropdown | Verify the exact folder, filename and extension, then refresh or restart ComfyUI. |
| R2V does not load | Confirm the workflow uses ref2va, not the fl2va checkpoint used by T2V/I2V. |
| CUDA out of memory | Use the fast preview size, shorten duration, close GPU-heavy apps, restart ComfyUI and retest the native workflow. |
| No, damaged or strange audio | Check the audio VAE and update to v0.31.0+, which includes H3 audio-sampler, offload and EasyCache fixes. |
| Sage Attention error | Remove the patch first. Then verify that its wheel matches the actual PyTorch and CUDA versions in ComfyUI’s Python environment. |
| Model reloads from disk every queue | Leave Dynamic VRAM enabled, reduce the job, close competing apps and check the console. The cache must evict components when combined VRAM and RAM are insufficient. |
The first load looks frozen
Watch the console rather than the browser spinner. If disk, CPU or GPU activity continues, the model may still be loading. If the console stops on an error, copy the first meaningful error message—not merely the final stack-trace line—before searching for a fix.
ComfyUI broke after an update
Disable custom nodes and test the native template in a clean portable ComfyUI copy. If the official workflow works there, H3 is not the problem; a custom node or dependency conflict is. Add extras back one at a time.
Frequently asked questions
Is MiniMax H3 open source?
It is open weight, but it is distributed under the custom MiniMax H3 Community License Agreement rather than a standard OSI-approved licence. The current agreement has territorial, commercial, distribution and use restrictions, so “downloadable” should not be confused with “unrestricted.”
Can MiniMax H3 run on 12GB or 6GB VRAM?
Community workflows report runs on 12GB and even 6GB cards using GGUF, aggressive offloading, reduced resolution and extra nodes. Those are useful experiments, not official minimums and not the native setup documented here. Expect compromises in speed, resolution, complexity or all three.
Can H3 generate audio locally?
Yes. H3-Base jointly predicts video and 32kHz stereo audio, and the native ComfyUI templates use a separate H3 audio VAE to decode it.
Is MiniMax H3 2K fully local?
No, not in the initial release. The open H3-Base path produces 768p-class output. The official 2K validation workflow also calls MiniMax’s hosted Context-IR and Regenerate-2K services.
Should I disable Dynamic VRAM?
Not as a first fix. Dynamic VRAM lets ComfyUI offload components when memory is tight. Disabling it can replace a slow reload with an out-of-memory failure. Update, reduce the job and reproduce the issue with the clean native workflow before changing memory flags.
Related local-AI guides
If you need a first frame before animating it, try the Leonardo AI image-generation guide. For a smaller audio-only local project, see the MusicGen guide or the Tortoise TTS setup guide. You can also browse more practical AI tutorials.
Official sources
- MiniMax H3 model repository and system documentation
- MiniMax H3 model card and weights
- MiniMax H3 Community License Agreement
- Official MiniMax H3 ComfyUI workflows
- ComfyUI v0.31.0 release notes and H3 fixes
- ComfyUI RAM-aware model-cache change
- Official Comfy Desktop Windows installation guide
- Official ComfyUI Portable Windows guide
Final thoughts
H3’s feature list is enormous, but a dependable installation is pleasantly unglamorous: current ComfyUI, the right diffusion checkpoint, three shared components and one conservative test. Prove that baseline first. Once it works, I2V, reference inputs and performance tuning become experiments instead of mysteries.