Feedback
AI Ad Video Example
Loading...
comfyui minimax h3
Produce 2K/24fps clips with perfectly aligned audio inside ComfyUI. The comfyui minimax h3 workflow accepts text, image, or reference inputs with open weights and no watermark.
All Tools
Discover our comprehensive AI-powered animation toolkit
MiniMax H3
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.

Veo3.1
Create Stunning Videos with Veo3.1
FLUX 3 Video Generator

Kling 3.0
Next-Gen AI Video Generator
Grok Video Generator
Create Videos from Text or Images with AI
MiniMax H3 video generator
MiniMax H3 AI Video Generator

Seedance 2.0
The Future of AI Video Is Here.
Key Benefits of the comfyui minimax h3 Workflow
The comfyui minimax h3 workflow connects MiniMax's open-weights omni-modal model with ComfyUI's node graph. It processes text, stills, clips, and sound together, generating video with naturally synced stereo audio — voice, effects, and music — all in one pass. Output can reach 2K resolution at 24fps for up to 15 seconds, while keeping every diffusion parameter adjustable.
- Synced Stereo SoundVoice, sound effects, and music are rendered together with the footage into one MP4, perfectly timed by the comfyui minimax h3 pipeline.
- Full Local ControlExecute the comfyui minimax h3 model on your own hardware, adjusting resolution, length, and each diffusion setting — no rate limits or API barriers.
- Mixed Reference TypesMix prompts, images, clips, and audio tracks in a single run — the comfyui minimax h3 nodes can pin down a character, style, movement, camera angle, or vocal tone.
Getting the Most from the comfyui minimax h3 Workflow
Follow this simple guide to produce open-weight clips with synchronized audio — the comfyui minimax h3 workflow gets you there in three steps.
Key Capabilities of the comfyui minimax h3 Workflow
From three prebuilt ComfyUI templates to open-weight multimodal generation, the comfyui minimax h3 workflow bundles native stereo sound, reference-driven control, and Sage Attention acceleration — everything needed for local video creation.
Three Built-In Template Variants
The comfyui minimax h3 template pack includes ready-made text-to-video, image-to-video, and reference-to-video setups, each handling a distinct generation mode immediately.
Unified Multimodal Understanding
The comfyui minimax h3 model processes text, stills, motion, and sound in a single context, allowing you to blend every input type within one render.
Guided Output from References
Use up to 9 images, 3 video clips, and 3 audio files through the comfyui minimax h3 R2V node to fix a character's look, visual style, action, camera movement, or voice.
Sharp Text and Brand Mark Output
The comfyui minimax h3 model renders written text and brand logos crisply, while natural-language instructions clearly describe how references relate to each other.
Boost Speed with Sage Attention
By inserting the Patch Sage Attention KJ node into the comfyui minimax h3 workflow, you can nearly halve render time without much quality trade-off.
Flexible Resolution and Length Settings
The comfyui minimax h3 Resolution Selector derives width and height from aspect ratio and megapixels, snapping to the 32-pixel grid and 17-frame-per-block timing at 24fps.
comfyui minimax h3: Common Questions
Answers to the most common questions about using the comfyui minimax h3 model within ComfyUI.
What does the comfyui minimax h3 workflow involve?
The comfyui minimax h3 workflow is ComfyUI's built-in support for MiniMax H3, an open-weight omni-modal model from MiniMax. It produces video with naturally synced stereo audio, taking text, images, clips, or sound references in one forward pass.
Which resolution and frame rate does it handle?
The comfyui minimax h3 workflow can deliver up to 2K at 24fps for around 15 seconds. The native canvas has a 768-pixel short edge, with a maximum of 768x1344 pixels, all rounded to multiples of 32.
What different ways can I generate video with this workflow?
The comfyui minimax h3 template library provides three example flows: text-to-video (T2V), image-to-video (I2V) with optional control over the first and last frames, and reference-to-video (R2V) that pins down character, style, motion, camera, or voice.
Is audio included in the generated video?
Yes — the comfyui minimax h3 model creates true stereo audio, covering speech, effects, and music, all modeled alongside the video and written into one MP4 with perfect sync.
What do I need to do to start using it?
Ensure ComfyUI is at least version 0.30.0, navigate to Template Library > Video, select a comfyui minimax h3 workflow, and use the pop-up to fetch models from the Hugging Face Comfy-Org/MiniMax-H3 repository.
Are there ways to make the workflow run faster?
Absolutely — after installing SageAttention and KJNodes, simply place a Patch Sage Attention KJ node between the UNETLoader and BasicGuider in the comfyui minimax h3 workflow, and you can roughly double the render speed.
Ready to Use the comfyui minimax h3 Workflow?
Run MiniMax H3 on your own machine through ComfyUI, with true stereo audio, open weights, and every parameter adjustable. Pick between text, image, or reference input paths and start right away.
