Run MiniMax H3 locally for 768p video with synced stereo audio
MiniMax H3 is an open-weight model that generates video and stereo audio together. Here's how to serve it with SGLang and where the hosted APIs fit in.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
MiniMax H3 is an open-weight video model from MiniMax. It generates video with native 32 kHz stereo audio, and it takes text, images, video and audio as input. H3 predicts video and audio latents together in one transformer, so sound is generated alongside the picture rather than added afterward. The open release covers the 768p generator (H3-Base). The prompt-preprocessing stage (H3-Context-IR) is not part of the open release, and MiniMax offers it as a hosted API instead. The 2K stage (H3-Regenerate-2K) regenerates the 768p clip in context rather than using a conventional super-resolution module. It is also API-only for now, and MiniMax says it will open-source it once it is ready.
Key specs
| Item | Value |
|---|---|
| Core model | H3-Omni-Transformer, 33B-parameter dense, single-stream |
| AdaLN parameters | About 13B. They can be precomputed and cached, so inference-only deployments don’t need to load them |
| Text/vision encoder | Full Qwen3-VL-32B weights, hidden states taken from layer 50 |
| Weights precision | BF16, CFG-distilled |
| Output duration | 4–15 seconds |
| Output resolution | Shorter side 768 px by default. 2K needs H3-Regenerate-2K (hosted only for now) |
| Frame rate | 24 FPS |
| Audio | 32 kHz stereo |
| Aspect ratios | 21:9, 16:9, 4:3, 1:1, 3:4 and 9:16, among others |
| Dialogue languages | Stable support for 11: Arabic, Chinese, English, French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish |
The model ships as two checkpoints for different tasks:
- FL2VA: text-to-audio-video (
t2va) and first/last-frame-to-audio-video (fl2va). Accepts zero, one or two images. - Ref2VA: reference-to-audio-video (
ref2va). Accepts up to 9 images, up to 3 video clips and up to 3 audio clips. Each clip must be 2–15 seconds, total video and total audio duration must each be 15 seconds or less, and a request can include at most 12 files.
Install
The card doesn’t include install commands for the serving frameworks. It points to four supported options:
- SGLang: cookbook
- vLLM: recipes
- diffusers: pipeline docs
- ComfyUI: tutorial
Install your framework using its own docs. Then download the weights with the hf CLI. The repo stores the original checkpoint and the diffusers format side by side, so download only the parts you need:
# Original checkpoint, both task families (SGLang, vLLM):hf download MiniMaxAI/MiniMax-H3 --include "model_index.json" "FL2VA/*" "Ref2VA/*" --local-dir MiniMax-H3
# Or a single task family:hf download MiniMaxAI/MiniMax-H3 --include "model_index.json" "FL2VA/*" --local-dir MiniMax-H3If you use diffusers, you can skip the manual download. ModularPipeline.from_pretrained("MiniMaxAI/MiniMax-H3") fetches only the components it needs. The card doesn’t show the rest of the diffusers call, so use the diffusers docs linked above for generation arguments.
Each task folder contains model_index.json, processor/, tokenizer/, text_encoder/, transformer/, visual_vae/ and audio_vae/.
Run it
The card uses SGLang as its deployment example. To start a server for the FL2VA checkpoint on 4 GPUs:
sglang serve \ --model-path MiniMaxAI/MiniMax-H3 \ --num-gpus 4 \ --ulysses-degree 4 \ --performance-mode speed \ --host 0.0.0.0 \ --port 30010 \ --model-variant fl2vaThis command points --model-path at the Hub ID, not at a local download directory. The card doesn’t show how to serve from a local path. If you need that, check the SGLang cookbook.
The card links the request payloads instead of including them inline. Use the reproducible 768p scripts in the repo to send your first request and compare your output with MiniMax’s reference clips:
- T2VA:
scripts/readme/reproducible-768p-t2va-request.sh, with expected output inassets/t2va.mp4 - FL2VA:
scripts/readme/reproducible-768p-fl2va-request.sh, with expected output inassets/fl2va.mp4 - Ref2VA:
scripts/readme/reproducible-768p-ref2va-request.sh, with expected output inassets/ref2va.mp4
Text or keyframes to a clip with sound
The FL2VA server covers three jobs:
- Text only: a scene with dialogue, ambience and music.
- One image: the image becomes the first or last frame of the clip.
- Two images: the model generates the motion between them.
MiniMax says H3-Context-IR is critical to H3’s output quality, so the prompt you send to H3-Base matters. The card’s T2VA example shows the prompt that H3-Context-IR produced and that was then passed to H3-Base. It is split into three labeled parts: integrated_multimodal_description (shots with timestamps, such as a cut at 00:04.500), overall_soundscape and non_diegetic_music. This is not a universal format. The Ref2VA example’s Context-IR output uses a different structure (subject_definitions, summary, retention_analysis, detailed_description and others). If you write prompts yourself, follow the repo’s prompt guides for each task (see Gotchas) instead of copying one layout everywhere.
To serve FL2VA, use the serve-fl2va.sh command above.
Editing and dubbing with reference media
Ref2VA is the more unusual checkpoint. In the card’s example, the input is a source video, its music track and a separate voice sample. The model keeps the framing, lighting and subject of the source video, reuses the original music, and makes the subject speak new lines in the voice from the sample. In the Context-IR output passed to H3-Base, dialogue is wrapped in <d> tags with a language marker, for example <d>[English] Follow the wind, live free.</d>.
The card’s example shows one kind of use: adding new dialogue to existing footage in a referenced voice, without reshooting. The example uses English dialogue with an English voice reference.
Serve Ref2VA on its own port:
sglang serve \ --model-path MiniMaxAI/MiniMax-H3 \ --num-gpus 4 \ --ulysses-degree 4 \ --performance-mode speed \ --host 0.0.0.0 \ --port 30011 \ --model-variant ref2vaAs written, each server command uses --num-gpus 4. Running FL2VA and Ref2VA at the same time therefore needs 8 GPUs or separate nodes. The card doesn’t show the two servers sharing a machine.
Stay within the input limits listed in Key specs: at most 9 images, at most 3 video and 3 audio clips of 2–15 seconds each, at most 15 seconds of video and 15 seconds of audio in total, and at most 12 files per request.
Full 2K output: local generation plus hosted APIs
To match the quality of MiniMax’s own API, the card describes a three-step pipeline:
- Send your raw inputs to the hosted H3-Context-IR API (docs). It returns a structured prompt.
- Generate the 768p clip on your local H3-Base server using that prompt.
- Send the 768p result together with the original context to the hosted H3-Regenerate-2K API (docs), with the video passed as
base_video. The card says reusing the original context is what lets the model recover details such as small text, which conventional super-resolution would have to guess.
The links above are API documentation pages. Get the exact request URLs and payloads from those docs or from the repo scripts, not from the page names. CN docs are on platform.minimaxi.com.
The card’s example scripts send local files as Base64 data URLs. For production, it recommends uploading the video to a public URL and passing that URL as base_video. Set up the environment first:
# URL of your SGLang deploymentSGLANG_DEPLOYMENT_URL="<sglang-deployment-url>"
# MiniMax API endpoint (choose one)# CNMINIMAX_API_BASE="https://api.minimaxi.com"# Global# MINIMAX_API_BASE="https://api.minimax.io"
# API token obtained from the MiniMax platformTOKEN="<token>"Each stage has a script in the repo, for example scripts/readme/full-2k-t2va-h3-context-ir.sh, full-2k-t2va-h3-base.sh and full-2k-t2va-h3-regenerate-2k.sh. I2VA has a matching set of all three. Ref2VA has H3-Context-IR and H3-Base scripts but no H3-Regenerate-2K script, so the repo doesn’t show the full three-step pipeline for Ref2VA. Each case also includes reference outputs from the hosted API at 768p and 2K.
Gotchas
- The recommended prompt stage is not open. MiniMax says H3-Context-IR is “critical to the quality of the final output” and strongly recommends using it. It is not included in the open-source release. MiniMax offers it as an API and publishes guides for building your own. If you don’t use the API, follow the repo’s prompt guides (
docs/VIDEO_PROMPT_WRITING_GUIDE_base_en.md,docs/VIDEO_PROMPT_WRITING_GUIDE_ref_en.md) and the skills, or expect weaker results than the demos. - 2K is API-only for now. Fully local deployments top out at 768p until MiniMax releases H3-Regenerate-2K, which it says it will do once the module is ready.
- Sparse attention is not released yet. The open release runs full attention only, so long, high-token sequences cost more compute than they will once the sparse implementation ships.
- Use the H3 tokenizer. It adds special tokens such as
<d>, so a stock Qwen3-VL tokenizer won’t work. - Two checkpoints, two servers. FL2VA and Ref2VA are separate weights, so serving both means running two processes, one with
--model-variant fl2vaand one with--model-variant ref2va. With the card’s commands, that is 4 GPUs each. - The model is large. The transformer has 33B parameters and the encoder is the full Qwen3-VL-32B, both in BF16. About 13B of the transformer’s parameters are AdaLN branches that, according to the card, don’t need to be loaded for inference-only use, so the memory needed at load time is lower than the headline totals suggest. The card doesn’t give per-GPU memory, but its reference deployment uses 4 GPUs with Ulysses parallelism. Plan for a multi-GPU node.
- The hosted stages moderate content. Text, images, videos and enhanced prompts sent to the APIs pass through automated moderation, and MiniMax notes it can produce false positives.
- License. The weights use the custom MiniMax H3 Community License, not an OSI license. The card also links an application form for users in the USA, EU, UK and South Korea. Read the license and its Q&A before shipping a product.
When to pick it
Pick H3 if you need open weights for video with native stereo audio, such as dialogue in one of its 11 stable languages, or if you need reference-driven generation and editing with mixed image, video and audio inputs. The full weights, including the AdaLN branches, are released to support further development such as fine-tuning. The card doesn’t include training code or a fine-tuning recipe, though.
Skip it if:
- You have one consumer GPU. This is our inference: the card gives no minimum hardware, but its example deployment uses 4 GPUs.
- You need 2K output without calling MiniMax’s API.
- Your license or compliance rules exclude custom community licenses.
- You only need silent clips.
If you won’t build your own prompt-preprocessing step and you’re fine with a hosted service, the MiniMax API already runs the full pipeline and is the simpler option.
Related
Self-Host MiMo-V2.6-Pro-RL, Xiaomi's 1T MoE Agent Model
Xiaomi's 1.02T-parameter MoE model has 1M context and omnimodal input, and was trained for agents. Here is how to serve it with vLLM or SGLang.
XiaomiMiMo/MiMo-V2.6-Pro-RL
DeepSeek-V4.1-Flash: a 1M-context multimodal MoE built for agent work
DeepSeek's 552B MoE activates 8B params on prefill and stores 890 bytes of KV cache per token. What the card says and how to start running it.
deepseek-ai/DeepSeek-V4.1-Flash
Run Gemma 4 31B for Image, Video and Reasoning Tasks
Google DeepMind's largest dense Gemma 4 model handles images, video frames and 256K-token context. Here's how to run it with Transformers.
google/gemma-4-31B-it