MiniMax open-sources H3, its next-generation general-purpose video model
MiniMax released the weights for H3, its general-purpose omni-modal generation model, on August 3, 2026, three days after first unveiling the model. "Today, we are officially open-sourcing MiniMax H3, our next-generation general-purpose video model," the company said, making the weights available on Hugging Face and ModelScope under a new community license.
What's new
H3 first launched on July 31 as a model that "understands unified context across text, images, video, and audio, generating video with native stereo sound, up to 15 seconds at 2K resolution." MiniMax positioned it around instruction-following, accurate text and brand rendering inside generated video, and V2V (video-to-video) motion transfer — capabilities aimed squarely at commercial content production: advertising, branding, e-commerce, product design, UI/UX, and gaming.
On pricing, MiniMax said H3's per-second API cost "at 2K resolution... is less than a third of mainstream models," and at 768p it runs "less than half the price of mainstream models' 720p" — an aggressive undercut aimed at rivals like ByteDance's Seedance, which shipped an update of its own within hours of H3's original release.
The August 3 open-source release makes two task-specific checkpoints, FL2VA and Ref2VA, available as downloadable Hugging Face repositories, with support across common inference frameworks including SGLang, vLLM, diffusers, and ComfyUI. Not everything shipped open: MiniMax noted that "due to the complexity of the system," its H3-Regenerate-2K module "is not yet open-sourced" and remains accessible only through the company's hosted API. The released weights fall under MiniMax's own "H3 Community License Agreement," with the company publishing a separate Q&A document to address licensing questions.
Context
H3 is the latest entry in a fast-moving contest among Chinese labs to ship cheap, capable, openly-licensed video generation models. ByteDance and MiniMax have been trading upgrades to their respective video models within days of each other, and MiniMax has increasingly leaned on open-sourcing model weights — as it previously did with its M2 and M3 text/agentic models — as a distribution strategy that undercuts closed rivals like Seedance on both price and access.
Why it matters
Open-weight releases of frontier-adjacent video models remain rare — most competitive video generation, from OpenAI's Sora to Google's Veo, stays behind an API. By open-sourcing H3's core checkpoints just three days after launch, MiniMax is betting that developer adoption and ecosystem lock-in (via Hugging Face, ModelScope, and integration into popular inference stacks) matters more than keeping the model fully proprietary. Combined with aggressive per-second pricing on the hosted API for the pieces that remain closed, it's a two-pronged play: capture open-source developers who can self-host, while still competing on cost for anyone who uses the managed service — intensifying the price and openness pressure already building in the video-generation market.
Corroborating sources
- Bloomberg
https://www.bloomberg.com/news/articles/2026-07-31/china-s-minimax-and-bytedance-release-dueling-ai-video-models
- Huggingface.co
https://huggingface.co/MiniMaxAI/MiniMax-H3
“MiniMax H3 is a general-purpose, omni-modal generative system”