AI/TLDR

MiniMax · 2026-07-31 · major

MiniMax H3 — open-weights video model does 2K, 15s, and native stereo sound

MiniMax launched H3, a full-modal video model that generates 15-second 2K clips with synchronized stereo audio. It edits video, transfers motion, and MiniMax says weights will drop on Hugging Face within days.

MiniMax H3 official launch cover image

A single open-weights video model does 2K clips, stereo audio, editing, and motion transfer up to 15 seconds long.

Quick facts

MakerMiniMax (Shanghai)
ModelH3 (Hailuo 3.0), full-modal video
OutputUp to 15s clips, native 2K (2560×1440), 24 FPS
AudioNative synchronised stereo sound, voice + SFX
Inputs per requestUp to 9 images + 3 video clips + 3 audio tracks
AvailabilityMiniMax API and hosted providers; open weights promised within days
Price vs peersUnder 1/3 of mainstream 2K video pricing

What is it?

MiniMax H3 (Hailuo 3.0) is a full-modal video generation model from the Shanghai lab MiniMax. It generates up to 15-second video at native 2K resolution with synchronized stereo audio, and unifies text-to-image, text-to-video, text-to-audio, video editing, and motion transfer inside one model.

How does it work?

The new H3-VAE tokenizer gives MiniMax roughly a 4× gain in effective sequence length, and long contextual prompts compress from ~100K tokens down to ~4K. A single H3 request accepts up to nine reference images, three video clips, and three audio tracks as context, so users can restyle a clip, transfer motion from a reference, or generate voice and sound effects in one call.

Why does it matter?

Open, weight-available video generation at 2K with native audio has not existed before at this quality level — most rivals either cap out at 720p, drop audio, or lock the weights behind a proprietary API. MiniMax is also pricing 2K generation at under a third of mainstream competitors and says weights will hit Hugging Face within days, which puts real pressure on Runway, Kling, and Sora clones.

Who is it for?

creators, ad agencies, and video-app builders wanting open 2K video with audio

Frequently asked questions

When will the MiniMax H3 weights actually be downloadable?
MiniMax says weights will land on Hugging Face within days of the July 31 announcement, following the same open-weights pattern it used for M2 and M3. The model is already usable through the MiniMax API and hosted providers such as fal.ai while the open release finishes.
What can MiniMax H3 do besides text-to-video?
MiniMax H3 unifies text-to-image, text-to-video, text-to-audio, video editing, and motion transfer in one model — no task-specific model swap. A single request can mix up to nine reference images, three video clips, and three audio tracks, letting users transfer motion from a reference clip or restyle an existing video from a prompt.
How does MiniMax H3 compare on price?
MiniMax says H3's per-second 2K price is under a third of what mainstream rivals charge, and its 768p tier is under half the cost of competing 720p offerings. That aligns with the company's earlier LLM strategy of undercutting Chinese and Western frontier peers on token cost.
How is MiniMax H3 doing this at 2K in one model?
MiniMax H3 uses a new H3-VAE tokenizer that MiniMax says gives a 4× gain in effective sequence length, and compresses long contextual prompts from roughly 100K tokens down to about 4K. That lets one model handle 15-second 2K generations with stereo audio without splitting into separate stages.

Try it

hailuoai.video (H3 preview) via the MiniMax API

Sources · 4 outlets

Tags

  • minimax
  • minimax-h3
  • video-generation
  • text-to-video
  • multimodal
  • audio-generation
  • chinese-labs
  • open-weights
  • hailuo

← All releases · Learn AI