MiniMax H3
Moving beyond traditional single-task pipelines, MiniMax H3 leverages the H3-Omni Transformer architecture to deliver unified understanding and generation across interleaved text, image, video, and audio contexts. The model generates smooth, continuous video clips from 4 to 15 seconds at up to 2K crisp resolution, natively accompanied by 32kHz stereo audio including lip-synced speech, ambient foley, and soundtrack effects.
All Models

MiniMax H3 Image-to-Video
MiniMax H3 Image-to-Video is MiniMax's next-generation multimodal AI video model. Supporting first-frame driving and first-to-last frame transitions, it generates up to 2K cinematic HD videos directly, with durations ranging from 5 to 15 seconds. Built on a unified Omni architecture, it natively supports integrated audio-video generation (sound effects, ambient audio, and multilingual lip-sync) alongside exceptional camera control, physics simulation, and subject consistency—ideal for e-commerce, commercial ads, and short drama production.

MiniMax H3 Text-to-Video
MiniMax H3 Text-to-Video is MiniMax's next-generation AI video generation model. Powered by a unified Omni architecture, it accurately parses complex prompt text to directly generate up to 2K cinematic-grade videos up to 15 seconds long. It natively supports integrated audio-video generation (ambient audio, sound effects, and multilingual lip-sync) alongside exceptional motion smoothness, physical simulation, and camera control—ideal for commercial advertising, short dramas, and social media video creation.

MiniMax H3
MiniMax H3 : generate a video that keeps the subject from a reference image, driven by a text prompt. Supports 2K, 5-15s.
MiniMax H3 Models API Pricing Details
| Model | Pricing (USD) | Our Pricing (USD) | Discount | |
|---|---|---|---|---|
| MiniMax H3 Image-to-Video | $0.1/SEC | Start from$0.06/SEC | -40% | |
| MiniMax H3 Text-to-Video | $0.1/SEC | Start from$0.06/SEC | -40% | |
| MiniMax H3 | $0.1/SEC | Start from$0.06/SEC | -40% |