Access the complete MiniMax model suite, engineered by the Shanghai-based team behind Hailuo. The API features the M3 model family—optimized for long-context code execution and autonomous software agents—alongside the Hailuo video generation engine, which ranks #1 on WorldModelBench for physical simulation. Hailuo delivers cutting-edge fluid dynamics and photorealistic motion control, generating commercial-grade video that strictly adheres to physical mechanics.
MiniMax
All Models
MiniMax H3
3 variants available

MiniMax H3 Image-to-Video
MiniMax H3 Image-to-Video is MiniMax's next-generation multimodal AI video model. Supporting first-frame driving and first-to-last frame transitions, it generates up to 2K cinematic HD videos directly, with durations ranging from 5 to 15 seconds. Built on a unified Omni architecture, it natively supports integrated audio-video generation (sound effects, ambient audio, and multilingual lip-sync) alongside exceptional camera control, physics simulation, and subject consistency—ideal for e-commerce, commercial ads, and short drama production.

MiniMax H3 Text-to-Video
MiniMax H3 Text-to-Video is MiniMax's next-generation AI video generation model. Powered by a unified Omni architecture, it accurately parses complex prompt text to directly generate up to 2K cinematic-grade videos up to 15 seconds long. It natively supports integrated audio-video generation (ambient audio, sound effects, and multilingual lip-sync) alongside exceptional motion smoothness, physical simulation, and camera control—ideal for commercial advertising, short dramas, and social media video creation.

MiniMax H3
MiniMax H3 : generate a video that keeps the subject from a reference image, driven by a text prompt. Supports 2K, 5-15s.
MiniMax
2 variants available

MiniMax M3
MiniMax-M3 is MiniMax’s latest M-series multimodal foundation model, built for agentic reasoning, tool use, coding, and long-context tasks. It supports text, image, and video inputs with text output, offering a 1M-token context window, extended thinking, function calling, and structured outputs. With strong capabilities in long-horizon agent workflows, software development, multimodal understanding, and extended response generation, MiniMax-M3 is ideal for autonomous agents, coding assistants, document and video analysis, and production-grade applications that require massive context at a competitive cost.

MiniMax M2.7
MiniMax-M2.7 is a next-generation large language model built for autonomous real-world productivity and continuous improvement. Through advanced agentic capabilities and multi-agent collaboration, it can plan, execute, evaluate, and refine complex tasks in dynamic environments while actively contributing to its own evolution. Optimized for production-grade workflows, M2.7 excels at live debugging, root-cause analysis, financial modeling, and end-to-end document creation across Word, Excel, and PowerPoint. It achieves 56.2% on SWE-Pro, 57.0% on Terminal Bench 2, and a 1495 ELO rating on GDPval-AA, establishing a new benchmark for multi-agent systems operating in real-world digital workflows.
MiniMax Models API Pricing Details
| Model | Pricing (USD) | Our Pricing (USD) | Discount | |
|---|---|---|---|---|
| MiniMax H3 Image-to-Video | $0.1/SEC | Start from$0.06/SEC | -40% | |
| MiniMax H3 Text-to-Video | $0.1/SEC | Start from$0.06/SEC | -40% | |
| MiniMax M2.7 | $0.3/M Tokens | Start from$0.3/M Tokens | — | |
| MiniMax M3 | $0.6/M Tokens | Start from$0.6/M Tokens | — | |
| MiniMax H3 | $0.1/SEC | Start from$0.06/SEC | -40% |