Alibaba

Integrate Alibaba's comprehensive model family through a single unified API endpoint. Access both the Qwen suite for state-of-the-art language and multimodal vision capabilities, and the Wan model series for high-definition video generation up to 1080p.

Alibaba

All Models

Qwen

3 variants available

Qwen 3.5 Omni Plus

Qwen 3.5 Omni Plus

Qwen 3.5 Omni Plus is Alibaba Cloud's next-generation native multimodal (Omni) flagship model. Powered by an end-to-end unified architecture, it supports real-time full-duplex interaction across text, image, audio, and video inputs and outputs. It features ultra-low latency streaming inference, expressive speech synthesis, joint vision-audio reasoning, and strong tool-use capability—ideal for real-time voice assistants, AI customer service, video interaction, and digital human agents.

OfficialAudio-to-TextInputAudioOutputText
$7.25/M Tokens
Qwen 3.7 Plus

Qwen 3.7 Plus

Qwen3.7-Plus is a cost-effective model in Alibaba’s Qwen3.7 series, supporting text and image inputs with text output. It combines enhanced vision-language capabilities with full-stack, agentic intelligence for coding, tool use, and productivity workflows. Its key strength is multimodal, interactive agent functionality: it can understand real-world scenes, interpret on-screen content, interact with graphical interfaces, generate code from visual references, and autonomously navigate mobile apps from end to end.

OfficialLLM
$0.4/M Tokens
Qwen 3.7 Max

Qwen 3.7 Max

Qwen3.7-Max is Alibaba’s flagship model in the Qwen3.7 series, designed for agentic, text-based workflows. It excels at coding, debugging, office automation, productivity tasks, tool use, and long-horizon autonomous execution. With a 1-million-token context window and support for outputs of up to 64K tokens, it is ideal for processing large documents, repository-scale coding, multi-step planning, structured content generation, and complex workflows requiring sustained reasoning across hundreds or even thousands of steps.

OfficialLLM
$2.5/M Tokens

Happy Horse

6 variants available

HappyHorse 1.1 Reference-to-Video Spicy

HappyHorse 1.1 Reference-to-Video Spicy

HappyHorse 1.1 Reference-to-Video Spicy is Alibaba's high-expressiveness, high-dynamism variant in the HappyHorse model series. Building upon multi-image subject locking and native audio-visual synchronization, the Spicy mode is a uncensored and fine-tuned for high-intensity action, aggressive camera tracking, strong visual impact, and dramatic visual effects. It leverages reference images to drive bold, highly expressive motion sequences and complex camera maneuvers while maintaining multilingual lip-sync and ambient audio—ideal for action-packed short dramas, viral social video, game CG FX, and high-impact commercial ads.

OfficialImage-to-VideoInputImageOutputVideo
$0.14/SEC
HappyHorse 1.1 Reference-to-Video

HappyHorse 1.1 Reference-to-Video

HappyHorse 1.1 Reference-to-Video is Alibaba's next-generation AI reference-guided video generation model. Supporting 1 to 9 reference images (covering character identity, outfits, product silhouette, and scene styles), it achieves high-precision multi-image fusion and subject locking to prevent visual drift across shots. It generates native 720P/1080P HD videos up to 15 seconds long per run.

OfficialImage-to-VideoInputImageOutputVideo
$0.14/SEC
HappyHorse 1.1 Text-to-Video

HappyHorse 1.1 Text-to-Video

HappyHorse 1.1 Text-to-Video is Alibaba's next-generation AI text-to-video model. Built on an integrated audio-video joint generation architecture, it generates native 720P/1080P HD videos directly from text, supporting up to 15 seconds of rendering. It natively supports multilingual lip-sync, ambient sound effects, and audio-visual synchronization without extra dubbing, while delivering exceptional motion smoothness, subject consistency, and camera control—ideal for short dramas, commercial ads, and social media video production.

OfficialText-to-VideoInputTextOutputVideo
$0.14/SEC
HappyHorse 1.1 Image-to-Video

HappyHorse 1.1 Image-to-Video

HappyHorse 1.1 Image-to-Video is Alibaba's next-generation AI image-to-video model. Supporting first-frame driving, first-to-last frame transitions, and short video extension, it generates native 720P/1080P HD videos up to 15 seconds per run. Built on an integrated audio-video joint generation architecture, it natively supports multilingual lip-sync, ambient sound effects, and audio-visual synchronization, while delivering exceptional motion smoothness, subject consistency, and camera control—ideal for e-commerce, short drama VFX, and social media production.

OfficialImage-to-VideoInputImageOutputVideo
$0.14/SEC
HappyHorse 1.1 Text-to-Video Spicy

HappyHorse 1.1 Text-to-Video Spicy

HappyHorse 1.1 Text to Video Spicy turns simple text prompts into short cinematic clips, blending impressive temporal stability with expressive, nuanced character movement.

SpicyText-to-VideoInputTextOutputVideo
$0.14/SEC
HappyHorse 1.1 Image-to-Video Spicy

HappyHorse 1.1 Image-to-Video Spicy

HappyHorse 1.1 Image to Video Spicy transforms a single starting image into a short cinematic video, combining rock-solid temporal consistency with smooth, highly expressive character movement.

SpicyImage-to-VideoInputImageOutputVideo
$0.14/SEC

Wan 2.7

4 variants available

Wan 2.7 Image-to-Video

Wan 2.7 Image-to-Video

Tongyi Wanxiang Wan 2.7 Image-to-Video is Alibaba Cloud's next-generation image-to-video model. It supports first-frame generation, first-to-last frame transitions, and short video extension, generating 720P/1080P HD videos up to 15 seconds per run. Featuring powerful camera control and realistic physics simulation, it natively supports audio-driven lip-sync and action alignment while adaptively supporting mainstream aspect ratios like 16:9 and 9:16—ideal for e-commerce, VFX, and film post-production.

OfficialImage-to-VideoInputImageOutputVideo
$0.1/SEC
Wan 2.7 Text-to-Video

Wan 2.7 Text-to-Video

Tongyi Wanxiang Wan 2.7 Text-to-Video is Alibaba Cloud's next-generation text-to-video model. Supporting long Chinese and English prompts with smart expansion, it generates native 720P/1080P HD videos directly from text, up to 15 seconds per run. Featuring native audio-visual coordination and sound effect sync, it adaptively supports multiple aspect ratios like 16:9 and 9:16 with strong motion continuity, realistic physics simulation, and lighting rendering—ideal for commercial ads, short videos, and anime creation.

OfficialText-to-VideoInputTextOutputVideo
$0.1/SEC
Wan 2.7 Text-to-Video Spicy

Wan 2.7 Text-to-Video Spicy

Wan 2.7 Text-to-Video Spicy turns simple text prompts into short cinematic clips, blending impressive temporal stability with expressive, nuanced character movement.

SpicyText-to-VideoInputTextOutputVideo
$0.1/SEC
Wan 2.7 Image-to-Video Spicy

Wan 2.7 Image-to-Video Spicy

Wan 2.7 Image-to-Video Spicy turns a first-frame image into short cinematic motion with stable temporal detail and expressive character movement.

SpicyImage-to-VideoInputImageOutputVideo
$0.1/SEC

Alibaba Models API Pricing Details

ModelPricing (USD)Our Pricing (USD)Discount
HappyHorse 1.1 Reference-to-Video Spicy$0.14/SECStart from$0.14/SEC
HappyHorse 1.1 Reference-to-Video$0.14/SECStart from$0.14/SEC
Qwen 3.5 Omni Plus$7.25/M TokensStart from$7.25/M Tokens
Wan 2.7 Image-to-Video$0.1/SECStart from$0.1/SEC
Wan 2.7 Text-to-Video$0.1/SECStart from$0.1/SEC
HappyHorse 1.1 Text-to-Video$0.14/SECStart from$0.14/SEC
HappyHorse 1.1 Image-to-Video$0.14/SECStart from$0.14/SEC
Qwen 3.7 Plus$0.4/M TokensStart from$0.4/M Tokens
Qwen 3.7 Max$2.5/M TokensStart from$2.5/M Tokens
HappyHorse 1.1 Image-to-Video Spicy$0.14/SECStart from$0.14/SEC
HappyHorse 1.1 Text-to-Video Spicy$0.14/SECStart from$0.14/SEC
Wan 2.7 Image-to-Video Spicy$0.1/SECStart from$0.1/SEC
Wan 2.7 Text-to-Video Spicy$0.1/SECStart from$0.1/SEC

Explore models from other providers