Wan 3.0 Text-to-Video

aliyun/wan3-0/text-to-video
OfficialText-to-Video

Wan 3.0 Text-to-Video is Alibaba's next-generation AI text-to-video model under the Tongyi Wanxiang series. It deeply parses complex prompt text to directly generate up to 30-second videos in up to 1080P/4K cinematic resolution. Featuring native audio-visual synchronization (ambient audio, sound effects, and multilingual lip-sync), it delivers exceptional motion smoothness, physical simulation, and precise camera control—ideal for commercial advertising, short dramas, film VFX, and social media content creation.

Read Me

Wan 3.0 Text-to-Video API

Wan 3.0 is Alibaba's Tongyi Lab's next-generation video generation model officially released on August 24, 2026. This document covers its Text-to-Video endpoint: generating high-quality video from text prompts alone, with native 30-second single-pass generation, 480P/720P/1080P resolution tiers, native synchronized audio and multilingual dubbing, and 30 FPS output. The model significantly improves micro-expressions, body linkage, spatial layout, physical motion, and text rendering to reduce the "AI feel" and approach real-shot quality.

On iCreat, developers call the aliyun/wan3-0/text-to-video endpoint through an asynchronous task-based REST API. Pricing: 480P $0.05/sec, 720P $0.10/sec, 1080P $0.20/sec, billed by output duration.

Model Positioning

Wan 3.0 Text-to-Video targets text-driven video creation — the T2V endpoint of the Wan 3.0 unified model. Compared to its predecessor Wan 2.7 (15-second max), single-pass duration doubles to 30 seconds, suited for short dramas, ads, and music videos requiring continuous narrative. Supports an optional thinking mode (reasoning before generation) for multi-event sequences and complex compositions. Wan 3.0 is a closed-source API model with no weight downloads, unlike the open-source Wan 2.2 series.

Core Capabilities

30-second single-pass generation

A single request can generate up to 30 seconds of continuous video — double Wan 2.7's 15-second limit. Supports intelligent duration recommendations and an extend function for continuous narrative.

Three resolution tiers

Supports 480P, 720P, and 1080P, defaulting to 1080P. 480P for low-cost draft iteration, 1080P for final delivery.

Native synchronized audio

Audio is generated in the same pass as the video (not added in post), including dialogue, BGM, and sound effects, ensuring temporal alignment between sound and motion. Supports multilingual dubbing.

Real-world fidelity

Significant improvements in micro-expressions, body linkage, spatial layout, physical motion, and on-screen text rendering, reducing the "AI feel" for a real-shot look.

Thinking mode

Optional reasoning mode — the model reasons about composition and motion before generating, suited for multi-event sequences and complex compositions.

Negative prompts

Supports negative_prompt to describe content that should not appear in the video, improving generation control.

Pricing

Resolution Unit Price (USD/sec) 5-Second Cost
480P $0.05 $0.25
720P $0.10 $0.50
1080P $0.20 $1.00

Total = unit price × output video duration. Billed by actual generated duration with no minimum-length threshold.

Note: Alibaba Cloud Bailian domestic list price is 480P ¥0.3/sec, 720P ¥0.6/sec, 1080P ¥1.2/sec; the USD prices above are iCreat platform prices, matching Alibaba's official overseas pricing.

Use Cases

  • Short dramas and film pre-visualization: 30-second single-pass generation for continuous narrative
  • Advertising and TVC: rapid dynamic ad production, iterate at 480P, deliver at 1080P
  • Music videos: native synchronized audio, visuals and sound in one pass
  • Tourism and brand shorts: text-driven generation with controlled per-clip cost
  • Social media short-form: multi-aspect (vertical 9:16, horizontal 16:9) for all platforms
  • Creative concept validation: low-cost rapid scene and composition testing

Model Comparison

Wan 3.0 T2V vs. Wan 2.7

Dimension Wan 3.0 T2V Wan 2.7
Release date August 24, 2026 2026 (earlier)
Max single-pass duration 30 seconds 15 seconds
Resolution 480P / 720P / 1080P 480P / 720P / 1080P
Native audio Supported (same pass) Not supported
Thinking mode Supported Not supported
Document input Supported (DOC/XLS/PPT/PDF/MD) Not supported
Reference assets Up to ~20 Limited
Open source Closed API Partially open (Wan 2.2)
480P price $0.05/sec
1080P price $0.20/sec
Best for Long narrative, native audio, doc-to-video Short video generation

Wan 3.0 T2V vs. HappyHorse 1.1 and Seedance 2.5

Dimension Wan 3.0 T2V HappyHorse 1.1 Seedance 2.5
Developer Alibaba Tongyi Alibaba ByteDance
Release date August 24, 2026 June 22, 2026 2026
Max duration 30 seconds 3–15 seconds 30 seconds
Resolution 480P / 720P / 1080P 720P / 1080P 720P / 1080P
Native audio Supported Supported Supported
Document input Supported (industry first) Not supported Not supported
Reference assets Up to ~20 Up to 9 reference images
480P price $0.05/sec
720P price $0.10/sec $0.14/sec
1080P price $0.20/sec $0.18/sec
Open source Closed Closed Closed
Best for Long narrative, doc-to-video Multi-reference consistency 30-second continuous video

Why Choose Wan 3.0 T2V?

  • 30-second single pass: double Wan 2.7, for continuous narrative and single-take long video
  • Native synchronized audio: dialogue, BGM, and sound effects in one pass, no post-production dubbing
  • Three tiers + per-second billing: 480P iteration at $0.05/sec, 1080P delivery at $0.20/sec, predictable cost
  • Industry-first document input: DOC/XLS/PPT/PDF/MD direct-to-video (unified model)
  • Real-world fidelity: micro-expressions, body linkage, and physics significantly improved
  • Thinking mode: pre-generation reasoning for complex compositions and multi-event sequences

Specifications

Category Description
Model name Wan 3.0 Text-to-Video
Developer Alibaba (Tongyi Lab)
Endpoint aliyun/wan3-0/text-to-video
Release date August 24, 2026 (public beta August 6)
Model type Video generation model (closed-source API)
Authentication API Key (Authorization: Bearer)
input.prompt string, required, describes the video content to generate
input.negative_prompt string, optional, content that should not appear in the video
parameters.resolution string, required; 480P, 720P, 1080P
parameters.ratio string, required; 1:1, 9:16, 16:9, 3:4, 4:3
parameters.duration integer, required; 4–30 seconds
parameters.watermark boolean, optional; default false
Output Video file; 30 FPS, native synchronized audio, up to 30 seconds
Frame rate 30 FPS
Billing unit USD per second
API mode Asynchronous task submission

Architecture

Wan 3.0 uses a unified model architecture, consolidating text-to-video, image-to-video, and reference-to-video into a single endpoint rather than separate models. The model introduces an "Omni-Reference" system supporting up to ~20 reference assets (text, images, video, audio, documents, web pages) for high-fidelity video synthesis. Audio is generated in the same pass as the video, ensuring temporal alignment.

Wan 3.0 entered public beta on August 6, 2026, and launched commercially on August 24. Unlike the open-source Wan 2.2 series, Wan 3.0 is a closed-source API model with no weight downloads. Available on Alibaba Cloud Bailian (Beijing, Singapore regions) and Qianwen AI platform, with a limited 30% discount (August 24 to September 23).

Notes

  • duration must be an integer between 4 and 30; out-of-range values return an error
  • prompt is required; negative_prompt is optional, used to exclude unwanted content
  • 480P tier for low-cost draft iteration ($0.05/sec), 1080P for final delivery ($0.20/sec)
  • 1080P costs more per second than 720P; estimate cost in advance for long videos (total = unit price × duration)
  • Native audio is default output, including dialogue, BGM, and sound effects, synchronized with visuals
  • Thinking mode is optional, suited for multi-event sequences and complex compositions
  • Fetch results via the result endpoint promptly after completion
  • Wan 3.0 is a closed-source API model, not supporting local deployment; for self-hosting, use the open-source Wan 2.2 series

FAQ

What is the difference between Wan 3.0 and Wan 2.7?

Wan 3.0 doubles single-pass duration to 30 seconds (Wan 2.7: 15 seconds), adds native synchronized audio, thinking mode, document input (DOC/XLS/PPT/PDF/MD), and up to ~20 reference assets. Wan 3.0 is a closed-source API; Wan 2.2 series is open-source.

What is the maximum single-pass duration?

Up to 30 seconds of continuous video — double Wan 2.7. Supports intelligent duration recommendations and an extend function.

Is native audio supported?

Yes. Audio is generated in the same pass as the video, including dialogue, BGM, and sound effects, ensuring temporal alignment. Supports multilingual dubbing.

How do I control generation cost?

480P at $0.05/sec for draft iteration, 1080P at $0.20/sec for final delivery. Recommended workflow: iterate prompts at 480P, then generate the final video at 1080P.

Can I self-host?

No. Wan 3.0 is a closed-source API model, not supporting local deployment. For self-hosting, use the open-source Wan 2.2 series (Apache 2.0), though with limited features and duration.

What is document input?

The Wan 3.0 unified model supports direct input of DOC/XLS/PPT/PDF/MD documents and web URLs. The model reads structured content and generates corresponding video — an industry-first direct "static materials to dynamic footage" conversion.