
Wan 3.0 Text-to-Video Spicy
Wan 3.0 Text-to-Video Spicy is Alibaba's high-expressiveness AI text-to-video model variant under the Tongyi Wanxiang series. Building upon the core Wan 3.0 architecture, the Spicy edition is fine-tuned for high-intensity physical motion, dramatic camera maneuvering, high-contrast lighting, and powerful visual impact, directly generating up to 30-second videos in up to 1080P cinematic resolution.
Read Me
Wan 3.0 Text-to-Video Spicy API
Wan 3.0 is Alibaba's Tongyi Lab's next-generation video generation model officially released on August 24, 2026. This document covers its Text-to-Video Spicy endpoint: generating high-quality video from text prompts alone, with native 30-second single-pass generation, 480P/720P/1080P resolution tiers, native synchronized audio and multilingual dubbing, and 30 FPS output. Spicy mode is the uncensored global variant of Wan 3.0 — it removes standard content restrictions while preserving the same generation capabilities, offering greater creative freedom.
On iCreat, developers call the aliyun/wan3-0/text-to-video-global endpoint through an asynchronous task-based REST API. Pricing: 480P $0.05/sec, 720P $0.10/sec, 1080P $0.20/sec, billed by output duration.
Model Positioning
Wan 3.0 Text-to-Video Spicy targets uncensored text-driven video creation — the Spicy global endpoint of Wan 3.0. Compared to the standard version (text-to-video), Spicy removes content moderation while maintaining the same generation quality, making it suitable for advertising, art, and concept design that demand freer creative expression. Single-pass duration up to 30 seconds (double Wan 2.7), with thinking mode and native synchronized audio. Wan 3.0 is a closed-source API model with no weight downloads.
Core Capabilities
30-second single-pass generation
A single request can generate up to 30 seconds of continuous video — double Wan 2.7's 15-second limit. Supports intelligent duration recommendations and an extend function for continuous narrative.
Three resolution tiers
Supports 480P, 720P, and 1080P, defaulting to 1080P. 480P for low-cost draft iteration, 1080P for final delivery.
Native synchronized audio
Audio is generated in the same pass as the video (not added in post), including dialogue, BGM, and sound effects, ensuring temporal alignment between sound and motion. Supports multilingual dubbing.
Real-world fidelity
Significant improvements in micro-expressions, body linkage, spatial layout, physical motion, and on-screen text rendering, reducing the "AI feel" for a real-shot look.
Thinking mode
Optional reasoning mode — the model reasons about composition and motion before generating, suited for multi-event sequences and complex compositions.
Negative prompts
Supports negative_prompt to describe content that should not appear in the video, improving generation control.
Pricing
| Resolution | Unit Price (USD/sec) | 5-Second Cost |
|---|---|---|
| 480P | $0.05 | $0.25 |
| 720P | $0.10 | $0.50 |
| 1080P | $0.20 | $1.00 |
Total = unit price × output video duration. Billed by actual generated duration with no minimum-length threshold.
Note: Alibaba Cloud Bailian domestic list price is 480P ¥0.3/sec, 720P ¥0.6/sec, 1080P ¥1.2/sec; the USD prices above are iCreat platform prices, matching Alibaba's official overseas pricing.
Use Cases
- Short dramas and film pre-visualization: 30-second single-pass generation for continuous narrative
- Advertising and TVC: rapid dynamic ad production, iterate at 480P, deliver at 1080P
- Music videos: native synchronized audio, visuals and sound in one pass
- Uncensored free creative: advertising, art, and concept design under Spicy mode
- Social media short-form: multi-aspect (vertical 9:16, horizontal 16:9) for all platforms
- Creative concept validation: low-cost rapid scene and composition testing
Model Comparison
Wan 3.0 T2V Spicy vs. Wan 3.0 T2V (Standard)
| Dimension | T2V Spicy | T2V (Standard) |
|---|---|---|
| Endpoint | aliyun/wan3-0/text-to-video-global |
aliyun/wan3-0/text-to-video |
| Content moderation | Uncensored (restrictions removed) | Standard moderation |
| Generation capability | Same | Same |
| Resolution | 480P / 720P / 1080P | 480P / 720P / 1080P |
| Duration | 4–30 seconds | 4–30 seconds |
| Native audio | Supported | Supported |
| Thinking mode | Supported | Supported |
| Price | Same | Same |
| Best for | Free creative, art, concept design | Standard commercial assets |
Wan 3.0 T2V Spicy vs. Wan 2.7 and HappyHorse 1.1
| Dimension | T2V Spicy | Wan 2.7 | HappyHorse 1.1 |
|---|---|---|---|
| Developer | Alibaba Tongyi | Alibaba Tongyi | Alibaba |
| Release date | August 24, 2026 | 2026 (earlier) | June 22, 2026 |
| Max duration | 30 seconds | 15 seconds | 3–15 seconds |
| Resolution | 480P / 720P / 1080P | 480P / 720P / 1080P | 720P / 1080P |
| Native audio | Supported | Not supported | Supported |
| Document input | Supported (unified model) | Not supported | Not supported |
| Content moderation | Uncensored (Spicy) | Standard | Standard |
| 480P price | $0.05/sec | — | — |
| 720P price | $0.10/sec | — | $0.14/sec |
| 1080P price | $0.20/sec | — | $0.18/sec |
| Open source | Closed | Partially open (Wan 2.2) | Closed |
| Best for | Uncensored long narrative, native audio | Short video generation | Multi-reference consistency |
Why Choose Wan 3.0 T2V Spicy?
- Uncensored creative freedom: Spicy mode removes content restrictions for advertising, art, and concept design
- 30-second single pass: double Wan 2.7, for continuous narrative and single-take long video
- Native synchronized audio: dialogue, BGM, and sound effects in one pass, no post-production dubbing
- Three tiers + per-second billing: 480P iteration at $0.05/sec, 1080P delivery at $0.20/sec, predictable cost
- Real-world fidelity: micro-expressions, body linkage, and physics significantly improved
- Thinking mode: pre-generation reasoning for complex compositions and multi-event sequences
Specifications
| Category | Description |
|---|---|
| Model name | Wan 3.0 Text-to-Video Spicy |
| Developer | Alibaba (Tongyi Lab) |
| Endpoint | aliyun/wan3-0/text-to-video-global |
| Release date | August 24, 2026 (public beta August 6) |
| Model type | Video generation model (uncensored global variant) |
| Authentication | API Key (Authorization: Bearer) |
input.prompt |
string, required, describes the video content to generate |
input.negative_prompt |
string, optional, content that should not appear in the video |
parameters.resolution |
string, required; 480P, 720P, 1080P |
parameters.ratio |
string, required; 1:1, 9:16, 16:9, 3:4, 4:3 |
parameters.duration |
integer, required; 4–30 seconds |
parameters.watermark |
boolean, optional; default false |
| Output | Video file; 30 FPS, native synchronized audio, up to 30 seconds |
| Frame rate | 30 FPS |
| Billing unit | USD per second |
| API mode | Asynchronous task submission |
Architecture
Wan 3.0 uses a unified model architecture, consolidating text-to-video, image-to-video, and reference-to-video into a single system. The model introduces an "Omni-Reference" system supporting up to ~20 reference assets (text, images, video, audio, documents, web pages) for high-fidelity video synthesis. Audio is generated in the same pass as the video, ensuring temporal alignment. Spicy mode removes the standard version's content moderation filter at the review layer, preserving the same generation backbone and reasoning capabilities.
Wan 3.0 entered public beta on August 6, 2026, and launched commercially on August 24. Unlike the open-source Wan 2.2 series, Wan 3.0 is a closed-source API model with no weight downloads. Available on Alibaba Cloud Bailian (Beijing, Singapore regions) and Qianwen AI platform.
Notes
durationmust be an integer between 4 and 30; out-of-range values return an errorpromptis required;negative_promptis optional, used to exclude unwanted content- 480P tier for low-cost draft iteration ($0.05/sec), 1080P for final delivery ($0.20/sec)
- 1080P costs more per second than 720P; estimate cost in advance for long videos (total = unit price × duration)
- Native audio is default output, including dialogue, BGM, and sound effects, synchronized with visuals
- Thinking mode is optional, suited for multi-event sequences and complex compositions
- Spicy mode is uncensored; generated content is not subject to the standard version's content restrictions — comply with local laws and regulations when using
- Fetch results via the result endpoint promptly after completion
- Wan 3.0 is a closed-source API model, not supporting local deployment; for self-hosting, use the open-source Wan 2.2 series
FAQ
What is the difference between Spicy and the standard version?
Spicy is the uncensored global variant of Wan 3.0 T2V, removing the standard version's content moderation filter while maintaining the same generation capability. Both share the same price, resolution, duration, and features. The endpoints differ: Spicy uses text-to-video-global, the standard version uses text-to-video.
What is the maximum single-pass duration?
Up to 30 seconds of continuous video — double Wan 2.7. Supports intelligent duration recommendations and an extend function.
Is native audio supported?
Yes. Audio is generated in the same pass as the video, including dialogue, BGM, and sound effects, ensuring temporal alignment. Supports multilingual dubbing.
How do I control generation cost?
480P at $0.05/sec for draft iteration, 1080P at $0.20/sec for final delivery. Recommended workflow: iterate prompts at 480P, then generate the final video at 1080P.
What is the difference from Wan 2.7?
Wan 3.0 Spicy doubles single-pass duration to 30 seconds (Wan 2.7: 15 seconds), adds native synchronized audio, thinking mode, document input, and up to ~20 reference assets. Wan 3.0 is a closed-source API; Wan 2.2 series is open-source.
Can I self-host?
No. Wan 3.0 is a closed-source API model, not supporting local deployment. For self-hosting, use the open-source Wan 2.2 series (Apache 2.0), though with limited features and duration.



