
Wan 3.0 Prime Text-to-Video Spicy
Wan 3.0 Prime Text-to-Video Spicy is Alibaba's high-speed, high-expressiveness AI text-to-video model variant under the Tongyi Wanxiang family. Combining Prime's ultra-fast generation inference with Spicy's high-dynamism visual tuning, it deeply parses complex text prompts to directly render up to 30-second 1080P HD videos with significantly reduced wait times. Fine-tuned for bold physical movement, high-contrast lighting, dramatic camera maneuvers, and intense visual impact, it natively supports integrated audio-video generation (ambient audio, sound effects, and multilingual lip-sync)—delivering rapid turnarounds and extreme visual tension for high-energy social media content, action sequences, and commercial ads.
Read Me
Wan 3.0 Prime Text-to-Video Spicy API
Wan 3.0 Prime is Alibaba's Tongyi Lab's premium video generation model released in late August 2026, as the Prime tier of the Wan 3.0 family. This document covers its Text-to-Video Spicy endpoint: generating high-quality video from text prompts alone, with 7× inference speedup over the standard version, native 30-second single-pass generation, 480P/720P/1080P resolution tiers, native synchronized audio and multilingual dubbing, and 30 FPS output. Spicy mode is the uncensored global variant of Wan 3.0 Prime — it removes standard content restrictions while preserving the same generation capabilities, offering greater creative freedom.
On iCreat, developers call the aliyun/wan3-0-prime/text-to-video-global endpoint through an asynchronous task-based REST API. Pricing: 480P $0.068/sec, 720P $0.14/sec, 1080P $0.28/sec, billed by output duration.
Model Positioning
Wan 3.0 Prime Text-to-Video Spicy targets uncensored premium text-driven video creation — the Spicy global T2V endpoint of Wan 3.0 Prime. Compared to the standard Wan 3.0 T2V Spicy, Prime offers 7× inference speedup, first/last-frame control and up to ~20 reference images (unified model capabilities), with stronger visual quality and consistency. Spicy removes content moderation, making it suitable for advertising, art, and concept design that demand freer creative expression. Wan 3.0 Prime is a closed-source API model with no weight downloads.
Core Capabilities
7× inference acceleration
Prime inference speed is 7× faster than standard Wan 3.0, significantly reducing wait times while maintaining generation quality — ideal for creative workflows requiring rapid iteration.
30-second single-pass generation
A single request can generate up to 30 seconds of continuous video — double Wan 2.7's 15-second limit. Supports intelligent duration recommendations and an extend function for continuous narrative.
Three resolution tiers
Supports 480P, 720P, and 1080P. 480P for low-cost draft iteration, 1080P for final delivery.
Native synchronized audio
Audio is generated in the same pass as the video (not added in post), including dialogue, BGM, and sound effects, ensuring temporal alignment between sound and motion. Supports multilingual dubbing.
Real-world fidelity
Significant improvements in micro-expressions, body linkage, spatial layout, physical motion, and on-screen text rendering, reducing the "AI feel" for a real-shot look.
Negative prompts
Supports negative_prompt to describe content that should not appear in the video, improving generation control.
Pricing
| Resolution | Unit Price (USD/sec) | 5-Second Cost |
|---|---|---|
| 480P | $0.068 | $0.34 |
| 720P | $0.14 | $0.70 |
| 1080P | $0.28 | $1.40 |
Total = unit price × output video duration. Billed by actual generated duration with no minimum-length threshold.
Note: Prime pricing is approximately 1.36–1.4× the standard Wan 3.0 (480P $0.068 vs $0.05, 720P $0.14 vs $0.10, 1080P $0.28 vs $0.20).
Use Cases
- Premium short dramas and film pre-visualization: 30-second single-pass generation for continuous narrative
- Advertising and TVC: rapid high-quality dynamic ad production, iterate at 480P, deliver at 1080P
- Music videos: native synchronized audio, visuals and sound in one pass
- Uncensored free creative: advertising, art, and concept design under Spicy mode
- Social media short-form: multi-aspect (vertical 9:16, horizontal 16:9) for all platforms
- Rapid iterative creation: 7× inference acceleration for multi-iteration workflows
Model Comparison
Wan 3.0 Prime T2V Spicy vs. Wan 3.0 T2V Spicy (Standard)
| Dimension | Prime T2V Spicy | Wan 3.0 T2V Spicy |
|---|---|---|
| Endpoint | aliyun/wan3-0-prime/text-to-video-global |
aliyun/wan3-0/text-to-video-global |
| Positioning | Premium tier | Standard tier |
| Inference speed | 7× standard | Baseline |
| Duration | 4–30 seconds | 4–30 seconds |
| Resolution | 480P / 720P / 1080P | 480P / 720P / 1080P |
| Native audio | Supported | Supported |
| First/last-frame control | Supported (unified model) | Not supported |
| Reference images | Up to ~20 (unified model) | Not supported |
| 480P price | $0.068/sec | $0.05/sec |
| 720P price | $0.14/sec | $0.10/sec |
| 1080P price | $0.28/sec | $0.20/sec |
| Content moderation | Uncensored (Spicy) | Uncensored (Spicy) |
| Best for | Premium long narrative, rapid iteration | Standard text-to-video |
Wan 3.0 Prime T2V Spicy vs. Wan 3.0 Prime I2V Spicy and HappyHorse 1.1
| Dimension | Prime T2V Spicy | Prime I2V Spicy | HappyHorse 1.1 |
|---|---|---|---|
| Workflow | Text-to-video | Image-to-video | T2V / I2V / R2V |
| Input | Text only | Reference image + text | Text + images/video/audio |
| Inference acceleration | 7× | 7× | Baseline |
| Duration | 4–30 seconds | 4–30 seconds | 3–15 seconds |
| Resolution | 480P / 720P / 1080P | 480P / 720P / 1080P | 720P / 1080P |
| Native audio | Supported | Supported | Supported |
| Reference images | Not supported (T2V endpoint) | Supported (up to ~20) | Up to 9 |
| 720P price | $0.14/sec | $0.14/sec | $0.14/sec |
| 1080P price | $0.28/sec | $0.28/sec | $0.18/sec |
| Content moderation | Uncensored | Uncensored | Standard |
| Best for | Uncensored premium T2V | Uncensored premium I2V | Multi-reference consistency |
Why Choose Wan 3.0 Prime T2V Spicy?
- 7× inference acceleration: dramatically faster generation for rapid iterative creation
- Uncensored creative freedom: Spicy mode removes content restrictions for advertising, art, and concept design
- 30-second single pass: double Wan 2.7, for continuous narrative and single-take long video
- Native synchronized audio: dialogue, BGM, and sound effects in one pass, no post-production dubbing
- Three tiers + per-second billing: 480P iteration at $0.068/sec, 1080P delivery at $0.28/sec, predictable cost
- Real-world fidelity: micro-expressions, body linkage, and physics significantly improved
Specifications
| Category | Description |
|---|---|
| Model name | Wan 3.0 Prime Text-to-Video Spicy |
| Developer | Alibaba (Tongyi Lab) |
| Endpoint | aliyun/wan3-0-prime/text-to-video-global |
| Release date | Late August 2026 |
| Model type | Premium video generation model (uncensored global variant) |
| Authentication | API Key (Authorization: Bearer) |
input.prompt |
string, required, describes the video content to generate |
input.negative_prompt |
string, optional, content that should not appear in the video |
parameters.resolution |
string, required; 480P, 720P, 1080P |
parameters.ratio |
string, required; 1:1, 9:16, 16:9, 3:4, 4:3 |
parameters.duration |
integer, required; 4–30 seconds |
parameters.watermark |
boolean, optional; default false |
| Output | Video file; 30 FPS, native synchronized audio, up to 30 seconds |
| Frame rate | 30 FPS |
| Inference acceleration | 7× standard Wan 3.0 |
| Max reference images | Not supported (T2V endpoint; unified model supports up to ~20) |
| Billing unit | USD per second |
| API mode | Asynchronous task submission |
Architecture
Wan 3.0 Prime uses the same unified model architecture as Wan 3.0, but with inference optimization delivering 7× speedup. The model consolidates text-to-video, image-to-video, and reference-to-video into a single system, with an "Omni-Reference" system supporting up to ~20 reference assets (text, images, video, audio, documents, web pages). Audio is generated in the same pass as the video, ensuring temporal alignment. Spicy mode removes the standard version's content moderation filter at the review layer, preserving the same generation backbone and reasoning capabilities.
Wan 3.0 Prime was released in late August 2026 as a closed-source API model with no weight downloads. Available on Replicate, Pika API, Atlas Cloud, and other platforms.
Notes
durationmust be an integer between 4 and 30; out-of-range values return an errorpromptis required;negative_promptis optional, used to exclude unwanted content- 480P tier Prime pricing is $0.068/sec, higher than standard $0.05/sec
- 1080P is the highest per-second rate ($0.28/sec); estimate cost in advance for long videos (total = unit price × duration)
- Native audio is default output, including dialogue, BGM, and sound effects, synchronized with visuals
- Spicy mode is uncensored; generated content is not subject to the standard version's content restrictions — comply with local laws and regulations when using
- Fetch results via the result endpoint promptly after completion
- Wan 3.0 Prime is a closed-source API model, not supporting local deployment; for self-hosting, use the open-source Wan 2.2 series



