Wan 3.0 Reference-to-Video Spicy

aliyun/wan3-0/reference-to-video-global
SpicyText-to-VideoImage-to-VideoVideo-to-VideoAudio-to-Video

Wan 3.0 Reference-to-Video Spicy is Alibaba's high-motion video synthesis model engineered for high-energy visual dynamics and stylized action generation. Built on an expanded spatio-temporal attention mechanism alongside robust reference feature anchoring, it synthesizes large-scale physical movements, aggressive camera trajectories, and dramatic temporal transitions while preserving strict subject identity and garment texture fidelity. Wan 3.0 Reference-to-Video Spicy powers dynamic action video production, cinematic FX pre-visualization, interactive gaming visual assets, and high-impact commercial advertising pipelines.

Read Me

Wan 3.0 Reference-to-Video Spicy

Wan 3.0 Reference-to-Video Spicy is an ultra-fast, high-quality AI reference-to-video model from Alibaba's Tongyi Wanxiang series, released in August 2026. Combining the ultra-fast rendering capabilities of the architecture with Wan 3.0's multimodal foundation, the model supports single-image first-frame guidance and smooth transitions between first and last frames. It can directly generate 1080p HD videos up to 30 seconds long. The model natively integrates joint audio-video generation, including ambient sound effects and multilingual character lip synchronization. It provides highly accurate motion physics simulation, camera movement control, and subject consistency, widely applicable to dynamic e-commerce product presentations, film and television visual effects, rapid short-drama iteration, and commercial advertising production.

The iCreat platform exposes its Spicy (global, unrestricted) version via the endpoint https://api.icreat.ai/v1/task/submit/aliyun/wan3-0/reference-to-video-global, operating in an asynchronous three-step mode with 480P at $0.05/second, 720P at $0.10/second, and 1080P at $0.20/second.

Model Positioning

Wan 3.0 Reference-to-Video Spicy is positioned as the global, unrestricted version of Wan 3.0's reference-to-video model, offering the same capabilities as the standard version with relaxed content restrictions for global users.

Compared to the standard R2V version, the Spicy version shares identical model capabilities and parameter specifications, served through a separate global endpoint. Compared to the Prime accelerated version, the Spicy version is priced lower, suitable for cost-sensitive scenarios.

On the iCreat platform, the model provides services through an asynchronous three-step workflow where users submit tasks, poll for status, and retrieve results.

Core Capabilities

Multimodal Reference Input

Supports combined input of reference images (reference_image), reference videos (reference_video), and reference audio (reference_audio), simultaneously anchoring visual style, motion patterns, and audio atmosphere for cross-reference consistency retention.

Native Audio-Visual Joint Generation

Natively supports ambient sound effects, background music, and multilingual lip-sync generation, completed in a single inference pass without post-production synthesis.

Multi-Resolution and Long Duration

Supports 480P, 720P, and 1080P resolution tiers, with a maximum single generation of 30 seconds at 30 FPS, meeting creation needs from short clips to complete scenes.

Flexible Aspect Ratio Control

Supports five aspect ratios: 1:1, 9:16, 16:9, 3:4, and 4:3, covering landscape, portrait, square, and special ratio display requirements.

Negative Prompt Support

Supports a negative_prompt parameter, allowing users to describe content that should not appear in the video, enabling precise control over generation results.

Relaxed Content Restrictions

The Spicy (global) version offers relaxed content restrictions compared to the standard version, providing greater creative freedom for global users while maintaining the same model capabilities and pricing.

Pricing

Resolution Unit Price Billing Example Cost
480P $0.05/second Unit price × output duration 5 seconds = $0.25
720P $0.10/second Unit price × output duration 5 seconds = $0.50
1080P $0.20/second Unit price × output duration 5 seconds = $1.00

Total cost = unit price × output video duration.

Note: The above are iCreat platform prices, consistent with Alibaba Cloud Model Studio official pricing (¥0.45/0.9/1.8 per second, 480P/720P/1080P). The iCreat platform price shall prevail.

Application Scenarios

  • E-commerce product dynamic visualization and style-consistent advertising video production
  • Short drama creation and film VFX segment rapid iteration
  • Character consistency-preserving cross-scene video generation
  • Audio-visual synchronized commercial advertising and brand content creation
  • Creative design concept dynamic visualization

Model Comparison

Comparison Table 1: Wan 3.0 R2V Versions (iCreat Endpoint)

Feature R2V Standard R2V Spicy (Global) R2V Prime
Endpoint Path reference-to-video reference-to-video-global reference-to-video (Prime)
Positioning Standard version Global, unrestricted version Accelerated version
480P Price $0.05/second $0.05/second $0.068/second
720P Price $0.10/second $0.10/second $0.14/second
1080P Price $0.20/second $0.20/second $0.28/second
Resolution 480P/720P/1080P 480P/720P/1080P 480P/720P/1080P
Single Duration 4–30 seconds 4–30 seconds 4–30 seconds
Frame Rate 30 FPS 30 FPS 30 FPS
Native Audio Supported Supported Supported
Content Restrictions Standard Relaxed Standard

Note: Standard and Spicy versions are identically priced, distinguished by endpoint path; Prime is priced at approximately 1.4x with faster generation speed.

Comparison Table 2: Wan 3.0 R2V vs T2V vs I2V (Spicy Endpoint)

Feature R2V (Reference-to-Video) T2V (Text-to-Video) I2V (Image-to-Video)
Input Type Text+Image+Video+Audio Text only Text+Image
Reference Image Supported Not supported Supported
Reference Video Supported Not supported Not supported
Reference Audio Supported Not supported Not supported
Resolution 480P/720P/1080P 480P/720P/1080P 480P/720P/1080P
Single Duration 4–30 seconds 4–30 seconds 4–30 seconds
Frame Rate 30 FPS 30 FPS 30 FPS
Price Same Same Same

Note: R2V, T2V, and I2V are all called via the iCreat endpoint at identical pricing; R2V supports the richest reference input types.

Why Choose Wan 3.0 Reference-to-Video Spicy?

  • Global, unrestricted version with relaxed content restrictions for greater creative freedom
  • Supports reference image + video + audio tri-modal combined input
  • Same pricing as the standard version, lower than Prime accelerated version
  • Supports 480P/720P/1080P three resolution tiers, from draft to final output
  • Up to 30 seconds single duration, exceeding the limit of most competing models
  • Native audio-visual joint generation with ambient sound and lip-sync in one pass

Specifications

Field Value
Model Name Wan 3.0 Reference-to-Video Spicy
Developer Alibaba (Tongyi Wanxiang)
Endpoint https://api.icreat.ai/v1/task/submit/aliyun/wan3-0/reference-to-video-global
Query Status Endpoint https://api.icreat.ai/v1/task/query-status
Get Result Endpoint https://api.icreat.ai/v1/task/get-result
Release Date 2026-08
Model Type Multimodal video generation (reference-to-video, global/unrestricted)
Authentication Authorization: Bearer ; X-ICREAT-AI-GROUP: wan-spicy
Input Type prompt (prompt text), negative_prompt (negative prompt), media (reference image/video/audio URL)
Parameters resolution (480P/720P/1080P), ratio (1:1/9:16/16:9/3:4/4:3), duration (4–30 seconds), watermark (default false)
Frame Rate 30 FPS
Native Audio Supported
Billing Unit Per second
API Mode Asynchronous three-step (submit → query-status → get-result)

Architecture

The iCreat platform's Wan 3.0 Reference-to-Video Spicy operates using an asynchronous three-step architecture. Users submit tasks via the submission endpoint, and the system returns a task ID. The task status is then polled via the query-status endpoint, and once complete, the final video result is retrieved via the get-result endpoint. Requests must include the API Key in the Authorization header for authentication, along with the X-ICREAT-AI-GROUP: wan-spicy header. The request body contains the input object (prompt, negative_prompt, media reference array) and the parameters object (resolution, ratio, duration, watermark) as sibling fields. The media array supports three types: reference_image, reference_video, and reference_audio. The model ID is included in the submission endpoint URL path; the request body does not use a model field.

Notes

  • The ratio parameter is required and must be specified when submitting a task
  • The duration parameter ranges from 4 to 30 seconds; out-of-range values will be rejected
  • Reference media (images/videos/audio) must be publicly accessible HTTPS URLs; private or authenticated addresses may cause task failure
  • The model natively supports audio generation; generated videos include audio-visual content by default
  • The Spicy (global) version offers relaxed content restrictions compared to the standard version
  • Please safeguard your API Key and avoid hardcoding it in client-side code or public repositories

Frequently Asked Questions

How is the billing calculated?

Billing is calculated as the unit price multiplied by the output video duration. 480P is $0.05/second, 720P is $0.10/second, and 1080P is $0.20/second. For example, generating a 5-second 720P video costs $0.50. The iCreat platform price shall prevail.

What is the difference between the Spicy and Standard versions?

The Spicy (global) version and the Standard version share identical model capabilities, parameter specifications, and pricing, distinguished by endpoint path (reference-to-video-global vs reference-to-video). The Spicy version offers relaxed content restrictions for global users, using the X-ICREAT-AI-GROUP: wan-spicy header.

What is the difference between R2V, T2V, and I2V?

R2V (reference-to-video) supports combined input of reference images, reference videos, and reference audio. T2V (text-to-video) requires only text prompt input. I2V (image-to-video) supports text prompts and reference images. All three are priced identically; choose R2V if you need to use reference videos or audio.

What values does the ratio parameter support?

Five aspect ratios: 1:1 (square), 9:16 (portrait), 16:9 (landscape), 3:4 (portrait special), and 4:3 (landscape special). This parameter is required.

What is the maximum single duration?

The video generation duration ranges from 4 to 30 seconds, with a maximum of 30 seconds. The duration parameter must be an integer within this range.

What reference types does the media array support?

Three reference types: reference_image (reference image), reference_video (reference video), and reference_audio (reference audio). They can be combined to simultaneously anchor visual style, motion patterns, and audio atmosphere.