
Wan 3.0 Prime Reference-to-Video
Wan 3.0 Prime Reference-to-Video is Alibaba's flagship controllable video generation model built for precise reference-conditioned visual synthesis. Powered by an upgraded spatio-temporal decoupling architecture and multi-reference feature fusion, it preserves character identity, garment textures, ambient lighting, and complex camera trajectories across extended sequences while generating native 4K high-frame-rate video. Operating with strict temporal consistency and fluid motion dynamics, Wan 3.0 Prime drives e-commerce video production, cinematic pre-visualization, digital human animation, and commercial advertising pipelines.
Read Me
Wan 3.0 Prime Reference-to-Video
Wan 3.0 Prime Reference-to-Video is a video generation model developed by Alibaba's Tongyi Wanxiang team, released in August 2026. The model combines the Prime architecture's ultra-fast rendering with the Wan 3.0 omni-modal foundation, supporting multimodal combined input of reference images, reference videos, and reference audio to generate dynamic videos with high-precision physical motion simulation, camera control, and subject consistency. The model supports 480P, 720P, and 1080P resolution output, with a maximum duration of 30 seconds at 30 FPS and native audio-visual joint generation (ambient sound effects, background music, and multilingual lip-sync).
The iCreat platform exposes its reference-to-video capability via the endpoint https://api.icreat.ai/v1/task/submit/aliyun/wan3-0-prime/reference-to-video, operating in an asynchronous three-step mode with 480P at $0.068/second, 720P at $0.14/second, and 1080P at $0.28/second.
Model Positioning
Wan 3.0 Prime Reference-to-Video is positioned as the reference-to-video model of the Wan 3.0 Prime series, significantly boosting generation speed while maintaining high-quality video generation capabilities.
Compared to image-to-video (I2V) which only supports reference images, reference-to-video (R2V) additionally supports reference video and reference audio input, enabling simultaneous anchoring of visual style, motion patterns, and audio atmosphere — suitable for production scenarios with higher requirements for visual consistency and audio-visual synchronization.
On the iCreat platform, the model provides services through an asynchronous three-step workflow where users submit tasks, poll for status, and retrieve results.
Core Capabilities
Multimodal Reference Input
Supports combined input of reference images (reference_image), reference videos (reference_video), and reference audio (reference_audio), simultaneously anchoring visual style, motion patterns, and audio atmosphere for cross-reference consistency retention.
Native Audio-Visual Joint Generation
Natively supports ambient sound effects, background music, and multilingual lip-sync generation, completed in a single inference pass without post-production synthesis.
Multi-Resolution and Long Duration
Supports 480P, 720P, and 1080P resolution tiers, with a maximum single generation of 30 seconds at 30 FPS, meeting creation needs from short clips to complete scenes.
Flexible Aspect Ratio Control
Supports five aspect ratios: 1:1, 9:16, 16:9, 3:4, and 4:3, covering landscape, portrait, square, and special ratio display requirements.
Negative Prompt Support
Supports a negative_prompt parameter, allowing users to describe content that should not appear in the video, enabling precise control over generation results.
Pricing
| Resolution | Unit Price | Billing | Example Cost |
|---|---|---|---|
| 480P | $0.068/second | Unit price × output duration | 5 seconds = $0.34 |
| 720P | $0.14/second | Unit price × output duration | 5 seconds = $0.70 |
| 1080P | $0.28/second | Unit price × output duration | 5 seconds = $1.40 |
Total cost = unit price × output video duration.
Note: The above are iCreat platform prices. Alibaba Cloud Model Studio official pricing is ¥0.45/0.9/1.8 per second (480P/720P/1080P). As a third-party relay platform, iCreat's price includes platform service fees. The iCreat platform price shall prevail.
Application Scenarios
- E-commerce product dynamic visualization and style-consistent advertising video production
- Short drama creation and film VFX segment rapid iteration
- Character consistency-preserving cross-scene video generation
- Audio-visual synchronized commercial advertising and brand content creation
- Creative design concept dynamic visualization
Model Comparison
Comparison Table 1: Wan 3.0 Prime R2V vs T2V vs I2V (iCreat Endpoint)
| Feature | R2V (Reference-to-Video) | T2V (Text-to-Video) | I2V (Image-to-Video) |
|---|---|---|---|
| Input Type | Text+Image+Video+Audio | Text only | Text+Image |
| Reference Image | Supported | Not supported | Supported |
| Reference Video | Supported | Not supported | Not supported |
| Reference Audio | Supported | Not supported | Not supported |
| Resolution | 480P/720P/1080P | 480P/720P/1080P | 480P/720P/1080P |
| Single Duration | 4–30 seconds | 4–30 seconds | 4–30 seconds |
| Frame Rate | 30 FPS | 30 FPS | 30 FPS |
| Price | Same | Same | Same |
Note: R2V, T2V, and I2V are all called via the iCreat endpoint at identical pricing; R2V supports the richest reference input types.
Comparison Table 2: Wan 3.0 Prime vs Standard (iCreat Endpoint)
| Feature | Wan 3.0 Prime R2V | Wan 3.0 Standard R2V |
|---|---|---|
| Positioning | Accelerated, faster generation | Standard, balances quality and cost |
| 480P Price | $0.068/second | $0.05/second |
| 720P Price | $0.14/second | $0.10/second |
| 1080P Price | $0.28/second | $0.20/second |
| Resolution | 480P/720P/1080P | 480P/720P/1080P |
| Single Duration | 4–30 seconds | 4–30 seconds |
| Frame Rate | 30 FPS | 30 FPS |
| Native Audio | Supported | Supported |
Note: Prime is priced at approximately 1.4x the standard version, trading cost for faster generation speed; both versions share the same resolution, duration, and audio capabilities.
Why Choose Wan 3.0 Prime Reference-to-Video?
- Supports reference image + video + audio tri-modal combined input, anchoring visual style, motion, and audio atmosphere
- Accelerated generation speed, suitable for efficiency-critical production scenarios
- Supports 480P/720P/1080P three resolution tiers, from draft to final output
- Up to 30 seconds single duration, exceeding the limit of most competing models
- Native audio-visual joint generation with ambient sound and lip-sync in one pass
- Negative prompt support for precise content control
Specifications
| Field | Value |
|---|---|
| Model Name | Wan 3.0 Prime Reference-to-Video |
| Developer | Alibaba (Tongyi Wanxiang) |
| Endpoint | https://api.icreat.ai/v1/task/submit/aliyun/wan3-0-prime/reference-to-video |
| Query Status Endpoint | https://api.icreat.ai/v1/task/query-status |
| Get Result Endpoint | https://api.icreat.ai/v1/task/get-result |
| Release Date | 2026-08 |
| Model Type | Multimodal video generation (reference-to-video) |
| Authentication | Authorization: Bearer ; X-ICREAT-AI-GROUP: wan-spicy |
| Input Type | prompt (prompt text), negative_prompt (negative prompt), media (reference image/video/audio URL) |
| Parameters | resolution (480P/720P/1080P), ratio (1:1/9:16/16:9/3:4/4:3), duration (4–30 seconds), watermark (default false) |
| Frame Rate | 30 FPS |
| Native Audio | Supported |
| Billing Unit | Per second |
| API Mode | Asynchronous three-step (submit → query-status → get-result) |
Architecture
The iCreat platform's Wan 3.0 Prime Reference-to-Video operates using an asynchronous three-step architecture. Users submit tasks via the submission endpoint, and the system returns a task ID. The task status is then polled via the query-status endpoint, and once complete, the final video result is retrieved via the get-result endpoint. Requests must include the API Key in the Authorization header for authentication, along with the X-ICREAT-AI-GROUP header. The request body contains the input object (prompt, negative_prompt, media reference array) and the parameters object (resolution, ratio, duration, watermark) as sibling fields. The media array supports three types: reference_image, reference_video, and reference_audio. The model ID is included in the submission endpoint URL path; the request body does not use a model field.
Notes
- The ratio parameter is required and must be specified when submitting a task
- The duration parameter ranges from 4 to 30 seconds; out-of-range values will be rejected
- Reference media (images/videos/audio) must be publicly accessible HTTPS URLs; private or authenticated addresses may cause task failure
- The model natively supports audio generation; generated videos include audio-visual content by default
- Prime is priced at approximately 1.4x the standard version; consider the standard version if speed is not a priority
- Please safeguard your API Key and avoid hardcoding it in client-side code or public repositories
Frequently Asked Questions
How is the billing calculated?
Billing is calculated as the unit price multiplied by the output video duration. 480P is $0.068/second, 720P is $0.14/second, and 1080P is $0.28/second. For example, generating a 5-second 720P video costs $0.70. The iCreat platform price shall prevail.
What is the difference between R2V, T2V, and I2V?
R2V (reference-to-video) supports combined input of reference images, reference videos, and reference audio, simultaneously anchoring visual style, motion patterns, and audio atmosphere. T2V (text-to-video) requires only text prompt input. I2V (image-to-video) supports text prompts and reference images. All three are priced identically; choose R2V if you need to use reference videos or audio.
What values does the ratio parameter support?
Five aspect ratios: 1:1 (square), 9:16 (portrait), 16:9 (landscape), 3:4 (portrait special), and 4:3 (landscape special). This parameter is required.
What is the maximum single duration?
The video generation duration ranges from 4 to 30 seconds, with a maximum of 30 seconds. The duration parameter must be an integer within this range.
What reference types does the media array support?
Three reference types: reference_image (reference image), reference_video (reference video), and reference_audio (reference audio). They can be combined to simultaneously anchor visual style, motion patterns, and audio atmosphere.
Does the model support audio generation?
Yes. The model natively supports audio-visual joint generation, including ambient sound effects, background music, and multilingual lip-sync, completed in a single inference pass without post-production synthesis.



