Wan 3.0

Text-to-VideoImage-to-Video

Wan 3.0 is Alibaba's next-generation multimodal AI video generation model. Supporting both text-to-video and image-to-video workflows, it generates up to 30-second videos in up to 1080P/4K cinematic resolution.

All Models

Wan 3.0 Prime Reference-to-Video Spicy
NEW
SpicyImage-to-Video

Wan 3.0 Prime Reference-to-Video Spicy

Wan 3.0 Prime Reference-to-Video Spicy is Alibaba's flagship uncensored high-motion video model engineered for native 4K visual synthesis. Delivering native 4K output at 60 fps, it processes up to 8 reference images or 2 reference videos simultaneously to preserve character identity and material fidelity across complex sequences. With unrestricted generation capabilities, it synthesizes explosive physical movements, dramatic camera trajectories, and physical interactions—powering high-energy cinematic pre-visualization, high-impact commercial ads, AAA gaming assets, and action pipelines.

Wan 3.0 Reference-to-Video Spicy
NEW
SpicyText-to-Video

Wan 3.0 Reference-to-Video Spicy

Wan 3.0 Reference-to-Video Spicy is Alibaba's high-motion video synthesis model engineered for high-energy visual dynamics and stylized action generation. Built on an expanded spatio-temporal attention mechanism alongside robust reference feature anchoring, it synthesizes large-scale physical movements, aggressive camera trajectories, and dramatic temporal transitions while preserving strict subject identity and garment texture fidelity. Wan 3.0 Reference-to-Video Spicy powers dynamic action video production, cinematic FX pre-visualization, interactive gaming visual assets, and high-impact commercial advertising pipelines.

Wan 3.0 Prime Reference-to-Video
NEW
OfficialText-to-Video

Wan 3.0 Prime Reference-to-Video

Wan 3.0 Prime Reference-to-Video is Alibaba's flagship controllable video generation model built for precise reference-conditioned visual synthesis. Powered by an upgraded spatio-temporal decoupling architecture and multi-reference feature fusion, it preserves character identity, garment textures, ambient lighting, and complex camera trajectories across extended sequences while generating native 4K high-frame-rate video. Operating with strict temporal consistency and fluid motion dynamics, Wan 3.0 Prime drives e-commerce video production, cinematic pre-visualization, digital human animation, and commercial advertising pipelines.

Wan 3.0 Reference-to-Video
NEW
OfficialText-to-Video

Wan 3.0 Reference-to-Video

Wan 3.0 Reference-to-Video is Alibaba's flagship high-efficiency video generation model engineered for precise reference-driven visual synthesis. Built on an upgraded multi-modal reference feature alignment and spatio-temporal attention architecture, it preserves strict subject identity, garment textures, and ambient lighting across dynamic video sequences. Combining high-throughput inference speeds with fluid camera movement, Wan 3.0 Reference-to-Video efficiently powers e-commerce product videos, digital human animation, social media marketing assets, and commercial advertising pipelines.

Wan 3.0 Prime Image-to-Video
NEW
OfficialImage-to-Video

Wan 3.0 Prime Image-to-Video

Wan 3.0 Prime Image-to-Video is Alibaba's high-speed AI image-to-video model under the Tongyi Wanxiang family. Leveraging the Prime architecture's rapid inference alongside the Wan 3.0 multimodal foundation, it animates source images (with optional end-frame guidance) into up to 30-second 1080P HD videos with significantly reduced rendering latency. Featuring native audio-visual synchronization (ambient audio, sound effects, and multilingual lip-sync), it delivers accurate physical motion simulation, precise camera control, and strong subject consistency—ideal for fast-turnaround e-commerce animation, film VFX, short dramas, and commercial advertising.

Wan 3.0 Prime Text-to-Video
NEW
OfficialText-to-Video

Wan 3.0 Prime Text-to-Video

Wan 3.0 Prime Text-to-Video is Alibaba's high-speed AI text-to-video model under the Tongyi Wanxiang family. Combining the Prime architecture's rapid inference with the core Wan 3.0 multimodal foundation, it deeply parses complex text prompts to directly render up to 30-second 1080P HD videos with significantly reduced generation latency. Featuring native audio-visual synchronization (ambient audio, sound effects, and multilingual lip-sync), it delivers exceptional physical motion simulation, seamless temporal coherence, and precise camera control—providing rapid turnarounds and high-quality visual output for commercial advertising, short dramas, film VFX, and high-frequency social media content creation.

Wan 3.0 Image-to-Video Spicy
NEW
SpicyImage-to-Video

Wan 3.0 Image-to-Video Spicy

Wan 3.0 Image-to-Video Spicy is Alibaba's high-dynamic image-to-video model under the Tongyi Wanxiang framework. Specially engineered for large-scale motion and high visual intensity, it transforms a single static image into up to 30-second videos in cinematic 1080P resolution. While maintaining strict character identity and background consistency from the input image, the "Spicy" edition delivers a dramatic boost in action amplitude, complex physical collision simulation, and expressive camera movement, alongside native audio-visual generation capabilities. It is ideally suited for high-energy commercial advertising, game CG animation, short drama production, and advanced visual effects workflows.

Wan 3.0 Prime Image-to-Video Spicy
NEW
SpicyImage-to-Video

Wan 3.0 Prime Image-to-Video Spicy

Wan 3.0 Prime Image-to-Video Spicy is Alibaba's high-speed, high-expressiveness AI image-to-video model variant under the Tongyi Wanxiang family. Combining Prime's ultra-fast generation inference with Spicy's high-dynamism visual tuning, it animates source images (with optional end-frame guidance) into up to 30-second 1080P HD videos in a single pass.

Wan 3.0 Prime Text-to-Video Spicy
SpicyText-to-Video

Wan 3.0 Prime Text-to-Video Spicy

Wan 3.0 Prime Text-to-Video Spicy is Alibaba's high-speed, high-expressiveness AI text-to-video model variant under the Tongyi Wanxiang family. Combining Prime's ultra-fast generation inference with Spicy's high-dynamism visual tuning, it deeply parses complex text prompts to directly render up to 30-second 1080P HD videos with significantly reduced wait times. Fine-tuned for bold physical movement, high-contrast lighting, dramatic camera maneuvers, and intense visual impact, it natively supports integrated audio-video generation (ambient audio, sound effects, and multilingual lip-sync)—delivering rapid turnarounds and extreme visual tension for high-energy social media content, action sequences, and commercial ads.

Wan 3.0 Text-to-Video Spicy
NEW
SpicyText-to-Video

Wan 3.0 Text-to-Video Spicy

Wan 3.0 Text-to-Video Spicy is Alibaba's high-expressiveness AI text-to-video model variant under the Tongyi Wanxiang series. Building upon the core Wan 3.0 architecture, the Spicy edition is fine-tuned for high-intensity physical motion, dramatic camera maneuvering, high-contrast lighting, and powerful visual impact, directly generating up to 30-second videos in up to 1080P cinematic resolution.

Wan 3.0 Text-to-Video
OfficialText-to-Video

Wan 3.0 Text-to-Video

Wan 3.0 Text-to-Video is Alibaba's next-generation AI text-to-video model under the Tongyi Wanxiang series. It deeply parses complex prompt text to directly generate up to 30-second videos in up to 1080P/4K cinematic resolution. Featuring native audio-visual synchronization (ambient audio, sound effects, and multilingual lip-sync), it delivers exceptional motion smoothness, physical simulation, and precise camera control—ideal for commercial advertising, short dramas, film VFX, and social media content creation.

Wan 3.0 Image-to-Video
NEW
OfficialImage-to-Video

Wan 3.0 Image-to-Video

Wan 3.0 Image-to-Video is Alibaba's next-generation AI image-to-video model under the Tongyi Wanxiang series. Supporting single first-frame driving and smooth first-to-last frame transitions, it directly generates up to 30-second videos in up to 1080P/4K cinematic resolution.

Sample Works

Prompt

15 seconds, 16:9, text-to-video only. 6 scenes total, 2.5 seconds each. All transitions are clean hard cuts; no special effects. No reference images, starting frames, or storyboards used. Includes generated natural ambient sound. [Art Style] Original Japanese hand-drawn 2D cooking animation. Fully hand-drawn aesthetic. Retains fluid, clear pencil line art with subtle variations in line weight and the characteristic slight jitter of hand-drawing. Food outlines are fine and warm; outlines for hands and tools are slightly thicker. Uses clean, 2–3 layer full-color cel-style shading in warm, neutral tones, avoiding overly heavy black shadows. Kitchen scenes feature visible brushstrokes, shallow depth of field, and atmospheric perspective; retains moderate analog film grain, with slight color bleeding at cel edges. Strong specular highlights appear only on the ganache and wet dough. Does not imitate existing studios, artists, films, or characters. [Subject & Rules] Subject: Cocoa chocolate bread. Preparation involves making a ganache slab (dark chocolate and cream) to be cooled, kneading cocoa dough, and performing the initial fermentation. Only the hands and forearms of an experienced adult woman are visible. She wears ivory linen cuffs; skin tone is a warm peach-beige; nails are short and unpolished; five fingers per hand; no jewelry. No face, head, torso, full body, or second person appears. Focus is strictly on the food and the cooking process. Each scene contains only one primary action, and the state of ingredients progresses only forward. Never show the finished loaf, baked bread, shaped dough logs, or the chocolate filling. [Scene Composition] Scene 1 (0.00–2.50) - Shot: High-angle close-up, looking straight down at a worn oak work surface; very slow push-in. Start: A block of dark chocolate and a knife with a light oak handle. Action: Left hand holds the chocolate steady; right hand makes four cuts, lifting the knife completely after each stroke. Transformation: A pile of uniformly sized fragments. Texture: Dry, with sharp cut surfaces and fine crumbs; no signs of melting or gloss. Transition: All fragments move to Scene 2. Scene 2 (2.50–5.00) *Texture emphasis - Shot: Low-angle macro view just above the rim of the bowl, moving very slowly to the right. Start: Fragments from Scene 1 begin to melt into hot cream, starting from the edges. Action: A dark brown silicone spatula stirs three times, pivoting from the center with an expanding radius. Transformation: Fragments disappear and merge into a uniform black ganache; the final traces slowly close up. Texture: Thick and splatter-free, with narrow, bright highlights along the edges. Transition: Cream and flour are mixed into the ganache; the mixture is flattened to a 3mm thickness and refrigerated. Scene 3 (5.00–7.50) - Shot: Fixed high-angle (45-degree) medium shot overlooking the work surface. Start: A chilled ganache sheet (22x8cm, 3mm thick) resting on parchment paper. Action: Hands grasp both ends of the paper and lift it once. Transformation: The ganache sheet bends slightly without sagging or breaking. Texture: A cold, resilient solid with a narrow, hard sheen along the edges. Transition: The sheet remains chilled. Scene 4 (7.50–10.00) - Shot: Fixed high-angle close-up, positioned near the rim of a stainless steel bowl. Start: A mixture of uniform brown dry powder and cold dairy in the center of the bowl. Action: A spatula folds the mixture from bottom to top four times. Transformation: Clumps of rough, torn cocoa dough form, with minimal loose flour. Texture: Dark, moist areas with an uneven surface and a few dry flour particles along the edges. Transition: The dough and a piece of butter move to Scene 5. Scene 5 (10.00–12.50) *Texture emphasis - Shot: Fixed low-angle medium shot, level with the work surface. Start: A block of butter at the center of a dough mass. Action: Push forward with both hands and fold inward three times. Transformation: Streaks of butter and cracks disappear, blending into a smooth, elastic, dark cocoa dough; edges thin out until translucent without tearing. Texture: Glossy with elasticity rather than wetness, demonstrating gluten strength. Transition: Place the dough in a bowl to proof until roughly doubled in size. Scene 6 (12.50–15.00) - Shot: High-angle (45°) macro view toward the center of the bowl, moving extremely slowly toward the point of finger pressure. Start: A smooth dome of dough, proofed to roughly double its size. Action: Index finger presses the center and is then fully withdrawn. Transformation: Edges slowly spring back; a shallow fingerprint remains at the center. Texture: Fine fermentation bubbles, soft elasticity, no baked crust. Transition: Ends at this fully proofed state. [Shot] For each scene, change at least two of the following: distance, angle, or height. Only static shots or extremely slow movements are permitted within a scene. No rapid panning, camera shake, rotation, or quick zooming. Points of contact must remain clearly visible at all times. Macro lenses are used only in Scene 2 and Scene 6. [Texture and Movement] Progressive changes: Chocolate chunks → thick emulsion → cooled, firm sheet; dry powder → rough dough → smooth gluten → proofed dough. Hand movements convey weight and inertia; liquids and dough adhere to the laws of viscosity, elasticity, and gravity. No abrupt deformations or reversals of state. [Lighting and Setting] A clean, warm morning kitchen. White quartz countertop and a worn oak workspace. Soft golden key light enters from an upper-left window; faint cool fill light comes from the opposite side; shadows cast on the floor beneath all objects align consistently. Only objects mentioned in the current scene are shown. [Physical Continuity] Maintain consistency in chocolate quantity, ganache batches, cocoa dough mass, hands, cuffs, bowls, the kitchen environment, and lighting sources. Maintain consistency in the quantity, size, and texture of tools and ingredients. Adhere to the actual sequence of cooking steps and physical laws. [Audio] Include only low-register kitchen ambient sounds and contact noises: dry chopping, the sound of a spatula scraping against a moist surface, the rustling of parchment paper, the clinking of bowls, and the sticky sound of dough. No background music, dialogue, narration, or human voices. [Negative Prompts] Text, titles, captions, subtitles, numbers, labels, logos.

Wan 3.0

Wan 3.0

Seedance 2.0

Seedance 2.0

Seedance 2.5

Seedance 2.5

Wan 3.0 Models API Pricing Details

ModelPricing (USD)Our Pricing (USD)Discount
Wan 3.0 Prime Reference-to-Video Spicy$0.068/SECStart from$0.061/SEC-10%
Wan 3.0 Reference-to-Video Spicy$0.05/SECStart from$0.04/SEC-20%
Wan 3.0 Prime Reference-to-Video$0.068/SECStart from$0.061/SEC-10%
Wan 3.0 Reference-to-Video$0.05/SECStart from$0.04/SEC-20%
Wan 3.0 Prime Image-to-Video$0.068/SECStart from$0.061/SEC-10%
Wan 3.0 Prime Text-to-Video$0.068/SECStart from$0.061/SEC-10%
Wan 3.0 Image-to-Video Spicy$0.05/SECStart from$0.04/SEC-20%
Wan 3.0 Prime Image-to-Video Spicy$0.068/SECStart from$0.061/SEC-10%
Wan 3.0 Prime Text-to-Video Spicy$0.068/SECStart from$0.061/SEC-10%
Wan 3.0 Text-to-Video Spicy$0.05/SECStart from$0.04/SEC-20%
Wan 3.0 Text-to-Video$0.05/SECStart from$0.04/SEC-20%
Wan 3.0 Image-to-Video$0.05/SECStart from$0.04/SEC-20%