
Qwen Image 3.0 Pro Spicy
Qwen Image 3.0 Pro Spicy is a flagship uncensored image generation model built upon Alibaba's advanced Qwen vision architecture. Operating as a Spicy uncensored model, it completely bypasses standard content filters to unlock unrestricted artistic freedom across photorealistic portraiture, complex visual storytelling, and high-detail stylized concept art. Leveraging Pro-tier enhancements, it delivers ultra-sharp textures, master-level lighting physics, fine typography rendering inside images, and exceptional prompt adherence. Qwen Image 3.0 Pro Spicy is the premier choice for professional designers, digital creators, marketing agencies, and visual artists requiring unconstrained, studio-quality image synthesis.
Read Me
Qwen Image 3.0 Pro Spicy API
Qwen Image 3.0 Pro Spicy is Alibaba's flagship third-generation image generation model, released with the Qwen-Image-3.0 series by the Qwen team on July 21, 2026, and built on a Diffusion Transformer (DiT) architecture, engineered for high information density and professional production. It accepts ultra-long prompts of up to 4,500 tokens, generating complex layouts like newspaper front pages, multi-panel storyboards, academic papers, and detailed UI interfaces in a single pass. It delivers micro-level detail — micro-expressions, pores, hair strands — with near-photographic fidelity, and industry-leading typography rendering legible small text down to 10px across 12 native languages and 20+ fonts. As of August 5, 2026, it ranked fifth globally on Arena.ai's Text-to-Image human-preference arena at 1263±11, the highest-ranked Chinese image generation model (Preliminary).
The model is offered through the iCreat platform as an API using a two-step asynchronous task flow: submit a task to obtain a task_id, then poll the result endpoint until a terminal status. It supports ten output sizes across the 1K and 2K tiers with differentiated per-image pricing: $0.04 per image at 1K and $0.075 at 2K; the output watermark can be turned off via parameter.
Model Positioning
Qwen Image 3.0 Pro targets flagship-tier professional image production. Unlike the standard Qwen Image 3.0, which balances speed, quality, and price, the Pro tier is purpose-built for complex content and fine rendering: long-text posters, dense infographics, multilingual layouts, UI interfaces, knowledge-heavy content, and first drafts of formal commercial work. Alibaba summarizes this generation as "rich content, authentic details, deep knowledge" — Pro is the full expression of that capability set, while Standard is the balanced variant for everyday batch production on the same architecture. On Arena.ai's human-preference Text-to-Image leaderboard, Pro ranks fifth globally and first among Chinese models at 1263±11, making it the go-to choice for complex layouts and small-text rendering.
Core Capabilities
Ultra-Long Complex Instructions
Prompts run up to 4,500 tokens — describe copy, multi-level headings, layout structure, module positions, chart data, and font/color requirements the way you would write a design brief. Newspaper front pages, multi-panel storyboards, and academic pages generate in one pass.
Micro-Level Detail Fidelity
Micro-expressions, pores, and hair strands are faithfully rendered with material and lighting fidelity approaching real photography — suited to formal commercial output of fine products and people.
Industry-Grade Multilingual Typography
Text renders legibly down to 10px across 12 native languages and 20+ fonts; small text stays legible, so text in multilingual posters, packaging, and infographics is usable as generated.
Dense Layouts and Interface Simulation
Supports dense nested layouts; mainstream web, game, and livestream interfaces can be simulated; product feature grids, knowledge cards, and selling-point graphics generate in a single pass.
Unified Generation and Editing
Text-to-image and image-to-image instruction editing share one model: add a reference image to content and edit by text instruction — outfit swaps, background changes, restoration, style transfer — with the flagship rendering capability carrying over to editing.
Pricing
| Size | Unit Price (USD/Image) |
|---|---|
1K: 1024x1024, 1280x720, 720x1280, 1024x768, 768x1024 |
0.04 |
2K: 2048x2048, 2560x1440, 1440x2560, 2048x1536, 1536x2048 |
0.075 |
Billed per generated image with differentiated 1K/2K pricing. The costUSD field in the response returns the actual cost once the task succeeds.
Application Scenarios
- Complex commercial posters: first drafts of formal layouts dense with long text, modules, and brand information
- Multilingual localized marketing: native text rendering for cross-language posters, packaging, and cross-border assets
- Publishing and education layout: production-grade newspaper pages, exam sheets, academic pages, and knowledge infographics
- UI and product prototyping: simulated generation of web, game, and livestream interfaces
- Instruction-based editing: outfit swaps, background changes, restoration, and style transfer
Model Comparison
Same-Series Comparison
| Model | Positioning | Instruction Limit | Output Resolution |
|---|---|---|---|
| Qwen Image 3.0 Pro (this model) | Flagship: complex layouts and fine rendering | 4,500 tokens | 1K / 2K |
| Qwen Image 3.0 (Standard) | Balanced: batch daily production, $0.03 flat for 1K/2K | 4,500 tokens | 1K / 2K |
| Qwen Image 2.0 Pro | Previous flagship: typography and realistic texture | ~1,000 tokens | 512×512 to 2048×2048 |
| Qwen Image 1.0 | First-gen foundation, 20B, Apache 2.0 open weights | ~1,000 tokens | — |
Cross-Model Comparison
| Model | Developer | Supported Resolutions | Billing |
|---|---|---|---|
| Qwen Image 3.0 Pro | Alibaba | 1K / 2K | Per image, differentiated 1K/2K |
| Nano Banana Pro | 1K / 2K / 4K | Per image ($0.134 at 1K/2K, $0.24 at 4K) | |
| Nano Banana 2 | 0.5K / 1K / 2K / 4K | Per image ($0.067 at 1K) | |
| GPT Image 2 | OpenAI | Up to 4K | Per image by quality tier ($0.006–$0.211) |
Same-series data comes from Alibaba's official releases; cross-model prices are the vendors' official API prices; this model's pricing and specifications follow the iCreat channel documentation.
Why Choose Qwen Image 3.0 Pro?
- 4,500-token instruction window plus dense nested layouts — complex pages generate in one pass, no stitching
- 10px small text across 12 languages and 20+ fonts: text in multilingual commercial assets is usable as generated, skipping post-layout
- Micro-expression, pore, and hair-strand level detail raises the fidelity floor for formal commercial output
- Fifth globally and first among Chinese models on Arena.ai's human-preference arena (1263±11, Preliminary) — proven on complex-layout tasks
- Unified generation and instruction editing, with flagship rendering applied to editing scenarios
Specifications
| Item | Description |
|---|---|
| Base URL | https://api.icreat.ai |
| Submit endpoint | POST /v1/task/submit/aliyun/qwen-image-3-0-pro-global |
| Query result endpoint | POST /v1/task/result |
| Model ID | aliyun/qwen-image-3-0-pro-global |
| Authentication | Authorization: Bearer header |
| Call pattern | Two-step asynchronous task (submit → poll) |
input.messages |
Required, array of conversation messages carrying the text prompt and optional reference image |
messages[].role |
Required, user |
messages[].content |
Required, array of content parts: each item is {"text": "..."} or {"image": "..."}, combined as needed |
content[].text |
Text prompt or editing instruction |
content[].image |
Reference image URL for image-to-image or instruction editing |
parameters.size |
Required, output image size, e.g. "1024*1024"; ten values across the 1K/2K tiers |
parameters.watermark |
Optional, whether to add a watermark to the output image; default false |
| Task status | SUBMITTED / SUCCEEDED / FAILED |
| Output resource | type is Image, includes url and download_url |
| Cost field | costUSD, present only on SUCCEEDED |
Architecture
Qwen Image 3.0 Pro is built on a Diffusion Transformer (DiT) architecture as the third-generation flagship of the Qwen image series. It carries the unified generation-editing design: text-to-image and image-to-image instruction editing share the same weights, so typography, micro-level detail, and world knowledge accumulated on the generation side carry directly over to the editing side. The instruction window extends to 4,500 tokens, specifically strengthened for high-information-density layouts, supporting nested layouts and composed multi-type visuals in a single generation. Output covers ten sizes across the 1K and 2K tiers with differentiated pricing.
Notes
- Submit and query must be chained with the same
task_id - While processing,
resultis[]; keep polling, with a suggested interval of 2–5 seconds FAILEDis terminal; check request parameters and reference media URLs before resubmittinginputandparametersare sibling top-level fields; do not nest them- Each item in
messages[].contentmust be either{"text": "..."}or{"image": "..."}, combined as needed - Reference image URLs must be publicly accessible
parameters.sizeis required, written aswidth*height(e.g.1024*1024), not joined withxwatermarkdefaults to off; passtrueexplicitly when a watermark is neededcostUSDis returned only when the task succeeds; failed tasks do not produce a cost field
FAQ
How do I get started with the Qwen Image 3.0 Pro API?
Register on the iCreat platform and obtain an API Key from the console, send a generation request to the submit endpoint, then poll the result endpoint with the returned task_id. All requests are authenticated with the Authorization: Bearer header. Keep your API Key safe and never expose it in client code or public repositories.
How is the cost calculated?
Billed per generated image with differentiated pricing: $0.04 per image at 1K and $0.075 per image at 2K. For example, generating 10 1K images costs $0.40, while 10 2K images cost $0.75. The costUSD field in the response gives the actual cost once the task succeeds.
Does it support generation or editing?
Both. With only {"text": "..."} in content, it generates from the prompt; adding an {"image": "..."} reference enables instruction-based editing (outfit swaps, background changes, restoration, style transfer, and more). Generation and editing share the same model and endpoint.
What output sizes are supported?
parameters.size supports the 1K tier (1024*1024, 1280*720, 720*1280, 1024*768, 768*1024) and the 2K tier (2048*2048, 2560*1440, 1440*2560, 2048*1536, 1536*2048) — ten values in total, with differentiated pricing between tiers. Pick the right size for each delivery platform.
How do I query the task result?
Send a POST request to https://api.icreat.ai/v1/task/result with the task_id in the body. A status of SUBMITTED means processing, with result as [] — keep polling. When status is SUCCEEDED, result returns the image resource array; read url or download_url to retrieve the image.
What if the task fails?
FAILED is a terminal status. Check in order: whether messages contains at least one text item, whether reference image URLs are publicly accessible, whether size is one of the ten values written as width*height, and whether input and parameters are sibling top-level fields. Fix any issue and resubmit. The cost field is returned only when the task succeeds.



