Qwen Image 3.0 Pro Spicy

aliyun/qwen-image-3-0-pro-global
SpicyText-to-ImageImage-to-Image

Qwen Image 3.0 Pro Spicy is a flagship uncensored image generation model built upon Alibaba's advanced Qwen vision architecture. Operating as a Spicy uncensored model, it completely bypasses standard content filters to unlock unrestricted artistic freedom across photorealistic portraiture, complex visual storytelling, and high-detail stylized concept art. Leveraging Pro-tier enhancements, it delivers ultra-sharp textures, master-level lighting physics, fine typography rendering inside images, and exceptional prompt adherence. Qwen Image 3.0 Pro Spicy is the premier choice for professional designers, digital creators, marketing agencies, and visual artists requiring unconstrained, studio-quality image synthesis.

Read Me

Qwen Image 3.0 Pro Spicy API

Qwen Image 3.0 Pro Spicy is Alibaba's flagship third-generation image generation model, released with the Qwen-Image-3.0 series by the Qwen team on July 21, 2026, and built on a Diffusion Transformer (DiT) architecture, engineered for high information density and professional production. It accepts ultra-long prompts of up to 4,500 tokens, generating complex layouts like newspaper front pages, multi-panel storyboards, academic papers, and detailed UI interfaces in a single pass. It delivers micro-level detail — micro-expressions, pores, hair strands — with near-photographic fidelity, and industry-leading typography rendering legible small text down to 10px across 12 native languages and 20+ fonts. As of August 5, 2026, it ranked fifth globally on Arena.ai's Text-to-Image human-preference arena at 1263±11, the highest-ranked Chinese image generation model (Preliminary).

The model is offered through the iCreat platform as an API using a two-step asynchronous task flow: submit a task to obtain a task_id, then poll the result endpoint until a terminal status. It supports ten output sizes across the 1K and 2K tiers with differentiated per-image pricing: $0.04 per image at 1K and $0.075 at 2K; the output watermark can be turned off via parameter.

Model Positioning

Qwen Image 3.0 Pro targets flagship-tier professional image production. Unlike the standard Qwen Image 3.0, which balances speed, quality, and price, the Pro tier is purpose-built for complex content and fine rendering: long-text posters, dense infographics, multilingual layouts, UI interfaces, knowledge-heavy content, and first drafts of formal commercial work. Alibaba summarizes this generation as "rich content, authentic details, deep knowledge" — Pro is the full expression of that capability set, while Standard is the balanced variant for everyday batch production on the same architecture. On Arena.ai's human-preference Text-to-Image leaderboard, Pro ranks fifth globally and first among Chinese models at 1263±11, making it the go-to choice for complex layouts and small-text rendering.

Core Capabilities

Ultra-Long Complex Instructions

Prompts run up to 4,500 tokens — describe copy, multi-level headings, layout structure, module positions, chart data, and font/color requirements the way you would write a design brief. Newspaper front pages, multi-panel storyboards, and academic pages generate in one pass.

Micro-Level Detail Fidelity

Micro-expressions, pores, and hair strands are faithfully rendered with material and lighting fidelity approaching real photography — suited to formal commercial output of fine products and people.

Industry-Grade Multilingual Typography

Text renders legibly down to 10px across 12 native languages and 20+ fonts; small text stays legible, so text in multilingual posters, packaging, and infographics is usable as generated.

Dense Layouts and Interface Simulation

Supports dense nested layouts; mainstream web, game, and livestream interfaces can be simulated; product feature grids, knowledge cards, and selling-point graphics generate in a single pass.

Unified Generation and Editing

Text-to-image and image-to-image instruction editing share one model: add a reference image to content and edit by text instruction — outfit swaps, background changes, restoration, style transfer — with the flagship rendering capability carrying over to editing.

Pricing

Size Unit Price (USD/Image)
1K: 1024x1024, 1280x720, 720x1280, 1024x768, 768x1024 0.04
2K: 2048x2048, 2560x1440, 1440x2560, 2048x1536, 1536x2048 0.075

Billed per generated image with differentiated 1K/2K pricing. The costUSD field in the response returns the actual cost once the task succeeds.

Application Scenarios

  • Complex commercial posters: first drafts of formal layouts dense with long text, modules, and brand information
  • Multilingual localized marketing: native text rendering for cross-language posters, packaging, and cross-border assets
  • Publishing and education layout: production-grade newspaper pages, exam sheets, academic pages, and knowledge infographics
  • UI and product prototyping: simulated generation of web, game, and livestream interfaces
  • Instruction-based editing: outfit swaps, background changes, restoration, and style transfer

Model Comparison

Same-Series Comparison

Model Positioning Instruction Limit Output Resolution
Qwen Image 3.0 Pro (this model) Flagship: complex layouts and fine rendering 4,500 tokens 1K / 2K
Qwen Image 3.0 (Standard) Balanced: batch daily production, $0.03 flat for 1K/2K 4,500 tokens 1K / 2K
Qwen Image 2.0 Pro Previous flagship: typography and realistic texture ~1,000 tokens 512×512 to 2048×2048
Qwen Image 1.0 First-gen foundation, 20B, Apache 2.0 open weights ~1,000 tokens —

Cross-Model Comparison

Model Developer Supported Resolutions Billing
Qwen Image 3.0 Pro Alibaba 1K / 2K Per image, differentiated 1K/2K
Nano Banana Pro Google 1K / 2K / 4K Per image ($0.134 at 1K/2K, $0.24 at 4K)
Nano Banana 2 Google 0.5K / 1K / 2K / 4K Per image ($0.067 at 1K)
GPT Image 2 OpenAI Up to 4K Per image by quality tier ($0.006–$0.211)

Same-series data comes from Alibaba's official releases; cross-model prices are the vendors' official API prices; this model's pricing and specifications follow the iCreat channel documentation.

Why Choose Qwen Image 3.0 Pro?

  • 4,500-token instruction window plus dense nested layouts — complex pages generate in one pass, no stitching
  • 10px small text across 12 languages and 20+ fonts: text in multilingual commercial assets is usable as generated, skipping post-layout
  • Micro-expression, pore, and hair-strand level detail raises the fidelity floor for formal commercial output
  • Fifth globally and first among Chinese models on Arena.ai's human-preference arena (1263±11, Preliminary) — proven on complex-layout tasks
  • Unified generation and instruction editing, with flagship rendering applied to editing scenarios

Specifications

Item Description
Base URL https://api.icreat.ai
Submit endpoint POST /v1/task/submit/aliyun/qwen-image-3-0-pro-global
Query result endpoint POST /v1/task/result
Model ID aliyun/qwen-image-3-0-pro-global
Authentication Authorization: Bearer header
Call pattern Two-step asynchronous task (submit → poll)
input.messages Required, array of conversation messages carrying the text prompt and optional reference image
messages[].role Required, user
messages[].content Required, array of content parts: each item is {"text": "..."} or {"image": "..."}, combined as needed
content[].text Text prompt or editing instruction
content[].image Reference image URL for image-to-image or instruction editing
parameters.size Required, output image size, e.g. "1024*1024"; ten values across the 1K/2K tiers
parameters.watermark Optional, whether to add a watermark to the output image; default false
Task status SUBMITTED / SUCCEEDED / FAILED
Output resource type is Image, includes url and download_url
Cost field costUSD, present only on SUCCEEDED

Architecture

Qwen Image 3.0 Pro is built on a Diffusion Transformer (DiT) architecture as the third-generation flagship of the Qwen image series. It carries the unified generation-editing design: text-to-image and image-to-image instruction editing share the same weights, so typography, micro-level detail, and world knowledge accumulated on the generation side carry directly over to the editing side. The instruction window extends to 4,500 tokens, specifically strengthened for high-information-density layouts, supporting nested layouts and composed multi-type visuals in a single generation. Output covers ten sizes across the 1K and 2K tiers with differentiated pricing.

Notes

  • Submit and query must be chained with the same task_id
  • While processing, result is []; keep polling, with a suggested interval of 2–5 seconds
  • FAILED is terminal; check request parameters and reference media URLs before resubmitting
  • input and parameters are sibling top-level fields; do not nest them
  • Each item in messages[].content must be either {"text": "..."} or {"image": "..."}, combined as needed
  • Reference image URLs must be publicly accessible
  • parameters.size is required, written as width*height (e.g. 1024*1024), not joined with x
  • watermark defaults to off; pass true explicitly when a watermark is needed
  • costUSD is returned only when the task succeeds; failed tasks do not produce a cost field

FAQ

How do I get started with the Qwen Image 3.0 Pro API?​

Register on the iCreat platform and obtain an API Key from the console, send a generation request to the submit endpoint, then poll the result endpoint with the returned task_id. All requests are authenticated with the Authorization: Bearer header. Keep your API Key safe and never expose it in client code or public repositories.

How is the cost calculated?​

Billed per generated image with differentiated pricing: $0.04 per image at 1K and $0.075 per image at 2K. For example, generating 10 1K images costs $0.40, while 10 2K images cost $0.75. The costUSD field in the response gives the actual cost once the task succeeds.

Does it support generation or editing?​

Both. With only {"text": "..."} in content, it generates from the prompt; adding an {"image": "..."} reference enables instruction-based editing (outfit swaps, background changes, restoration, style transfer, and more). Generation and editing share the same model and endpoint.

What output sizes are supported?​

parameters.size supports the 1K tier (1024*1024, 1280*720, 720*1280, 1024*768, 768*1024) and the 2K tier (2048*2048, 2560*1440, 1440*2560, 2048*1536, 1536*2048) — ten values in total, with differentiated pricing between tiers. Pick the right size for each delivery platform.

How do I query the task result?​

Send a POST request to https://api.icreat.ai/v1/task/result with the task_id in the body. A status of SUBMITTED means processing, with result as [] — keep polling. When status is SUCCEEDED, result returns the image resource array; read url or download_url to retrieve the image.

What if the task fails?​

FAILED is a terminal status. Check in order: whether messages contains at least one text item, whether reference image URLs are publicly accessible, whether size is one of the ten values written as width*height, and whether input and parameters are sibling top-level fields. Fix any issue and resubmit. The cost field is returned only when the task succeeds.