Qwen Image 3.0 Spicy

aliyun/qwen-image-3-0-global
SpicyText-to-ImageImage-to-Image

Qwen Image 3.0 Spicy is a flagship uncensored image generation model built upon Alibaba's advanced Qwen vision architecture. Operating as a Spicy uncensored model, it completely bypasses standard content filters to unlock maximum creative freedom across photorealistic portraiture, complex visual storytelling, and high-detail stylized concept art without content restrictions. Leveraging enhanced rendering capacities, it delivers sharp textures, realistic lighting physics, and exceptional prompt adherence. Qwen Image 3.0 Spicy is the premier choice for professional designers, digital creators, marketing agencies, and visual artists requiring unconstrained, high-quality image synthesis.

Read Me

Qwen Image 3.0 Spicy API

Qwen Image 3.0 Spicy is Alibaba's third-generation image generation and editing foundation model, released by the Qwen team on July 21, 2026, and built on a Diffusion Transformer (DiT) architecture. Its keyword is "Real" — Rich Content, Authentic Details, Deep Knowledge. It accepts instructions up to 4,500 tokens (roughly 4.5× the previous generation) and can generate a 3×3 grid of nine complex infographics in a single pass; text rendering reaches 10px precision across 12 native languages and 20+ fonts. Text-to-image (T2I) and image-to-image instruction editing (I2I) are unified in one model, with typography and knowledge capabilities carrying over naturally from generation to editing.

The model is offered through the iCreat platform as an API using a two-step asynchronous task flow: submit a task to obtain a task_id, then poll the result endpoint until a terminal status. It supports ten output sizes across the 1K and 2K tiers, all billed at $0.03 per image; the output watermark can be turned off via parameter.

Model Positioning

Qwen Image 3.0 targets image production where the output must be deployable, not just decorative. It pushes image generation from "produce a pretty picture" to "produce a usable artifact": newspaper pages, exam sheets, film storyboards, nested UI mockups, and knowledge infographics with formulas can all be generated ready-to-use in one pass. Within the Qwen image series it is the third-generation foundation — the previous flagship, Qwen Image 2.0 Pro, ranked fifth overall on Alibaba's own Qwen-Image-Bench (behind GPT Image 2 and Google's Nano Banana models), and generation 3.0 upgrades with a 4.5× instruction window plus deeper semantic juxtaposition and spatial control. Typical needs include localized advertising, product UI prototyping, e-commerce graphics, and creative publishing at scale.

Core Capabilities

Long Instruction Window

Prompts run up to 4,500 tokens — roughly 4.5× the previous generation. Describe layout structure, text content, visual style, and typography the way you would brief a designer, instead of compressing everything into one sentence; fewer regenerations, less manual fixing.

Native Multilingual Text Rendering

Text renders legibly down to 10px across 12 native languages and 20+ fonts, supporting complex multi-column layouts, UI designs, and graphic documents — sharply cutting the cost of producing readable multilingual posters and storyboard material.

Rich Content and Deep Knowledge

Generates a 3×3 grid of nine complex infographics in one pass; understands complex instructions and nested visual logic, composing webpages, software interfaces, chat windows, and posters within a single image from one instruction.

Unified Generation and Editing

Text-to-image (T2I) and image-to-image instruction editing (I2I) share one model: typography, detail fidelity, and world knowledge accumulated in generation carry over naturally to editing — antique-painting restoration, panorama generation, hand-drawn sketch to PPT, and multi-shot storyboards complete in one step.

Authentic Detail Fidelity

Hyper-realistic detail with convincing skin, paper, and material textures; scenes GPT-Image-2 is known for, such as livestream pages, are covered as well.

Pricing

Size Unit Price (USD/Image)
1K: 1024x1024, 1280x720, 720x1280, 1024x768, 768x1024 0.03
2K: 2048x2048, 2560x1440, 1440x2560, 2048x1536, 1536x2048 0.03

Billed per generated image; 1K and 2K are the same price. The costUSD field in the response returns the actual cost once the task succeeds.

Application Scenarios

  • Localized advertising: batch production of multilingual posters and promo images with readable text
  • Product UI prototyping: fast generation and iteration of interfaces, icons, and operation graphics
  • E-commerce graphics: product pages, livestream layouts, and detail-page material at scale
  • Creative publishing: newspaper pages, exam sheets, paper layouts, hand-drawn sketches to finished art
  • Complex editing: restoration, panoramas, and multi-shot storyboards via instruction editing

Model Comparison

Same-Series Comparison

Model Released Positioning Text Rendering Instruction Limit
Qwen Image 3.0 (this model) July 2026 Third-gen foundation, productivity-focused 10px / 12 languages / 20+ fonts 4,500 tokens
Qwen Image 2.0 Pro 2025 Previous flagship Strong ~1,000 tokens
Qwen Image 1.0 August 2025 First-gen foundation, 20B parameters, Apache 2.0 open weights Strong ~1,000 tokens

Cross-Model Comparison

Model Developer Supported Resolutions Billing
Qwen Image 3.0 Alibaba 1K / 2K Per image
Nano Banana 2 Lite Google 1K Per image
GPT Image 2 OpenAI Up to 4K Per image by quality tier
Seedream 4.0 ByteDance Up to 4K Per image

Same-series data comes from Alibaba's official releases; this model's pricing and specifications follow the iCreat channel documentation.

Why Choose Qwen Image 3.0?

  • 4,500-token instruction window — brief the model like a designer instead of compressing prompts
  • 10px text rendering across 12 languages and 20+ fonts: text in multilingual posters and graphics is usable as generated
  • 3×3 infographic grids in a single pass — no stitching complex layouts together
  • Unified generation and instruction editing with zero switching cost for restoration, retouching, and storyboards
  • 1K and 2K at the same price — high-resolution delivery without a premium

Specifications

Item Description
Base URL https://api.icreat.ai
Submit endpoint POST /v1/task/submit/aliyun/qwen-image-3-0-global
Query result endpoint POST /v1/task/result
Model ID aliyun/qwen-image-3-0-global
Authentication Authorization: Bearer header
Call pattern Two-step asynchronous task (submit → poll)
input.messages Required, array of conversation messages carrying the text prompt and optional reference image
messages[].role Required, user
messages[].content Required, array of content parts: each item is {"text": "..."} or {"image": "..."}, combined as needed
content[].text Text prompt or editing instruction
content[].image Reference image URL for image-to-image or instruction editing
parameters.size Required, output image size, e.g. "1024*1024"; ten values across the 1K/2K tiers
parameters.watermark Optional, whether to add a watermark to the output image; default false
Task status SUBMITTED / SUCCEEDED / FAILED
Output resource type is Image, includes url and download_url
Cost field costUSD, present only on SUCCEEDED

Architecture

Qwen Image 3.0 is built on a Diffusion Transformer (DiT) architecture as the third-generation foundation of the Qwen image generation and editing series. It adopts a unified generation-editing design: text-to-image and image-to-image instruction editing share the same weights, so typography, detail fidelity, and world knowledge accumulated on the generation side carry over naturally to the editing side. The instruction window extends to 4,500 tokens, with enhanced semantic juxtaposition and spatial control supporting 3×3 grid layouts and composed multi-type visuals in a single generation. Output covers ten sizes across the 1K and 2K tiers.

Notes

  • Submit and query must be chained with the same task_id
  • While processing, result is []; keep polling, with a suggested interval of 2–5 seconds
  • FAILED is terminal; check request parameters and reference media URLs before resubmitting
  • input and parameters are sibling top-level fields; do not nest them
  • Each item in messages[].content must be either {"text": "..."} or {"image": "..."}, combined as needed
  • Reference image URLs must be publicly accessible
  • parameters.size is required, written as width*height (e.g. 1024*1024), not joined with x
  • watermark defaults to off; pass true explicitly when a watermark is needed
  • costUSD is returned only when the task succeeds; failed tasks do not produce a cost field

FAQ

How do I get started with the Qwen Image 3.0 API?​

Register on the iCreat platform and obtain an API Key from the console, send a generation request to the submit endpoint, then poll the result endpoint with the returned task_id. All requests are authenticated with the Authorization: Bearer header. Keep your API Key safe and never expose it in client code or public repositories.

How is the cost calculated?​

Billed per generated image: $0.03 per image for both 1K and 2K. For example, generating 10 2K images costs $0.30. The costUSD field in the response gives the actual cost once the task succeeds.

Does it support generation or editing?​

Both. With only {"text": "..."} in content, it generates from the prompt; adding an {"image": "..."} reference enables instruction-based editing (outfit swaps, background changes, restoration, and more). Generation and editing share the same model and endpoint.

What output sizes are supported?​

parameters.size supports the 1K tier (1024*1024, 1280*720, 720*1280, 1024*768, 768*1024) and the 2K tier (2048*2048, 2560*1440, 1440*2560, 2048*1536, 1536*2048) — ten values in total, both tiers at the same price. Pick the right size for each delivery platform.

How do I query the task result?​

Send a POST request to https://api.icreat.ai/v1/task/result with the task_id in the body. A status of SUBMITTED means processing, with result as [] — keep polling. When status is SUCCEEDED, result returns the image resource array; read url or download_url to retrieve the image.

What if the task fails?​

FAILED is a terminal status. Check in order: whether messages contains at least one text item, whether reference image URLs are publicly accessible, whether size is one of the ten values written as width*height, and whether input and parameters are sibling top-level fields. Fix any issue and resubmit. The cost field is returned only when the task succeeds.