Qwen 3.8 Max

qwen3.8-max
OfficialLLM

Qwen3.8-Max is Alibaba’s flagship model in the Qwen3.8 series, designed for agentic, text-based workflows. It excels at coding, debugging, office automation, productivity tasks, tool use, and long-horizon autonomous execution. With a 1-million-token context window and support for outputs of up to 64K tokens, it is ideal for processing large documents, repository-scale coding, multi-step planning, structured content generation, and complex workflows requiring sustained reasoning across hundreds or even thousands of steps.

Read Me

Qwen 3.8 Max API

Qwen 3.8 Max is Alibaba's flagship large language model officially released on August 3, 2026, built on the Qwen 3.5 architecture with a sparse MoE architecture and hybrid attention mechanism. It has 2.4 trillion (2.4T) total parameters with 95B activated parameters, supporting a 1-million-token (1M) context length. It achieves comprehensive improvements in coding, office productivity, research, and long-horizon tasks, capable of independent autonomous programming for up to about 16 days. In Arena third-party rankings, it placed second in vision and fourth in coding. Model weights were open-sourced on August 12, 2026, under Apache 2.0.

On iCreat, developers call qwen3.8-max through an OpenAI Chat Completions-compatible API. Pricing: input $2.00/million tokens, output $7.00/million tokens, cached read $0.25/million tokens, cached write $2.50/million tokens.

Model Positioning

Qwen 3.8 Max is the most powerful model in the Qwen family to date and the first Max-tier model to be open-sourced. With 2.4T total / 95B activated parameters, it comprehensively surpasses its predecessor in coding, office productivity, long-horizon autonomous tasks, and multimodal agents. It suits production-grade tasks requiring top-tier intelligence: long-horizon coding, complex research, enterprise compliance analysis. Overseas pricing is only about 40% (input) and 24% (output) of Claude Opus 5, offering significant value.

Core Capabilities

Long-horizon autonomous coding

Qwen 3.8 Max can start from an empty folder and independently deliver real projects spanning multiple days — writing code, running tests, and fixing bugs without human intervention. In testing, it autonomously coded for about 16 days, building the self-evolving agent framework oh-my-cli; in another case, it independently designed a GCD/RSA cryptographic hardware accelerator, optimizing from 8,298 gates to 678.

1M ultra-long context

Supports a 1-million-token context window (native 262,144 tokens, expandable to 1M). Entire code repositories, massive logs, or multiple long documents can be input at once for project-level refactoring, root-cause analysis, and clause-conflict detection.

Native multimodal

Supports text and visual understanding (native multimodal), ranking second on the Vision Arena leaderboard with strong visual reasoning.

Thinking mode and reasoning

Enable thinking mode via thinking: {"type": "enabled"} and control reasoning depth with reasoning_effort (low/medium/high) — save latency and cost on simple tasks, get deep reasoning on complex ones.

Function calling and agentic capability

Supports function calling and structured output, compatible with OpenAI and Anthropic API protocols, integrable into mainstream agent frameworks. In office scenarios, it delivers production-grade results as enterprise compliance lawyers, structural engineers, and sports data analysts.

Cache optimization

Cached reads billed at $0.25/million tokens (one-eighth of the standard $2.00 input price); cached writes at $2.50/million tokens, significantly reducing costs for repeated-context workloads.

Pricing

Token Type Price
Input $2.00/million tokens
Output $7.00/million tokens
Cached Read $0.25/million tokens
Cached Write $2.50/million tokens

Total cost = input tokens × $2.00/million + output tokens × $7.00/million (cached reads at $0.25/million, cached writes at $2.50/million).

Note: Alibaba Cloud Bailian domestic pricing is ¥12/million input, ¥36/million output, implicit cache hit ¥1.5/million (CNY); the USD prices above are iCreat platform prices. Official overseas pricing is $2/million input, $6/million output, cache hit $0.25/million; iCreat's output price of $7 is slightly above the official $6, at the iCreat platform rate.

Use Cases

  • Long-horizon autonomous coding: start from an empty folder and independently deliver multi-day projects — write code, run tests, fix bugs
  • Coding and development: multi-file refactoring, repository-level bug hunting, long-context code review
  • Office and productivity: enterprise compliance analysis, structural engineering calculations, sports data analysis — production-grade deliverables
  • Research and long-horizon tasks: research paper reproduction (~5 days / 125 hours of independent work), cryptographic hardware accelerator design
  • Visual reasoning: native multimodal visual understanding for mixed text-image analysis
  • Cost-sensitive high-frequency calls: cached reads at only $0.25/million tokens — ideal for RAG and multi-turn conversations with shared system prompts

Model Comparison

Qwen 3.8 Max vs. Qwen 3.7 Max

Dimension Qwen 3.8 Max Qwen 3.7 Max
Release date August 3, 2026 May 21, 2026
Positioning Flagship (first Max-tier open-source) Previous flagship
Architecture Sparse MoE + hybrid attention MoE
Total parameters 2.4T (2.4 trillion)
Activated parameters 95B
Context window 1M tokens (native 262K, expandable to 1M) 1M tokens
Input modalities Text + vision (native multimodal) Text
Long-horizon coding ~16 days autonomous coding
Vision Arena 2nd
CodeArena 4th
Official input price $2.00/million tokens $1.25/million tokens
Official output price $7.00/million tokens $3.75/million tokens
Open-source license Apache 2.0 (open-sourced 8/12) Closed
Best for Long-horizon coding, research, multimodal agents General flagship tasks

Qwen 3.8 Max vs. Qwen 3.7 Plus and Qwen 3.5 Omni Plus

Dimension Qwen 3.8 Max Qwen 3.7 Plus Qwen 3.5 Omni Plus
Release date August 3, 2026 June 3, 2026 2026 (earlier)
Positioning Flagship High value Omni-modal
Context window 1M tokens 1M tokens
Input modalities Text + vision Text + image + PDF Text + image + video + audio
Output modalities Text Text Text + audio
Official input price $2.00/million tokens $0.32/million tokens $7.25/million tokens
Official output price $7.00/million tokens $1.28/million tokens $5.50/million tokens
Open-source license Apache 2.0
Best for Long-horizon coding, research, top-tier intelligence Cost-effective daily tasks Omni-modal input/output

Why Choose Qwen 3.8 Max?

  • 2.4T flagship: the most powerful Qwen model to date — 4th in CodeArena, 2nd in Vision Arena
  • Long-horizon autonomous coding: starts from an empty folder and independently delivers ~16-day cross-day projects without human intervention
  • 1M context: supports entire code repositories and massive long documents in a single input
  • Native multimodal: text + visual understanding, visual reasoning ranked 2nd
  • Apache 2.0 open source: first Max-tier open weights — self-deployable, commercially unrestricted
  • High value: overseas input price only ~40% of Claude Opus 5; cached reads at only $0.25/million

Specifications

Category Description
Model name Qwen 3.8 Max
Developer Alibaba (Qwen / Tongyi Qianwen)
Model ID qwen3.8-max
Release date August 3, 2026 (preview July 19)
Model type Flagship LLM (native multimodal)
Architecture Sparse MoE + hybrid attention mechanism
Total parameters 2.4T (2.4 trillion)
Activated parameters 95B
Context window 1M tokens (native 262,144, expandable to 1M)
Input modalities Text + vision (native multimodal)
Output modalities Text
Reasoning capability Supported (thinking.type control, reasoning_effort adjustable)
Streaming output Supported (SSE)
Context caching Supported (cached read $0.25/million, cached write $2.50/million)
Function calling Supported
Structured output Supported
Open-source license Apache 2.0 (open-sourced 2026-08-12, alongside Qwen3.8-27B)
iCreat API capability Compatible with OpenAI Chat Completions and Anthropic Messages API
Primary tasks Long-horizon coding, research, office, multimodal agents, high-throughput workloads

Architecture

Qwen 3.8 Max is built on the Qwen 3.5 architecture with a sparse MoE architecture and hybrid attention mechanism. Total parameters were expanded to 2.4 trillion (2.4T), with 95B activated per token. The model achieves inference optimization through software-hardware co-design (T-Head chips), supporting long-horizon autonomous coding from an empty folder.

A preview version was released on July 19, 2026; the official version launched on August 3; weights were open-sourced on August 12 under Apache 2.0 (Qwen3.8-2.4T-A95B), alongside the Qwen3.8-27B model — the first Max-tier Qwen model to be open-sourced. The model is compatible with OpenAI and Anthropic APIs and supports integration into mainstream agent frameworks.

Notes

  • The 1M-token context is a shared request space; reserve room for system instructions, conversation state, and the final answer
  • When thinking mode is enabled, responses include reasoning token consumption; for latency- and cost-sensitive high-throughput scenarios, use reasoning_effort: "low"
  • Cached-read discount applies only when the input prefix hits the model cache; cached writes are billed at $2.50/million tokens
  • iCreat platform output price of $7/million is slightly above Alibaba's official overseas price of $6/million; the iCreat platform rate applies
  • Agent loops that repeatedly resend large context blocks should leverage caching: cached reads are only $0.25/million tokens
  • Long-horizon autonomous coding tasks consume large numbers of tokens — estimate cost in advance

FAQ

What is the difference between Qwen 3.8 Max and Qwen 3.7 Max?

Qwen 3.8 Max (released August 3, 2026) is the flagship model (2.4T total / 95B activated), supporting long-horizon autonomous coding, native multimodal, 1M context, Apache 2.0 open source. Qwen 3.7 Max (released May 21) is the previous flagship (1M context, closed). 3.8 Max comprehensively surpasses it in coding, vision, and long-horizon tasks.

What is Qwen 3.8 Max's context window?

1M tokens — native 262,144, expandable to 1M. Suitable for entire code repositories and massive long documents in a single input.

How do I enable thinking mode?

Set "thinking": {"type": "enabled"} in the request body, optionally adjusting reasoning_effort (low/medium/high). Lower or disable thinking on simple tasks to save latency and cost; raise the tier for complex tasks.

What is the difference between cached read and cached write?

Cached read (cache hit) is the price for reading when the input prefix hits the model cache: $0.25/million tokens (one-eighth of the standard $2.00 input). Cached write is the cost of writing context to the cache: $2.50/million tokens — subsequent requests that hit the cache enjoy the read discount.

Can I self-host?

Yes. Model weights were open-sourced on August 12, 2026, under Apache 2.0 on Hugging Face / ModelScope (Qwen3.8-2.4T-A95B), alongside Qwen3.8-27B, commercially unrestricted. Deployable via vLLM and similar inference frameworks.

Can I call it directly with the OpenAI SDK?

Yes. Set base_url to https://api.icreat.ai/llm/openai/v1 and api_key to your iCreat API Key — no extra adaptation needed. The model also supports the Anthropic Messages API protocol.