DeepSeek

Gain production-grade API access to DeepSeek’s flagship DeepSeek V4 model family. Setting industry standards for open-weights performance, the DeepSeek V4 suite features DeepSeek-V4-Pro (Flagship Reasoning) and DeepSeek-V4-Flash (High-Throughput Speedster).

All Models

Deepseek V4.1 Flash
NEW
OfficialLLM

Deepseek V4.1 Flash

DeepSeek V4.1 Flash is DeepSeek's open-weights, ultra-efficient vision-language Mixture-of-Experts (MoE) foundation model. Built on a novel Causal Encoder-Decoder (CED) architecture, it distributes 552B total parameters while activating only 8B parameters during prefill and 16B parameters during decode. Featuring native visual-text joint embedding, Compressed Sparse Attention, Engram memory, and a 4x KV-cache memory reduction, it effortlessly processes up to a 1M-token context window at high throughput. Engineered for extreme cost efficiency and frontier agentic coding, DeepSeek V4.1 Flash powers autonomous software engineering, multimodal document RAG, UI browser automation, and high-concurrency API integrations.

Deepseek V4 Pro 0813
OfficialLLM

Deepseek V4 Pro 0813

DeepSeek-V4-Pro-0813 is a flagship MoE model released on August 13, 2026, serving as the official GA release built for complex logical reasoning, code construction, and AI Agent collaboration. Architecture & Performance: 1.6T parameters (49B active), built-in DSpark speculative decoding, with significantly enhanced inference throughput. Compatible with OpenAI/Anthropic APIs, deeply integrated with DeepSeek Harness, and delivers outstanding performance on Terminal Bench.

Deepseek V4 Flash 0731
OfficialLLM

Deepseek V4 Flash 0731

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts (MoE) model from DeepSeek, featuring 284B total parameters with 13B activated per token and a 1-million-token context window. Designed for fast inference and high-throughput workloads, it delivers strong reasoning and coding capabilities while maintaining excellent cost efficiency. Its hybrid attention architecture enables efficient long-context processing. The model supports **high** and **xhigh** reasoning levels, with **xhigh** representing the maximum reasoning effort. It is ideal for coding assistants, conversational systems, and agentic workflows where responsiveness, scalability, and cost efficiency are essential.

DeepSeek Models API Pricing Details

ModelPricing (USD)Our Pricing (USD)Discount
Deepseek V4.1 Flash$0.3/M TokensStart from$0.24/M Tokens-20%
Deepseek V4 Pro 0813$1.32/M TokensStart from$0.792/M Tokens-40%
Deepseek V4 Flash 0731$0.44/M TokensStart from$0.264/M Tokens-40%