iCreat AI

Grok 4.6 Released: Benchmarks, API Pricing, and What the List-Price Comparisons Miss

Last UpdateAugust 13, 2026
Generate with
Grok 4.6 benchmarks, API pricing, and alternatives

Opening Notes

xAI (SpaceXAI) released Grok 4.6 on August 12, 2026, built for long-running agents and coding. API pricing is $2 / $6 per 1M input/output tokens, rising to $4 / $12 for prompts of 200K+ tokens. A fast variant costs double.

It scores 61 on the Artificial Analysis Intelligence Index, tied with GPT-5.6 Sol Max, one point behind Claude Fable 5 Max (62), two behind Claude Opus 5 (63).

On xAI's own published evals, Grok 4.6 wins GDPVal-AA v2, AA-Briefcase, and Harvey LAB. It loses clearly on DeepSWE v1.1 (65.9% vs Sol's 73%) and Terminal-Bench v3.0 (26% vs ~34% for both rivals).

Every comparison this week uses official list prices. Through API aggregators the picture shifts. Claude Fable 5 runs $1 / $5 per 1M tokens on iCreat, below Grok 4.6's own official rate, while scoring higher on the AA Index. GPT-5.6 Sol runs $0.50 / $3.

Grok 4.6 is a genuinely strong value release. If you pick a model by cost-per-intelligence this week, though, the math points somewhere else.

1. What actually shipped

On August 12, xAI announced Grok 4.6, the successor to Grok 4.5 released just five weeks earlier. xAI pitches it as a model for long-horizon agent work. Per the announcement, it stays with complex tasks across many steps, whether that's researching a topic across a codebase or turning a rough idea into a working application.

The confirmed facts from the official announcement.

  • Available now in Cursor, Grok Build, the xAI API, and via OpenRouter, Vercel, and Cloudflare.
  • Pricing starts at $2 per 1M input tokens and $6 per 1M output tokens. A fast variant costs double. Per xAI's API docs, prompts of 200K+ tokens are billed at $4 / $12.
  • The context window is 500K tokens, unchanged from Grok 4.5, per Artificial Analysis. Cache hits cost $0.50 per 1M tokens.
  • The gains come from post-training. xAI describes a longer supplemental training run, regenerated SFT trajectories, and broad agentic RL. No claims about scale.
  • Launch promo, 2x included usage in Grok Build and Cursor for the first week.

One widely circulated claim deserves a flag. The "1.5T parameters, same V9 foundation as Grok 4.5" figure comes from leaks and pre-launch chatter. It appears nowhere in xAI's announcement or docs. Treat it as unconfirmed.

2. Benchmarks, where it wins and where it loses

xAI published a ten-eval comparison against Grok 4.5 High, GPT-5.6 Sol Max, and Claude Fable 5 Max. Competitor figures come from published system cards and leaderboards. All of it is self-reported, so read the numbers as directional rather than gospel. Here is the full table, including the rows where Grok loses.

BenchmarkGrok 4.6 HighGrok 4.5 HighGPT-5.6 Sol MaxFable 5 Max
AA Intelligence Index61566162
GDPVal-AA v2 (Elo)1753152617281741
CursorBench v3.269.9%66.7%67.2%70.5%
DeepSWE v1.165.9%54%73%70%
FrontierCode v1.1 (Ext)61.3%56.6%60.6%63.6%
APEX-Agents57.5%47.1%56.7%59.2%
APEX-SWE56.4%53.6%58.8%
Terminal-Bench v3.026%15.7%34.6%34.1%
AA-Briefcase (Elo)1577131315021574
Harvey LAB (Vals)15.8%12.9%2.5%11.3%
Artificial Analysis Intelligence Index chart showing Claude Opus 5 at 63, Claude Fable 5 at 62, GPT-5.6 Sol and Grok 4.6 tied at 61

*Grok 4.6 joins the frontier tier, one point behind Fable 5. Chart by Artificial Analysis.*

Three things stand out.

Grok 4.6 crushes its predecessor. Every single eval improves, most by wide margins, five weeks after Grok 4.5 shipped. The cadence itself is the story.

Against the frontier, it trades blows. It leads on real-world knowledge work (GDPVal-AA, AA-Briefcase, Harvey LAB) and trails on software engineering depth. The DeepSWE gap, 65.9% against Sol's 73%, and the Terminal-Bench v3.0 gap, 26% against 34.6%, are not rounding errors. In agentic coding pipelines, failed runs mean retries, and retries raise your effective cost per completed task.

Version numbers matter. Artificial Analysis separately reports Grok 4.6 scoring 88.4% on Terminal-Bench v2.1, level with the leaders, while xAI's own table shows 26% on v3.0, well behind. Same benchmark family, opposite conclusion. Any comparison you read this week that cites "Terminal-Bench" without a version number is noise.

Independent analysis backs the mixed-but-real picture. Artificial Analysis measured Grok 4.6 at $0.84 per task on their Intelligence Index and highlighted its turn efficiency. It resolves AA-Briefcase tasks in ~53 turns and ~0.5B input tokens on average, versus ~103 turns and ~2.0B input tokens for Claude Opus 5 Max. That efficiency is an advantage xAI has earned, and it places Grok 4.6 on the intelligence-vs-cost Pareto frontier, at least when models are compared at official list prices.

That qualifier is doing more work than it looks.

3. The cost math beyond list prices

Every "Grok 4.6 is the value king" take this week compares official list prices. Here they are.

ModelOfficial price (per 1M in/out)Source
Grok 4.6$2.00 / $6.00 ($4 / $12 for ≥200K prompts)xAI
Claude Opus 5$5.00 / $25.00Anthropic
GPT-5.6 Sol$5.00 / $30.00 ($10 / $45 for >272K input)OpenAI
Claude Fable 5$10.00 / $50.00Anthropic

At list prices the pitch holds up. Sol-level composite intelligence at 40% of the input price and 20% of the output price.

Intelligence Index vs cost per task scatter plot placing Grok 4.6 on the Pareto frontier at official list prices

*At official list prices, Grok 4.6 sits on the intelligence-vs-cost Pareto frontier. Chart by Artificial Analysis. Note the x-axis uses list-price-based cost per task, which is exactly the assumption the next table breaks.*

List prices aren't the only prices, though. API aggregators buy capacity at scale and resell below list. Here is the same table with iCreat pricing added. These are standing rates, not launch promos.

ModelOfficial (in/out per 1M)On iCreatvs. official
Claude Fable 5$10.00 / $50.00$1.00 / $5.0090% lower
GPT-5.6 Sol$5.00 / $30.00$0.50 / $3.0090% lower
Grok 4.6$2.00 / $6.00

Cache pricing follows the same pattern. Fable 5 cache reads are $0.10 per 1M on iCreat versus $1.00 official, and Sol cache reads are $0.05 versus $0.50.

*Claude Fable 5 pricing on iCreat, $1 input / $5 output per 1M tokens. Live rates on the model page.*

The second table changes the conclusion of the first. Fable 5 outscores Grok 4.6 on the AA Index and beats it head-to-head on seven of the ten published evals, and on iCreat it costs half of Grok 4.6's own official rate. These are the same models compared in xAI's benchmark table. Reasoning effort is fully adjustable through the API, so the Max-tier configurations cited above are reachable.

A concrete scenario

Say you run an agentic coding workflow burning 10M input and 2M output tokens per day, roughly a mid-sized team's agent traffic.

Daily costMonthly (30d)
Fable 5 @ official$200.00$6,000
GPT-5.6 Sol @ official$110.00$3,300
Grok 4.6 @ official$32.00$960
Fable 5 @ iCreat$20.00$600
GPT-5.6 Sol @ iCreat$11.00$330

At list prices, Grok 4.6 saves you thousands per month over the frontier leaders. The value story is legitimate. Route the same workload through iCreat and the stronger model costs $360/month less than Grok 4.6's official rate, and Sol costs a third of it. This assumes prompts under each provider's long-context threshold; heavy-cache workloads shift the numbers further toward the discounted cache rates above.

4. So should you integrate Grok 4.6?

Depends on which of these is you.

You need maximum agent reliability today. Fable 5 Max leads APEX-Agents, CursorBench, FrontierCode, and APEX-SWE, and Terminal-Bench v3.0 shows an 8-point gap over Grok 4.6. For long-horizon pipelines where a failed run costs you retries and review time, the stronger model at $1/$5 is the boring, correct choice.

You're cost-optimizing hard on coding tasks. GPT-5.6 Sol at $0.50/$3.00 is the cheapest path to a 61 AA Index score in this post, cheaper than Grok 4.6 itself. Its DeepSWE 73% is also the best software-engineering score in the table.

You specifically want Grok 4.6's profile. The turn efficiency and knowledge-work Elo are real, and inside Cursor or Grok Build the first-week 2x usage promo is genuinely worth trying. Nothing here says the model is bad. It is a strong release at an honest price. Its headline advantage is that price, though, and the advantage doesn't survive contact with aggregator rates.

You'd rather not bet on one lab at all. Frontier leadership has changed hands three times this summer. Opus 5 led in July, Sol tied days later, now Grok 4.6 joins them, and a larger Grok 4.7 is reportedly weeks away (unconfirmed). A single API key that routes to whichever model is on top this month is the only position that doesn't require predicting the next flip.

5. Switching costs about 90 seconds

iCreat exposes Anthropic- and OpenAI-compatible endpoints, so existing SDK code needs only a base URL and key change.

# Claude Fable 5 — standard Anthropic SDK
import anthropic

client = anthropic.Anthropic(
    base_url="https://api.icreat.ai/v1/llm",
    api_key="ICREAT_API_KEY",
)
message = client.messages.create(
    model="claude-fable-5",
    messages=[{"role": "user", "content": "Hello"}],
)
# GPT-5.6 Sol — standard OpenAI SDK
from openai import OpenAI

client = OpenAI(
    base_url="https://api.icreat.ai/v1/llm",
    api_key="ICREAT_API_KEY",
)
completion = client.chat.completions.create(
    model="gpt-5.6-sol",
    messages=[{"role": "user", "content": "Hello"}],
)

There's no free tier, but top-ups start at $1. At iCreat's Fable 5 rates that buys about 1M input tokens, enough to run your own eval before committing to anything. The same key covers image, video, audio, and 3D model APIs, one integration for every modality.

FAQ

What is Grok 4.6's context window?
500K tokens, unchanged from Grok 4.5, per Artificial Analysis. Prompts of 200K+ tokens are billed at doubled rates ($4 / $12 per 1M).
How much does the Grok 4.6 API cost?
$2 per 1M input tokens and $6 per 1M output tokens. The fast variant is $4 / $12. Cache hits are $0.50 per 1M.
Is Grok 4.6 better than Claude Fable 5?
On xAI's own published table, Fable 5 Max scores higher on the AA Intelligence Index (62 vs 61) and beats Grok 4.6 head-to-head on seven of ten evals, including most coding and agent benchmarks. Grok 4.6 leads on GDPVal-AA v2, AA-Briefcase, and Harvey LAB. Fable 5 is stronger overall while Grok 4.6 is cheaper at list price.
Is Grok 4.6 better than GPT-5.6 Sol?
They tie on the AA Index (61). Sol is clearly ahead on DeepSWE (73% vs 65.9%) and Terminal-Bench v3.0 (34.6% vs 26%). Grok 4.6 leads on knowledge-work evals like GDPVal-AA and Harvey LAB.
How is Grok 4.6 different from Grok 4.5?
Same 500K context and same $2/$6 pricing, with improvements concentrated in post-training (SFT and RL). It beats Grok 4.5 on all ten published benchmarks, released five weeks apart.
Is there a cheaper way to use frontier models than Grok 4.6?
Yes. Through API aggregators, Claude Fable 5 runs $1 / $5 per 1M tokens on iCreat and GPT-5.6 Sol runs $0.50 / $3.00, both below Grok 4.6's official $2 / $6.
How many parameters does Grok 4.6 have?
xAI hasn't said. The widely repeated "1.5 trillion parameters" figure comes from pre-launch leaks and remains unconfirmed.