ModelsHot

Grok 4.7 Is Out: Same $2/$6 Price, a Bigger Base Model, and a 500K Context

xAI released Grok 4.7 on September 21 at the same $2/$6 pricing as Grok 4.6, built on a larger base model with longer reinforcement learning for long-horizon tasks. It nearly doubles its Terminal-Bench score but still trails Fable 5.1 on several tests, and it burns far more tokens per task.

Evan BrooksEvan Brooks
Heat: 1,150
Grok 4.7, xAI's new model release

xAI's newest model is bigger under the hood but costs exactly the same as the one it replaces. Grok 4.7 targets work that runs for hours instead of seconds, and it's shipping into Cursor and Grok Build on day one. Here's what changed, where it wins, and where rival models still lead.

What landed on September 21

xAI released Grok 4.7 on September 21, 2026, calling it its most capable model for coding and knowledge work. The pitch is straightforward: a larger base model, trained longer on harder problems, and priced identically to Grok 4.6.

The price is the selling point. Grok 4.7 runs $2 per million input tokens and $6 per million output tokens, the same numbers xAI charged for 4.6. Cache hits drop to $0.50 per million tokens. Fable 5.1 Max costs 5x more on input and roughly 8.3x more on output.

You can call it from the xAI API, and it also turns up in Cursor, Grok Build, OpenRouter, Vercel, and Cloudflare.

What changed under the hood

xAI lists four changes from Grok 4.6.

The first is a new base model. Grok 4.7 doesn't reuse the one behind 4.6, which is unusual for a point release and explains the size of the jumps on some tests.

The second is a longer reinforcement learning run, weighted toward problems that take many hours to finish. Reinforcement learning is the training stage where a model gets feedback and learns which approaches pay off, so spending it on long tasks pushes the model toward jobs it previously abandoned partway through.

The third change is better self-verification and long-context handling. xAI says the model checks its own work more carefully, which is the kind of claim that matters most on agent runs that chain dozens of steps.

The fourth is native support for the Grok Bot framework, the agent setup xAI uses to coordinate roles inside a project.

Grok 4.7 specs and API details

Property

Value

Model name

grok-4.7

Context window

500,000 tokens

Knowledge cutoff

May 2026

Modalities

Text and image in, text out

Reasoning effort

low, medium, high (default), xhigh

APIs

Responses API, Chat Completions

Tools

Function calling, web search, X search, code execution

Price

$2 input / $6 output per 1M tokens

Cache hits

$0.50 per 1M tokens

The 500,000-token context window and the May 2026 knowledge cutoff leaked ahead of launch and are now listed in xAI's developer docs, so they've moved from rumor to confirmed. Note that the context window didn't grow from Grok 4.6, which carried the same figure.

The company also offers a faster variant called Grok 4.7 Fast, the same model on quicker infrastructure with double the output speed at double the price. It runs only in Cursor and Grok Build, not on the public API, and it's excluded from Grok Build's free tier.

Where Grok 4.7 gains, and where it doesn't

The launch table compares Grok 4.7 at xHigh effort against Grok 4.6 High, GPT-5.6 Sol Max, and Fable 5.1 Max. Every figure is vendor-reported, so the rankings hold until independent labs run their own tests.

Benchmark

Grok 4.7 xHigh

Grok 4.6 High

GPT-5.6 Sol Max

Fable 5.1 Max

CursorBench 4.0

46.3%

40.4%

41.7%

51.8%

DeepSWE v1.1

71.0%*

65.2%

72.7%

70.0%

EEBench

64.0%

53.0%

39.4%

56.4%

Terminal-Bench 4.0

38.0%

20.3%

37.3%

57.9%

Harvey Legal Agent

19.6%

15.8%

2.5%

6.7%

HealthBench Professional

56.7%

48.5%

60.5%

62.1%

*Run at high effort.

Grok 4.7 beats Grok 4.6 on every row. The biggest jump is Terminal-Bench 4.0, nearly doubling from 20.3% to 38.0%. EEBench rose 11 points to 64.0%, the best figure in the table.

But xAI's own numbers show it doesn't lead across the board. Fable 5.1 Max wins four of the seven tests, including Terminal-Bench at 57.9%. GPT-5.6 Sol Max holds the top DeepSWE score at 72.7%.

That leaves developers with a workload-shaped decision rather than a clear winner. Grok 4.7 takes the terminal and legal tasks. Fable 5.1 Max takes the longer coding and health evaluations.

What independent testing adds

Artificial Analysis ran Grok 4.7 at xhigh effort and put it at 46 on its Intelligence Index, up two points from Grok 4.6 and enough to move xAI into the top four labs on that measure. Claude Fable 5.1 and GPT-6 both score 53.

The gains show up most in agentic knowledge work. On AA-Briefcase, a private benchmark for long-horizon professional tasks, Grok 4.7 gained 111 Elo over Grok 4.6 to reach 1,657, landing just behind Claude Opus 5 and Fable 5.1. On GDPval-AA it scored 1,695 Elo, ahead of 4.6's 1,605.

Coding agents improved too. Paired with Grok Build, it scored 56 on the Coding Agent Index, up nine points, ranking fourth behind Fable 5.1, GPT-6 Astra, and Claude Opus 5.

There's a catch buried in the same report. Grok 4.7 burns through roughly 81,000 output tokens per Intelligence Index task, against about 36,000 for Grok 4.6 and 27,000 for GPT-6 Astra. Cheaper tokens only help if you don't need two to three times as many of them.

Reliability and safety

xAI shipped Grok 4.7 with a new safeguard stack and calls it the strongest model it has tested on refusals and jailbreak resistance. It topped LatchBio's biosafety benchmark at 62.4%.

On HackerBench v0.3, xAI's own test for risky cyber tasks, 3.3% of dual-use prompts got through. The company says it rarely blocks legitimate security work, and select cybersecurity partners now get invite-only red-team access.

Independent testing adds one more data point. Artificial Analysis measured Grok 4.7's hallucination rate at 29%, down from 34% for Grok 4.6, while accuracy stayed flat at 47% against 48%.

Timing, delays, and the parameter claim

This release was late. Elon Musk postponed Grok 4.7 at least five times starting in late July. On September 11 he said it needed a few more days, citing premature abandonment of difficult tasks. Two days later xAI showed off Grok 4.8 instead, which looked a lot like a signal that 4.7 still wasn't done.

One claim floating around deserves caution. Reports have described Grok 4.7 as a 2.1 trillion parameter model. Musk hasn't confirmed that figure, and xAI's documentation doesn't publish a parameter count. Treat it as unverified until the company says otherwise.

If you already pay for Cursor or use Grok Build, Grok 4.7 is worth a test on your longest-running tasks, since that's where the gains cluster. If your work is short prompts and quick answers, the model offers little over 4.6 at the same price, and the higher token consumption could make it slower in practice.

Share This Story