AI
Grok 4.7 Just Launched at the Same Price, and the Independent Scoreboard Told the Honest Story
SpaceXAI shipped Grok 4.7 on Monday, September 21, at $2 per million input tokens and $6 per million output tokens for requests under 200,000 input tokens, rising to $4 in and $12 out above that line, exactly Grok 4.6's prices. The context window stays at 500,000 tokens. The knowledge cutoff moved forward three months to May 2026. Reasoning effort now runs across four levels, low, medium, high and xhigh, with published figures quoted at xhigh. A fast variant runs at twice the output speed for twice the price. In a market where every flagship launch seems to rewrite the pricing page, a stronger model at the same price is the kind of upgrade developers celebrate.
Under the hood, the company describes three changes rather than a new architecture. The base model is larger than Grok 4.6's. The reinforcement learning run was longer and weighted toward tasks that take many hours to complete. And the model was trained to verify its own work and to hold long context better, including native understanding of the Grok Bot harness. The vendor's own one line summary reads "twice as fast, at half the price of comparable models", a positioning claim about rivals, worth reading as exactly that.
The vendor's benchmark table deserves its label, vendor reported, on the vendor's chosen evaluations, published on launch day, when launch day numbers are at their least reliable. Read that way, the card is still informative. Grok 4.7 beats Grok 4.6 on all eight listed evaluations, the minimum a successor owes its users, and posts a mixed card against the field, ahead of GPT-5.6 Sol on CursorBench (46.3 against 41.7), Terminal-Bench (38.0 against 37.3) and EEBench (64.0 against 39.4), behind it on DeepSWE (71.0 against 72.7); behind Claude Fable 5.1 on most rows, ahead on EEBench and the Harvey legal agent eval (19.6 against 6.7).
The more useful numbers arrived the same day from Artificial Analysis, the independent evaluator in this story. On the Intelligence Index, Grok 4.7 scores 46 on the v4.3.2 scale, two points above Grok 4.6's 44, enough to bring SpaceXAI into the index's top four labs, with Claude Fable 5.1 at 53, GPT-6 Astra at 53, Claude Opus 5 at 51 and GPT-5.6 Sol at 47. Then the second number, the Coding Agent Index, measured in SpaceXAI's own Grok Build harness, lands at 56 against Grok 4.6's 47, a nine point jump, fourth among native harness models, ahead of GPT-5.6 Sol. This is the first outside evidence that the longer reinforcement learning run on multi hour tasks did something real, and it is the number a team choosing a coding agent should weigh hardest.
Every gain has a price, and Artificial Analysis found this one's receipt. Its evaluation run generated 240 million output tokens on the Intelligence Index against a median of 94 million across tracked models, more than double, putting the cost of one index task at $3.74. At $6 per million output tokens, verbosity is the line item that decides whether an agentic loop is affordable. So the honest reading is a trade rather than a coronation, nine points of coding agent gain purchased with far more output tokens, and general intelligence roughly level with the predecessor once every token is priced. Claude Fable 5.1, listed at $10 in and $50 out per million tokens against Grok 4.7's $2 and $6, sits seven index points ahead but costs roughly five times as much to run. Which of those facts decides your choice depends on whether the workload is a coding agent or a knowledge task, and that split is more useful than any single rank.
One more ledger entry for the rumor mill. The pre launch chatter got the price exactly right, though only because a catalog pull request copied Grok 4.6's rates, luck dressed up as a leak. The circulating parameter count, roughly 2.1 trillion, appeared on specification sheets at launch, only in the rumor mill. The practical read for developers, test the coding harness, watch the token bill, and enjoy the frontier's most honest price in months.
Quick answers
What is this story about?
SpaceXAI shipped Grok 4.7 on Monday, September 21, at $2 per million input tokens and $6 per million output tokens for requests under 200,000 input tokens, rising to $4 in and $12 out above that line, exactly Grok 4.6's prices. The context window stays at 500,000 tokens. The knowledge cutoff moved forward three months to May 2026. Reasoning effort now runs across four levels, low, medium, high and xhigh, with published figures quoted at xhigh. A fast variant runs at twice the output speed for twice the price. In a market where every flagship launch seems to rewrite the pricing page, a stronger model at the same price is the kind of upgrade developers celebrate.
Why does this story matter?
One more ledger entry for the rumor mill. The pre launch chatter got the price exactly right, though only because a catalog pull request copied Grok 4.6's rates, luck dressed up as a leak. The circulating parameter count, roughly 2.1 trillion, appeared on specification sheets at launch, only in the rumor mill. The practical read for developers, test the coding harness, watch the token bill, and enjoy the frontier's most honest price in months.
Sources
New to crypto? Read the crypto glossary, browse frequent questions, read our story, or explore the story archive.