Realtime AI News
Grok 4.6 arrives, topping GPT-5.6 Sol and Fable 5 Max on agent benchmarks at $2 per million tokens
SpaceXAI has released Grok 4.6, which official benchmarks show beating GPT-5.6 Sol and Fable 5 Max on GDPVal-AA v2 and other agentic tests while priced at just $2 per million input tokens and $6 per million output tokens. The model is live in Grok Build, Cursor, Grok Bot and the API — with OpenRouter, Vercel and Cloudflare also carrying it — and is tuned for long-horizon agent tasks.
SpaceXAI has released Grok 4.6, its new flagship model that the company's official benchmarks show beating GPT-5.6 Sol and Fable 5 Max — a return to the frontier tier for Musk's AI operation. The model is live in Grok Build, Cursor, Grok Bot and the open API, with OpenRouter, Vercel and Cloudflare also onboard from day one.
On GDPVal-AA v2, a benchmark of real-world working ability, Grok 4.6 scored a top 1,753, ahead of GPT-5.6 Sol's 1,728 and Fable 5 Max's 1,741; it also led AA-Briefcase (1,577) and Harvey LAB (15.8%).
On the AA Intelligence Index the model scored 61, tied with GPT-5.6 Sol, five points above Grok 4.5 and one behind Fable 5 Max. Its weakness is pure terminal work: Terminal-Bench v3.0 came in at 26%, below GPT-5.6 Sol's 34.6%.
Price is the big differentiator. Grok 4.6 costs $2 per million input tokens and $6 per million output tokens, flat versus Grok 4.5, while GPT-5.6 Sol Max lists $5/$30 and Claude Opus 5 lists $5/$25. The fine print: requests with over 200K tokens of context double to $4/$12 for the entire request, and a low-latency variant costs twice the standard price.
The release is explicitly aimed at long-horizon agent tasks — researching unfamiliar domains, editing across codebases, iterating on feedback for many rounds without restarting — and SpaceXAI is offering double usage for the first week in Cursor and Grok Build to pull developers in.
Grok 4.6 also powers Grok Bot, the 24/7 agent team launched the day before: each bot gets its own cloud environment, can log into tools and sites the user already uses, runs in parallel around the clock, and learns reusable routines from watching the user. Grok Bot's download and signup pages still sit on cursor.com infrastructure, a visible mark of the Cursor acquisition.
Technically, Grok 4.6 continues training from Grok 4.5 using self-generated reasoning and engineering data, plus optimizer tuning and reinforcement learning, running on SpaceXAI's Colossus supercomputer cluster.
Musk has already said Grok 4.7's initial training is complete, with large amounts of SpaceX data being added and a launch expected in three to four weeks, claiming it will surpass every current model — a timeline that will test whether SpaceXAI can hold its place at the top of the table.
Why it matters
With aggressive pricing and an agent-first focus across Grok Bot and Cursor, Grok 4.6 puts SpaceXAI back in the flagship race while pressing rivals on cost.
Nearby Updates
All08/13, 19:53
Embodied-data startup SCALEFORCE closes two funding rounds in 40 days to build physical-AI infrastructure
Embodied-intelligence data infrastructure startup Yuanpoint Technology (SCALEFORCE) has completed a new funding round worth tens of millions of yuan, with investors including Hengxu Capital, Kailian Capital and a top domestic embodied-AI industry player — its second round within 40 days. The proceeds will fund the MatrixOS physical-AI operating system, a large-scale data production network, and team expansion.
08/13, 20:34
Edge AI chip startup Acrab raises $130M Series B as GΞLIX 1 heads to production
Acrab, a Singapore-based AI compute infrastructure company, has closed a $130 million Series B with Vertex Growth among the participants, bringing total funding past $480 million for production capacity, ecosystem and next-generation platform development. The startup, founded in 2024, has moved its first edge AI chip GΞLIX 1 into customer onboarding and scaled production preparation alongside the Agent Box personal AI center.
08/13, 19:29
Claude clears all Hadamard matrices below order 2000, striking another open problem off the math list
Chinese tech outlet QbitAI reports that Anthropic's Claude has cleared all Hadamard matrices below order 2000, removing a long-standing open problem in combinatorics from the to-solve list. The report credits the mathematician's approach as much as the model itself, a fresh sign that large models can contribute to serious mathematical research.
08/13, 19:01
Maverick Payments integrates Findustry AI agent to automate chargeback rebuttals
Payment processor Maverick Payments has integrated Findustry's AI agent to automate chargeback rebuttals, as reported by PYMNTS. The move brings an agent into merchant dispute handling, promising lower operational burden and adding another real-world AI agent deployment in payments.