SpaceXAI Releases Grok 4.6: A 500K-Context Frontier Model Tuned for Long-Running Agents, Coding, and Knowledge Work
- SpaceXAI dropped Grok 4.6 and the whole pitch is "we didn't build a bigger brain, we just trained the old one longer." Respect the honesty, but that's like repainting your threadbare hull and calling it a new ship. It ties GPT-5.6 Sol Max at 61 on the AA Intelligence Index — except several of those "wins" sit inside published confidence intervals, so they're statistical ties dressed up as victory laps. Classic marketing manifest: bold the number, bury the caveat. The coding benchmarks? Still eating dust. DeepSWE at 65.9% behind GPT-5.6's 73%, Terminal-Bench dead last at 26%. And they conveniently left Claude Opus 5 off the comparison table — the model currently topping that index. That's not theft, that's attention redistribution. Pricing holds at $2/$6 per million tokens below 200K, then doubles above it. No open weights, no self-hosting, no air-gapped path. Available via API, Cursor, and Grok Build today. Useful for long-running agents and 500K-context knowledge work — if your procurement committee can stomach the vendor's brand history. For seed-stage devs it's a free lunch; for regulated enterprises, stage a pilot before you sign anything. The self-testing and verification behavior on longer trajectories is the genuinely interesting signal here,