[release] 5 min · Sep 22, 2026

Grok 4.7 — SpaceX Owns the IDE; Benchmarks Still Lag

Grok 4.7 ships with 2.1T parameters trained on SpaceX telemetry at unchanged pricing. Fable 5.1 still leads on CursorBench. The real risk is the data flywheel.

Grok 4.7 ↗ Sep 21, 2026
#grok#cursor#ai-models#spacex#vendor-lock-in

SpaceX now controls both the IDE and the model your team codes with — and that ownership matters more than Grok 4.7’s incremental benchmark gains. xAI released Grok 4.7 on September 21, 2026: a 2.1-trillion-parameter model trained partly on SpaceX Starlink telemetry, manufacturing records, and engineering fault logs. It is live in Cursor, Grok Build, the xAI API, and GitHub Copilot across Pro, Pro+, Max, Business, and Enterprise tiers. Pricing is unchanged from Grok 4.6. The capability jump is real but modest, and on the benchmarks that matter for day-to-day coding, Fable 5.1 still leads.

TL;DR

  • What: Grok 4.7 ships with 2.1T parameters (40% up from 4.6), trained on SpaceX proprietary engineering data, at identical pricing
  • Benchmarks: Fable 5.1 leads on CursorBench 4.0 (51.8% vs 46.3%), Terminal-Bench, and enterprise agent evals. Grok 4.7 narrowly wins DeepSWE v1.1 at high effort (71.0% vs 70.0%)
  • Hidden cost: X Search moved to per-fetch pricing the same day — $5/1k posts, $10/1k profiles — not reflected in headline token pricing
  • Risk: Every Cursor workflow on Grok 4.7 routes through SpaceX-owned infrastructure, feeding a model optimized for SpaceX’s engineering objectives

Grok 4.7 — What Happened

The core change is a bigger base model — 2.1 trillion parameters versus Grok 4.6’s 1.5 trillion — combined with a longer reinforcement learning run on a harder task mix weighted toward problems that take many hours to complete. xAI also folded in supplemental training data from SpaceX: satellite telemetry, manufacturing records, and engineering failure logs. The result is a model that is meaningfully better at sustained reasoning tasks but not categorically different from its predecessor.

Pricing stays at $2/$6 per million input/output tokens under 200K context, and $4/$12 above 200K. A Fast variant runs at twice the output speed for twice the price, available in Cursor and Grok Build but not on the public xAI API. GitHub Copilot integration uses usage-based billing at provider list pricing — if your team has budget controls set for previous models, verify whether Grok 4.7 triggers a new billing category or falls under existing spend caps.

X Search pricing changed the same day. Effective September 21 at 12

PM Pacific, xAI moved from $5 per 1,000 API calls to $5 per 1,000 posts fetched and $10 per 1,000 profiles fetched. Any Grok 4.7 agent workflow that invokes X Search now incurs per-fetch charges on top of token costs. The headline “$2/$6 per million tokens” does not reflect this.

Why This Matters

The benchmark story is more nuanced than the xAI announcement suggests, and more nuanced than the reflexive “Grok still loses” take.

On CursorBench 4.0, Grok 4.7 scores 46.3% at xhigh effort. That is a genuine improvement — but the comparison xAI highlights is misleading. They benchmark Grok 4.6 at 40.4% at high effort, not xhigh, making the headline gain appear larger than it is on equal footing. More importantly, Fable 5.1 scores 51.8% on the same benchmark. That is a 5.5-point gap on the test most directly relevant to Cursor users — the people xAI most wants to convert. On Terminal-Bench, the gap is even wider: Fable 5.1 leads by nearly 20 points.

Where Grok 4.7 does win is DeepSWE v1.1, scoring 71.0% at high effort versus Fable 5.1’s 70.0%. That is a real but razor-thin lead on a software engineering benchmark — one percentage point, and only at high effort. On the enterprise agent evaluations managed by Artificial Analysis, Grok 4.7 logs 1,695 Elo on GDPval and 1,657 on AA Briefcase v1.1. Both trail Fable 5.1 significantly: 1,853 on GDPval and 1,694 on AA Briefcase. Artificial Analysis’s first independent read scores Grok 4.7 at 46 on its Intelligence Index — rank 16 of 655 models. Credible, but not the top tier xAI’s marketing implies.

I don’t trust self-reported benchmark tables from a company that controls the evaluation harness and the IDE the benchmark runs in. The cross-effort comparison on CursorBench is the kind of statistical sleight-of-hand that would get a grad student’s paper rejected. When the company publishing the benchmarks also owns the IDE where the model is deployed, the incentive structure is obvious: make the numbers look good enough to flip the default model in Cursor, then let inertia do the rest.

If you are evaluating Grok 4.7 for your team, run your own evals on your actual codebase. The public benchmarks test synthetic tasks. Your daily coding involves your frameworks, your conventions, your repo structure. A model that scores 5 points lower on CursorBench might still outperform on your specific workload — or it might not. The point is that xAI’s published numbers do not answer this question for you.

The deeper concern is structural, not statistical. SpaceX now owns both sides of the coding workflow: the model and (through its acquisition) the IDE. The Grok 4.5 → 4.6 → 4.7 training pipeline uses a longer RL run each generation, with tasks increasingly weighted toward the kind of sustained multi-step reasoning that Cursor’s agentic mode demands. xAI explicitly trained Grok 4.7 on the Grok Bot agent harness, meaning the model is optimized not just for code completion but for the specific tool-use patterns inside SpaceX’s own infrastructure.

Every Cursor workflow you run on Grok 4.7 routes through SpaceX-owned infrastructure. The supplemental training on SpaceX telemetry and engineering logs is not just a capability claim — it is a statement about what this model is for. SpaceX is building a data flywheel where developer workflow data improves a model whose primary training objective is SpaceX’s own engineering challenges. That is a legitimate compute moat, but it is a moat built around you, not for you. The model gets better at SpaceX problems first and your problems second.

This pattern — vertically integrated model training within a captive IDE ecosystem — is exactly what we flagged when Grok 4.5 launched with Cursor data integration. Grok 4.7 does not change the trajectory. It accelerates it.

The Take

Grok 4.7 is a better model than Grok 4.6. It is not a better model than Fable 5.1 for coding, which is the only thing that matters if you are choosing a default model in Cursor. The CursorBench gap is 5.5 points. The Terminal-Bench gap is nearly 20 points. The DeepSWE win is one percentage point at high effort. If you are making a model decision based on benchmarks, the benchmarks say Fable 5.1.

But benchmarks are not why I am concerned. Before you make Grok 4.7 your team’s default, ask a simpler question: who benefits most from millions of developers running agentic workflows through SpaceX-owned infrastructure? The answer is not “developers.” The answer is a company training successive model generations on your coding patterns while simultaneously locking you into an IDE it controls. The Copilot default enablement timeline adds another distribution channel to the same flywheel.

My recommendation: use Fable 5.1 in Cursor for coding tasks. If you need Grok 4.7 for a specific capability — sustained multi-hour reasoning, SpaceX-adjacent engineering domains — use it through the xAI API directly, where at least you control the integration point. Do not hand SpaceX both the model and the IDE without understanding what you are giving up in the exchange.