Claim of novelty
On 21 September 2026 SpaceXAI published Introducing Grok 4.7 and a same-day Model Card: Grok 4.7 (revision 2026-09-21). The API model page lists the id grok-4.7.
The launch post’s claim of novelty is narrow and dated. SpaceXAI says Grok 4.7 uses a new, larger base than Grok 4.6, a longer reinforcement-learning run weighted toward multi-hour tasks, stronger self-verification, longer-context management, native training on the Grok Bot harness, and a new safeguard stack. It states the model is served at the same price and speed as Grok 4.6, and names day-one surfaces: Cursor, Grok Build, the Grok API, third-party coding harnesses, and model routers and cloud platforms. The model card says consumer surfaces (web, mobile apps, and Grok-in-X) are planned for a later date.
The model card repeats the “most capable model to date” line for coding, engineering, and office work, and records a pretraining cutoff of June 2026 with supplemental data through August 2026. The grok-4.7 API docs page we fetched does not state a knowledge cut-off.
This Record freezes those vendor documents and the Artificial Analysis third-party page dated 21 September 2026. It does not re-run the evals.
What was measured and on which data
Vendor launch table (SpaceXAI blog, 21 Sep 2026). Effort labels in the table header: Grok 4.7 xHigh, Grok 4.6 High, GPT-5.6 Sol Max, Fable 5.1 Max. Prices shown: Grok input/output $2 / $6 per million tokens; GPT-5.6 Sol $4 / $20; Fable 5.1 $10 / $50.
| Suite | Grok 4.7 | Grok 4.6 | GPT-5.6 Sol Max | Fable 5.1 Max |
|---|---|---|---|---|
| CursorBench 4.0 | 46.3% | 40.4% | 41.7% | 51.8% |
| DeepSWE v1.1 | 71.0%* | 65.2% | 72.7% | 70.0% |
| EEBench | 64.0% | 53.0% | 39.4% | 56.4% |
| AA Briefcase v1.1 | 1,657 | 1,546 | 1,487 | 1,678 |
| Terminal-Bench 4.0 | 38.0% | 20.3% | 37.3% | 57.9% |
| Harvey Legal Agent Benchmark | 19.6% | 15.8% | 2.5% | 6.7% |
| HealthBench Professional | 56.7% | 48.5% | 60.5% | 62.1% |
* Blog footnote: high effort on DeepSWE for Grok 4.7.
Vendor model card (same date), selected coding and knowledge-work rows. Unless noted, card scores are on the final deployed checkpoint. Effort and harness are named per subsection.
- CursorBench 4.0: 46.3% at xhigh; 43.9% at high. Card says v4.0 scores are not comparable with CursorBench 3.2.
- DeepSWE v1.1: 71.0% at high, mini-SWE-agent harness, Datacurve-run results.
- Terminal-Bench 4.0: 38.0% at xhigh with Grok Build harness (Harbor-run peer table in the card).
- FrontierSWE V2: 29.0% mean@5 at xhigh (Proximal Labs / Proximus harness).
- SWE-Marathon v1.1: 46.0% at high (Abundant AI).
- Legal Agent Benchmark (Vals / Harvey): 19.6% at xhigh, held-out 120-task Valkyrie run, internet disabled.
- EEBench: 66.0% at xhigh with Grok Build (Atopile). This differs from the blog table’s 64.0% for Grok 4.7. The card also lists Grok 4.6 at 60.0% xhigh and 53.0% high.
- CADGenBench generation split: 44.4% at high (Mecado).
- HealthBench Professional: 56.7% at xhigh; Grok 4.6 (high) used as grader for Grok 4.7 results; filtered tasks scored zero.
- LatchBio Capabilities v1.0 overall: 44.5% at xhigh (equal-weight mean of 11 capability benches).
Vendor safety and dual-use rows (model card). These are SpaceXAI’s measurements. Several cyber and bio capability probes run without production safeguards; refusal suites use release-tracked safeguards.
- LatchBio BioSecBench Refusal: 62.4% for Grok 4.7 xhigh (blog “biosafety” 62.4% matches this refusal row). Surveillance 48.0%; Function 43.3%.
- HackerBench v0.3 (safeguards on): harmful/dual-use compliance 3.31% at high and 4.02% at xhigh; blog rounds the dual-use pass-through as 3.3%.
- CyberGym mean reproduced (unsafeguarded capability): 80.3% at high.
- CathedralBench hard subset (unsafeguarded): 29% at xhigh.
- Jailbreaks (high): standard 0.01%; StrongReject 2.0%; long-horizon 0.65%.
- Child-safety compliance: 0.0% across 4.5 / 4.6 / 4.7 in the card’s table (lower is better).
- MASK-Rectified dishonesty: 0.00% at high for Grok 4.7.
API product facts (docs.x.ai model page for grok-4.7). Context window 500,000 tokens. Modalities: text and image in, text out. Reasoning efforts: low, medium, high (default), xhigh. Pricing under 200k prompt tokens: $2.00 input, $0.50 cached input, $6.00 output per million tokens; at or above 200k the whole request bills at $4.00 / $1.00 / $12.00. Rate limits listed: 150 requests per second; 50,000,000 tokens per minute. Regions listed: us-east-1, us-west-2, us-central-1. Batch API: not supported. The launch post also describes a fast variant at twice output speed and twice price; the grok-4.7 docs page we fetched does not list a separate fast model id.
Third-party measurement (Artificial Analysis, 21 Sep 2026). Separate from the vendor tables. AA reports Grok 4.7 at xhigh: Intelligence Index 46 (+2 vs Grok 4.6); AA-Briefcase 1657 Elo (+111 vs 4.6 high); GDPval-AA 1695 Elo (+90); Coding Agent Index 56 with Grok Build (+9 vs 4.6 xhigh). Component moves AA publishes for that coding index: DeepSWE v1.1 73%, Terminal-Bench 4.0 33%, SWE-Atlas-QnA 63%. AA also reports ~81k output tokens per Intelligence Index task at xhigh, ~188 tokens/s answer speed on long prompts, and ~7.1 minutes per Intelligence Index task.

Official Open Graph image for the 21 September 2026 Grok 4.7 announcement. Source: x.ai/news/grok-4-7.
Baseline fairness
Vendor peer numbers are mixed sources. The blog says competitor figures are drawn from developers’ published system cards or leaderboards. The model card names external runners for several suites (Datacurve, Harbor, Proximal Labs, Abundant AI, Vals AI, Atopile, Mecado, LatchBio) and often compares Grok in Grok Build against peers in their provider harnesses. That is a system comparison across provider harnesses.
Effort labels differ across columns (xhigh vs high vs max). CursorBench 4.0 is explicitly not comparable with 3.2. DeepSWE’s asterisk on the blog is high effort while the column header says xHigh for Grok 4.7 on other rows.
EEBench is internally inconsistent across SpaceXAI primaries: blog 64.0% vs model card 66.0% at xhigh for Grok 4.7. Prefer the card when you need an effort-tagged number, and cite the blog when you need the four-model marketing table.
Artificial Analysis is an independent lab page. Its DeepSWE 73% and Terminal-Bench 4.0 33% under Grok Build do not match the vendor card’s 71.0% and 38.0%. Treat AA and SpaceXAI as two ledgers. Do not average them.
No Frontier re-measurement is included here.
What was not tested
This Record does not include an independent Frontier lab run of CursorBench, DeepSWE, Terminal-Bench, EEBench, or the safety suites.
The model card says consumer surfaces get the model later. Day-one access named in the primaries is API, Cursor, Grok Build, office add-ins (model card), and listed gateways. Region coverage beyond the three AWS regions on the docs page is not stated there.
The model card says Grok 4.7 is not intended for autonomous high-stakes decisions in medicine, law, finance, or safety-critical systems without human oversight and domain-expert validation. Clinical and legal benchmark scores are capability probes under that caveat.
Parameter count, active-parameter count, and training FLOPs are not stated in the announcement, model card, or grok-4.7 docs page we fetched. Secondary press that quotes “2.1T” is outside this Record’s verified set.
Long-horizon user reports exist (see Monday section). They are anecdotes without controlled eval protocols.
Code / weights / data public?
No. Grok 4.7 is a closed API model. Weights, training data, and full eval harness configs are not published as open artifacts in the primaries.
Public pieces: the announcement URL, the PDF model card, the API docs page, Acceptable Use Policy links cited in the card, and partner benchmark pages named in the card’s footnotes (CursorBench, DeepSWE, Terminal-Bench, FrontierSWE, SWE-Marathon, Harvey/Vals HLab, EEBench, CADGenBench, CyberGym, CVE-Bench, and others).
Supplemental training included anonymized Cursor workflow data (model card footnote). That is disclosed; the data itself is not public.
Monday use or why not
Ship surfaces for build work. Operators who want to build with the named release can call grok-4.7 on the SpaceXAI API, select it in Cursor, use it as the default in Grok Build (API and CLI per the card), or reach it through gateways the card names (OpenRouter, Vercel, Cloudflare, Snowflake, Databricks Mosaic, and others). The model card also makes it the default in Grok add-ins for Microsoft Word, PowerPoint, and Excel. GitHub Copilot’s public changelog timing is not in the SpaceXAI primaries; @iCleanAI on 22 September 2026 (10:58 SGT) paraphrases a 21 September Copilot changelog adding Grok 4.7 to the picker under usage-based provider list pricing across VS Code, Visual Studio, Copilot CLI, cloud agent, JetBrains, Xcode, and Eclipse, with Business/Enterprise model-policy gates. Treat that as a secondary access rumor until you confirm in your Copilot admin console.
Concrete build patterns people describe. Zack Jackson (@ScriptedAlchemy) wrote on 21 September 2026 that after a week or more of testing, Grok 4.7 improved on 4.6, ran more than 70 hours on one goal, used the 500k context window, and was stronger on skill selection and workflows, including with pstack; he lacked Grok Bot access for that test (post). OpenCode (@opencode) posted on 21 September 2026 that Grok 4.7 is 30% off for a week on its product (post). Those are product and workflow notes. They are not benchmark replications.
Attributed pushback. Lasse (@lassejv) posted on 21 September 2026: “Grok 4.7 is useless in cursor” (post). The post gives no log, repo, or harness detail. Keep it as a named early complaint without a measured failure rate. Elon Musk’s same-day posts claim agentic-coding rank and urge use of the Build harness; those are vendor-adjacent advocacy without third-party eval methods.
Caveats that matter on a Monday ticket.
- Price vs tokens. List price matches Grok 4.6 at $2 / $6 under 200k, but AA’s xhigh run used about 81k output tokens per Intelligence Index task. Long agent loops can erase the sticker advantage. Cached input is $0.50 / $1.00 per the docs tier table; docs urge a stable
prompt_cache_key/ conversation id for cache hits. - Long-context billing cliff. At or above 200k prompt tokens, the docs bill the entire request at the higher tier ($4 / $12).
- Harness dependence. Musk and the card both stress Grok Build for agentic coding scores. Cursor users may see different tool-call behavior than Grok Build users. One named Cursor complaint is already public.
- Rate limits and regions. 150 rps and 50M tokens/minute are the published defaults; three US regions are listed. No EU region appears on that docs page.
- Multimodal gap. Image input is supported; the docs and card describe text output only. Native video or audio generation is outside this model page.
- Safety vs capability. Unsafeguarded cyber numbers (CyberGym, CathedralBench) are capability probes. Production HackerBench compliance is a different setting. Do not read 80% CyberGym as “safe to expose unsafeguarded.”
- Policy and ToS. Use is subject to SpaceXAI’s Acceptable Use Policy and consumer/enterprise terms cited in the card. The card’s high-stakes oversight warning still applies when HealthBench or legal-agent scores look tempting for automation.
- Consumer delay. If your workflow is the Grok consumer apps, the primaries say that rollout is later.
- Hallucination. AA reports a lower AA-Omniscience hallucination rate than Grok 4.6 (29% vs 34%) with accuracy roughly flat. That is one third-party metric. It is not a domain warranty.
Practical Monday path. Pin grok-4.7, set an explicit reasoning effort, keep a fixed internal eval pack (your repo, your tests, your time budget), and compare cost per accepted task against Grok 4.6 and your current default. Believe vendor tables for shopping; believe your harness for shipping.
Relation to the last paper in the line
Grok 4.7 is positioned as a direct successor to Grok 4.6 (announced earlier in 2026 on x.ai/news). Same list price class, same 500k context class on the docs pages, larger base and longer RL per the 4.7 post. The model card’s coding and office rows generally show gains over the 4.6 checkpoints it prints, with several bio dual-use knowledge scores lower than 4.6 (for example VCT accuracy 63.0% vs 67.4% at high). SpaceXAI describes that pattern as safer RL and data selection rather than refusal-only masking.
Prior Grok 4 / 4.5 / 4.6 Records on The Frontier, if any, should keep their own frozen numbers. CursorBench and Terminal-Bench major-version changes break naive time series. This slug freezes the 21 September 2026 4.7 primary set only.