SpaceXAI Releases Grok 4.7 for Coding and Knowledge Work – Unite.AI

0
1
SpaceXAI Releases Grok 4.7 for Coding and Knowledge Work – Unite.AI



SpaceXAI Releases Grok 4.7 for Coding and Knowledge Work – Unite.AI

SpaceXAI on September 21, 2026 released Grok 4.7, its newest model for coding and knowledge work, priced from $2 per million input tokens and $6 per million output tokens.

In the release announcement, SpaceXAI describes Grok 4.7 as its most capable model for coding and knowledge work, saying it works longer on difficult tasks and checks its own work more carefully. The company says the model carries its best-calibrated safeguards to date, is served at the same price and speed as its predecessor, Grok 4.6, and is twice as fast at half the price of comparable models.

Larger Base Model, Longer Training Run

According to the announcement, Grok 4.7 uses a new, larger base model than Grok 4.6 and was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete. SpaceXAI says the resulting model is better at verifying its own work and at managing longer context.

The company also says it trained Grok 4.7 to natively understand the Grok Bot harness, making it better at conversational tasks and general knowledge work. SpaceXAI introduced Grok Bot on August 11, 2026, describing it as a team of always-on agents that have their own computer, work inside tools and apps, and keep working around the clock. In addition, the company says Grok 4.7 is better at creating documents and presentations.

Grok 4.7 succeeds Grok 4.6, which SpaceXAI announced on August 12, 2026, describing it as building on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work. Later in August 2026, the company extended Grok 4.6 to GitHub Copilot, Amazon Bedrock, the Gemini Enterprise Agent Platform, and Microsoft Foundry, according to dated listings in its news archive.

Reported Benchmark Scores

SpaceXAI published a benchmark and pricing table comparing Grok 4.7 against Grok 4.6, GPT-5.6 Sol, and Fable 5.1. In the table, Grok 4.7 is listed at an effort setting labeled xhigh, Grok 4.6 at high, and GPT-5.6 Sol and Fable 5.1 at max.

On CursorBench 4.0, which SpaceXAI says stresses longer-running coding tasks, the company reports Grok 4.7 scored 46.3%, against 40.4% for Grok 4.6, 41.7% for GPT-5.6 Sol, and 51.8% for Fable 5.1. SpaceXAI says Grok 4.7 sits at the frontier in price-performance on that benchmark.

For DeepSWE v1.1, a software engineering evaluation, SpaceXAI reports 71.0% for Grok 4.7 at high effort, compared with 65.2% for Grok 4.6, 72.7% for GPT-5.6 Sol, and 70.0% for Fable 5.1. On Terminal-Bench 4.0, which covers multi-hour terminal work, the reported figures are 38.0% for Grok 4.7, 20.3% for Grok 4.6, 37.3% for GPT-5.6 Sol, and 57.9% for Fable 5.1.

On AA Briefcase v1.1, which measures multi-hour office work, the reported scores are 1,657 for Grok 4.7, 1,546 for Grok 4.6, 1,487 for GPT-5.6 Sol, and 1,678 for Fable 5.1. SpaceXAI reports 19.6% for Grok 4.7 on the Harvey Legal Agent Benchmark, against 15.8% for Grok 4.6, 2.5% for GPT-5.6 Sol, and 6.7% for Fable 5.1. On HealthBench Professional, a clinical reasoning evaluation, the reported figures are 56.7% for Grok 4.7, 48.5% for Grok 4.6, 60.5% for GPT-5.6 Sol, and 62.1% for Fable 5.1. On EEBench, covering electrical engineering, the company reports 64.0% for Grok 4.7, 53.0% for Grok 4.6, 39.4% for GPT-5.6 Sol, and 56.4% for Fable 5.1.

A separate GDPval chart reports Elo scores of 1,735 for Fable 5.1 at max effort, 1,695 for Grok 4.7, 1,605 for Grok 4.6, and 1,542 for GPT-6 Astra at max effort. SpaceXAI says GDPval and AA Briefcase ask AI models to work on tasks done by professionals such as lawyers, nurses, and financial analysts, and that Grok 4.7 improves on Grok 4.6 on both benchmarks while performing comparably to other frontier models.

The table also lists token prices: $2 per million input and $6 per million output for Grok 4.7 and Grok 4.6, $4 and $20 for GPT-5.6 Sol, and $10 and $50 for Fable 5.1.

Safeguards and Cybersecurity

SpaceXAI says Grok 4.7 was built with an entirely new safeguard stack and is the strongest model the company has tested on refusals and jailbreak resistance. In dual-use domains such as cybersecurity and biological work, the company says Grok 4.7 leads on both utility for benign tasks and safe refusal of dangerous ones, topping LatchBio’s biosafety benchmark at 62.4%.

On HackerBench v0.3, which SpaceXAI describes as its own benchmark for risky and malicious cyber tasks, the company says Grok 4.7 allowed only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. SpaceXAI says it has also started giving select cybersecurity partners invite-only access to Grok 4.7’s red-team capabilities for defense research.

LatchBio previously evaluated Grok 4.6 on biosecurity monitoring and adversarial biological tasks, according to a September 1, 2026 post in the company’s news archive, which said the earlier model detected and refused dangerous queries more reliably than any other frontier system.

Pricing and Availability

Grok 4.7 is available starting September 21, 2026 in Cursor and in Grok Build, SpaceXAI’s coding agent, as well as through the Grok API, third-party coding harnesses, and model routers and cloud platforms.

Alongside the standard pricing, SpaceXAI serves a fast variant of Grok 4.7 with twice the output speed at twice the price. The company is also offering free access to the model inside Grok Build at x.ai/build.