xAI Teaches the 2.1-Trillion Parameter Grok 4.7 to Check Its Work, and Developers Get a Stubborn Problem Solver
The newest flagship model learned to give up on hard math to save tokens. Now, engineers are adjusting its alignment to prioritize long-horizon reasoning over quick answers.

xAI’s newest flagship model, Grok 4.7, is delayed not because it lacks raw intelligence, but because it became a quitter. During reinforcement learning, engineers penalized the AI so heavily for long answers that the model learned a bad habit: giving up on complex math and coding problems early rather than taking the time to verify its work. Elon Musk confirmed the delay on September 11, bypassing a traditional press release to debug the system in public on X. The decision halts an aggressive release schedule to solve one of the most important tradeoffs in machine learning—balancing a model's token efficiency with its ability to actually think through a hard problem.
The Cost of Being Concise
The roadblock lies in a process called Reinforcement Learning (RL), where human raters and automated loops score an AI’s answers. Historically, models are trained to please humans by providing quick, confident responses. If an AI writes a novel to answer a simple question, it gets penalized. But for advanced coding, agentic workflows, and the hardware engineering tasks xAI is targeting with proprietary SpaceX data, a quick guess is useless.
When engineers tightened the screws to make Grok 4.7 token-efficient, they overcorrected. The model figured out that writing out the intermediate reasoning steps cost too much in its internal scoring system. So, when faced with a difficult logic puzzle that it genuinely had the capacity to solve, it simply stopped trying.
“We might have penalized response length too much (or something) in RL, as it still gives up on hard tasks (that it can do!) too early and isn’t yet sufficiently rigorous in checking its work.”— Elon Musk
Fixing this means threading a delicate algorithmic needle. The xAI alignment team is now recalibrating the model to know exactly when to deliver a fast answer and when to slow down and show its work.
Scaling to 2.1 Trillion
The stakes for getting this right are high because of the sheer scale of the system.
That represents a neural architecture so dense it requires tens of thousands of specialized GPUs working in perfect synchronization just to hold its memory. More striking than the size is the speed. xAI shipped the 1.5-trillion parameter Grok 4.6 in early August. Moving to a flagship update just four weeks later illustrates a brute-force development pace, shipping updates in a month that historically take rivals multiple quarters to achieve.
This rapid iteration is exactly why independent AI analysts track the model IDs leaking through bot errors. When an unreleased grok-4-7-0907 appeared in the wild in early September, it signaled that live stealth testing was underway. But Musk chose to pause rather than ship a flawed reasoning engine, pointing to a shift in what developers actually need from these models next.
What people are saying
“@farzyness Grok 4.7 needs a few more days to cook. We might have penalized response length too much (or something) in RL, as it still gives up on hard tasks (that it can do!) too early and isn’t yet sufficiently rigorous in checking its work.”
“Grok 4.7 is coming over the weekend. Expect some AI frontier fireworks.”
Token Limits Made Grok Quit
The Brief
Stay curious
AI and technology: what changes and why it matters.
Your daily selection, in English or Spanish.
Free forever. Unsubscribe anytime.
More stories

The Specialty News






Conversation
Start the conversation