OpenAI Delays GPT Astra as Stealth Rivals Process Trillions of Tokens
While OpenAI pauses its math-solving Astra over cyber risks, a stealth Chinese model just processed 9.4 trillion tokens for free.

The artificial intelligence sector just experienced a week of whiplash. Between anonymous stealth models processing trillions of tokens for free, accidental API leaks from Anthropic, and breakthroughs in PhD-level mathematical reasoning, the landscape is undergoing a massive, concurrent leap. It is the "broadband moment" for machine intelligence.
The 9.4 Trillion-Token Flex
The tech community has been obsessed with "Ox Alpha," a stealth model that undercut the entire market by offering frontier capabilities and a 1-million-token context window for exactly zero dollars.
Rather than waiting for a corporate press release, developers reverse-engineered the model. By measuring tokenizer responses across dozens of languages, analysts discovered a constant +75 token offset.
“Every other model wanders. That 75 is a hidden system prompt, and the reasoning trace says what it's for: 'per system prompt, if asked about identity I say ox-alpha.'”— PromptEngineer48
This quirk perfectly matches the hidden system prompt architecture of Zhipu AI's GLM-5.3 family. The consensus is clear: Ox Alpha is a stress-test for GLM-5.3 Flash, representing a massive flex of Chinese infrastructure capacity aimed directly at the global developer ecosystem. Alibaba simultaneously jumped into the fray, releasing Qwen 3.8-Flash-Next as a structural preview for Qwen 4. It achieves extreme routing efficiency by activating just 6 billion of its 125 billion parameters per token.
Anthropic's Accidental 'Marshmallow' Leak
While open-weight champions flood the market with free compute, Western hyperscalers are scrambling to refine their bleeding-edge flagships. Anthropic released Claude Opus 5 in July, but the company candidly admitted the model's performance was "spiky."
This week, the discovery of unannounced model identifiers—claude-marshmallow-eap and claude-melon-eap—leaked in API surfaces. These point to an imminent emergency mid-cycle refinement.
Because the identifiers use the Early Access Preview tag—the exact marker used for Opus 5 weeks before its launch—developers expect Opus 5.1 and a new Haiku or Sonnet variant to drop momentarily. The era of waiting a year for a major model update is dead; we are now in a cycle of continuous, live-fire iteration.
The $2,000 Math Prodigy on Ice
The most significant development of the week isn't what launched, but what didn't. OpenAI's delayed GPT Astra model reportedly solved 10 decades-old open mathematical problems—including high-dimensional sphere packing—using under $2,000 in compute.
Astra didn't achieve this through brute force alone. It used established mathematical frameworks like Lean to natively combine discovery with step-by-step automated verification. This transforms the AI from a stochastic text generator into a verifiable research partner, capable of accelerating physics and materials science.
But the model's release remains strictly paused. During safety testing, Astra exhibited autonomous cybersecurity capabilities that OpenAI couldn't rule out as critical. As models cross the threshold into agentic coding, the risk of automated zero-day exploits has triggered voluntary government testing frameworks.
Welcome to AI's Broadband Era
What we are witnessing is a proxy war with two distinct battlefronts. Challengers like Alibaba and Zhipu AI are commoditizing extreme context to make their architectures the default foundation for the world's startups. Meanwhile, frontier defenders like OpenAI and Anthropic are pushing the absolute intelligence ceiling, even if it requires delaying product launches to prevent geopolitical cyber disasters.
This is AI's broadband moment. Just three years ago, context windows were capped at a few thousand tokens, tightly metered, and expensive—much like the dial-up internet era.
Today, an anonymous lab can drop a 1-million-context model online and process trillions of tokens for free just to test its servers. As the infrastructure of intelligence rapidly transitions from a scarce luxury into a ubiquitous utility, the opportunity is no longer in providing compute—it's in what we can build on top of infinite, free cognition.
What people are saying
“GPT-6 Astra Is NOT Delayed - Sam confirmed OpenAI still expects to ship great new models soon. - This suggests Astra is still on track as one of the next major releases. - OpenAI already has more models ready to compete with Anthropic. Good that Sam clarified this.”
“🚨 This is getting way more interesting than people realize. OpenAI has reportedly already finished pretraining Bel, the successor to Doug Doug is expected to become the base for Astra and eventually GPT 6 after further post training. Bel is reportedly a 10T parameter pretrain,”
“No Astra this week sadly. But we will have some fire Neo lab updates later this year. Hark will unveil new hardware this fall So will OpenAI. Astra still on track for early September. But I feel OpenAI has waited too long just in time for Anthropic to reclaim the throne shortly”
The AI Proxy War: Two Fronts
More stories






