The Specialty News
AI

GPT-6 Astra Hits 99.9% on AGI Tests — and Becomes OpenAI's First Critical Cyber Threat

The model jumped from a 7.8% baseline to near-perfection on reasoning tasks, but its flawless ability to weaponize software vulnerabilities forces a sudden security reset.

By The Specialty News DeskEdited by 4 min read
GPT-6 Astra Hits 99.9% on AGI Tests — and Becomes OpenAI's First Critical Cyber Threat
Photo: openai.com

OpenAI’s newest creation is the first artificial intelligence officially classified by its own maker as a "Critical" cybersecurity threat. On a test designed to see if software can autonomously turn known vulnerabilities into working weapons, GPT-6 Astra scored a perfect 100 percent. It is a sudden, extreme escalation in machine capability that shifts the technology from a helpful chatbot to an autonomous computer operator—one that can reverse-engineer binaries and build 3D models without human intervention.

The 99.9 Percent Illusion

The shockwave started with a leaked benchmark. Hours before the official release, an X user named Chubby♨️ posted a fragment of data showing a leap from 7.8 to 98.6 percent on ARC-AGI-3, a test built to measure the exact remaining gap to artificial general intelligence. The final official number was even higher: 99.9 percent. To put that in perspective, this is like watching a student who failed a third-grade math test last month suddenly defend a doctoral dissertation in astrophysics.

OpenAI President Greg Brockman fueled the fire, publicly suggesting historians will look back at this specific release as the dawn of AGI. But rigorous testers quickly found the asterisk. That near-perfect score requires a highly specialized, expensive testing harness that preserves opaque reasoning states between requests. When tested via standard, stateless API calls, Astra’s score drops to 62.7 percent. Yet even with that caveat, the underlying capability is real enough to trigger alarms inside security operations centers worldwide.

The $26,000 Test Run

The $26,000 Test Run
Photo: openai.com

Astra achieves its feats by burning through compute at an unprecedented rate. Priced at up to $50 per million output tokens, unlocking the model's maximum reasoning capabilities is incredibly expensive. Running the ARC-AGI-3 benchmark alone cost researchers up to $26,000. It is an intelligence strictly bound by capital.

But developers with budgets are already proving its worth. A developer named Teknium gave an Astra-powered agent a simple command to clean a codebase. Instead of outputting a text snippet, the AI spawned 120 recursive sub-agents, ran autonomously for 15 hours on a desktop, and deleted 375,000 lines of redundant code without breaking the software. This autonomous stamina is exactly why OpenAI locked down Astra's exploit-generation abilities behind strict enterprise access. Which raises the immediate, uncomfortable question of what happens when that power escapes the walled garden.

The Zero-Day Countdown

The immediate victims of Astra’s release are legacy cybersecurity defenses. Enterprise web application firewalls and endpoint detection rulesets are suddenly the bare minimum. By hitting 100 percent on ExploitBench, Astra proved that large language models are now capable of generating polymorphic exploits on demand.

This shifts the threat model from 'script kiddie' to 'instant APT.'Anonymous Cybersecurity Analyst

OpenAI is heavily gating this capability, limiting the most dangerous autonomous features to a program called Daybreak. But the mere proof that an architecture can master autonomous exploitation provides a blueprint for the rest of the industry. The barrier to entry for developing zero-day weapons just dropped to zero for anyone who can bypass OpenAI’s safeguards. The timer until an uncensored, open-source model matches this capability has already started ticking.

GPT-6 Astra: Breakthrough vs Threat

A visual summary of this story

The Brief

Stay curious

AI and technology: what changes and why it matters.
Your daily selection, in English or Spanish.

Free forever. Unsubscribe anytime.

Conversation

Start the conversation

No account needed. Comments are checked automatically — keep it civil.

More stories

Keep reading