OpenAI Skipped the Benchmarks and Had Its Next Model Solve Ten Math Problems Nobody Could Crack
Astra shipped Lean-verified proofs for decade-old open problems, and the whole point was that you don’t have to take OpenAI’s word for it. A Fields Medalist already checked.
A huge AI announcement came this weekend with a receipt you can run yourself.
🧮 The Model That Brought Proof, Not Promises
For three years every AI lab has claimed a math result, and almost every one burned out inside a week of people arguing about what actually happened. The company making the claim was always the only party that could check it. Benchmarks got gamed. Demos got curated. The outside world was left squinting at a number.
OpenAI just changed the register. On Friday, August 1, it announced that an internal version of Astra, the model it’s calling its next major release, produced new results on ten open problems in mathematics and theoretical computer science. Each had been open for at least a decade. The headline result is the first explicit construction of a non-sofic group, a question that has stood in group theory since 1999.
Here’s the part that matters. OpenAI shipped every proof as a Lean 4 machine-checkable certificate, published to GitHub alongside a 249-page manuscript and the model’s reasoning walkthroughs. Lean’s trusted kernel gives a binary verdict. The proof compiles or it doesn’t. No PhD required, no year-long peer review backlog, no panel of experts vouching for it. Anyone can run the check.
Translation: OpenAI stopped asking to be believed and started letting people verify.
The community reaction was warmer than the usual eye-roll. Fields Medalist Timothy Gowers said he’d recommend one of the proofs to a top journal without hesitation. Thomas Bloom, who curates the Erdős problem catalogue, called it big news. And the whole thing reportedly cost about $2,000 in compute.
Why the receipt is the story
Remember October 2025. OpenAI claimed GPT-5 had solved ten Erdős problems, and Bloom, the same person praising Astra now, tore it apart. The model hadn’t proved anything. It had retrieved existing solutions from literature he hadn’t catalogued yet. He called it a dramatic misrepresentation. DeepMind’s Demis Hassabis called it embarrassing. The VP who announced it was gone by April.
That’s the context that makes this launch land. OpenAI didn’t just deliver a capability. It delivered the exact thing it got caught faking last year, and it delivered it in the one format nobody can dispute. That’s not a demo. That’s a company that learned what a benchmark is worth in 2026, which is nothing.
The timing isn’t an accident either. OpenAI is in the middle of an IPO process targeting a valuation near a trillion dollars. Debuting your next model through verifiable frontier discovery instead of a self-reported score is the strongest possible signal to the market that the technology is real. This is the demo you run when investors are watching.
🔓 The Weekend’s Uglier AI Story
On the other end of the capability conversation, security researchers at Palo Alto’s Unit 42 detailed a Zhuhai-based actor who wired an open-source DeepSeek-powered agent into an attack framework and pointed it at more than 460 internet-facing systems. The agent enumerated targets, pulled public exploits, and ran the operation over Telegram.
That’s the flip side of the same coin Astra sits on. When a model gets good enough to do original math for $2,000, an open-weight version gets good enough to automate a real attack for less. The capability doesn’t care what you point it at.
📅 What’s Coming
SpaceX reports its first-ever quarterly earnings as a public company Tuesday, the first hard look at Starlink’s actual numbers. AMD reports Wednesday into a market that just spent a week rewarding proven AI revenue and punishing the rest. And the White House AI framework, plus Astra’s public release, could land any day now.
Because that’s the pattern this week. The models keep getting more capable, the proof keeps getting harder to argue with, and the stakes keep climbing on both sides of what the capability gets used for.
This is the week AI stopped grading its own homework and started handing you the answer key.
— The Bandicoots 📱🔌


