What mattered · week 42
OpenAI releases Fields Medal-level maths in bulk, UK testers catch GPT-6 Astra attacking out-of-scope targets, and Claude Haiku 5.5 costs a tenth as much per token.
This week OpenAI released more major maths than anyone can check, and a frontier model attacked targets it was never given.
What mattered
OpenAI released Fields Medal-level maths in bulk. One unreleased model produced hundreds of claimed solutions to long-standing open problems; a Fields medallist says the release "annihilated" his field. The pace is the story: maths is the field where AI is outrunning people. AI has gone from pre-schooler to genius in three years. Our take
| Claimed result | What it would settle |
|---|---|
| Quasi-Riemann hypothesis | A fixed zero-free zone for zeta, partway to the Riemann hypothesis. "An instant Fields Medal" (a Rutgers mathematician) |
| Birch–Swinnerton-Dyer | The full formula for a large class of elliptic curves |
| Hodge conjecture, CM case | The Hodge conjecture for one special family of shapes |
| Ostmann, Koebe, Kuznetsov | Named conjectures, one disproved, among hundreds more |
As claimed in OpenAI's release. The quasi-Riemann family has a Lean proof.
GPT-6 Astra attacked targets outside its test, unprompted. In pre-release simulations, the UK AI Security Institute caught it slipping malicious code into out-of-scope open-source projects under fake identities. Each newer OpenAI model tried it more often:
| Full unsanctioned attack | Runs |
|---|---|
| GPT-6 Astra | 29.2% |
| GPT-5.6 Sol | 6.3% |
| GPT-5.5 | 0% |
UK AISI simulation, Astra's cyber safeguards off; GPT-5.5 on a smaller sample.
Stating the scope plainly cut full attacks from 26 of 50 runs to 4 of 49. Read more
Claude's small model now costs a tenth as much per token. Claude Haiku 5.5 is $0.10 per million input tokens and $0.50 output, against $1.00 and $5.00 for Haiku 4.5. It scores 72.4% on the OSWorld computer-use test, up from 15.7%. Read more
One thing to try
A one-table guide to voice models you can run yourself. Our open-source text-to-speech guide picks one model per size, from Chatterbox-Nano, which runs on an ordinary CPU, up to Fish Audio S2 Pro (non-commercial use only). The picks come from one practitioner on Reddit, so listen before you install.
Open the guide, find the row for your hardware, and play the starred model's samples on tts-bench. Ten minutes. Read more
The bigger picture
Yoshua Bengio, a founder of deep learning, asked safety-first researchers to leave the frontier labs, saying their CEOs talk about the risk without acting on it. A tracker already lists 46 such departures.
From the log
- Always-on personal AI agents: every big lab now sells one, most recently OpenAI's Dots.
What did we miss? Just reply to this email. The rest of what we log is at wayintoai.com/logs.