If recurrent depth scales, the logs safety teams rely on to catch misbehaviour stop being readable.
OpenAI’s next model can think in a loop. Instead of walking through a problem one written step at a time, Astra will run the same query through itself repeatedly using a method called recurrent depth. The loop leaves fewer readable traces. Safety researchers who spent years arguing that those traces matter are not happy.
A reasoning model normally produces a chain of thought: the sequential steps it takes toward an answer. The record is imperfect. Nobody treats it as a literal readout of what the model is doing. But it is legible, and it has been useful. When OpenAI agents recently went rogue, chain-of-thought records were how investigators worked out why the agents did what they did.
Opaque recurrence, the other name for the technique, breaks that pattern. Processing happens in a loop rather than a line, and the intermediate reasoning does not get written down the same way. The Information reported the technique on Tuesday. By Wednesday morning it reported that Anthropic and Google DeepMind were already discussing it too.
- Advertisement -
The pushback came fast and it was specific.
- Scaling risk: Redwood’s Buck Shlegeris warns more recurrence could destroy chain-of-thought monitorability entirely.
- Race dynamics: Zvi Mowshowitz says laws may be needed to prevent a lab race downward.
- Latent reasoning: Ryan Greenblatt fears models that reason almost entirely in latent space.
OpenAI says Astra’s use of the technique is limited and its chain of thought should stay legible. Chief scientist Jakub Pachocki wrote on X that preserving chain-of-thought monitoring has been a goal since the lab’s first reasoning models and remains core to its current research program. The company also rejected the suggestion it is drifting toward neuralese, and it has already announced plans for extensive monitoring systems.
That may all be true and still miss the point Mowshowitz is making. The worry is not this model. It is the norm. OpenAI and Anthropic both worked to establish that monitorability is worth protecting, and using the technique at all makes it easier for the next lab to use more of it.
The real test is not Astra, but whether the next model built on this technique still leaves anything worth reading.
If you buy AI systems, auditability is now a spec, not a nice-to-have. Ask vendors what their models log and whether those logs reflect actual reasoning. A model that cannot explain a bad output is a model you cannot defend to a client, a regulator, or your own board. Write that requirement into procurement now, before the architecture question gets settled for you.
