
Anthropic's Natural Language Autoencoders: Reading a Model's Mind in English
Quick verdict
Anthropic released Natural Language Autoencoders (NLAs), a method that takes the internal activations of a language model at a given layer and decodes them back into human-readable English. Not features. Not probes. Sentences. The technique surfaced planning behavior researchers did not expect, and along the way it caught a real bug in Anthropic's training pipeline. It is also clearly not a finished tool. Ryan Greenblatt's early tests showed it failing to recover what should have been "internal chain of thought" on single-pass math problems.
For anyone tracking interpretability, NLAs are the most concrete step toward "just ask the model what it was thinking" since sparse autoencoders, and they come with a public release on Neuronpedia.
What an NLA actually does
The autoencoder is trained to compress an activation vector into a short natural-language description and then reconstruct the original activation from that text. If the reconstruction is faithful, the text is a real handle on what the activation encoded, written in English you can read.
Compare that to what came before:
| Method | Output | Trade-off |
|---|---|---|
| Linear probes | Yes/no for a hypothesis you supplied | You have to guess the right question |
| Sparse autoencoders | Thousands of atomic features | "Feature shattering" — meaning splits across many features |
| Natural Language Autoencoders | A sentence describing what the activation means | Only as faithful as the reconstruction loss allows |
The pitch is that NLAs sit at a more useful abstraction level than SAEs. You can read the output without writing a paper to interpret it.
What NLAs caught
Planning behavior. NLA decodes from intermediate layers showed the model holding what looked like multi-step plans in its activations before generating tokens. Not just predicting the next word, but staging future decisions. This kind of result has been hypothesized for years from circuit-level interpretability work. NLAs let researchers see it in plain text.
A training pipeline bug. While inspecting activations across training runs, the team found instances where the internal representation suggested the model was processing translation artifacts. Something in the data pipeline was injecting translation behavior where it should not have been. They fixed it. This is the kind of bug a sparse autoencoder might surface as "feature 47831 sometimes fires," useless until someone interprets it. NLAs surfaced it as a sentence.
Where NLAs break
Ryan Greenblatt ran a sharp test: take single-forward-pass math problems where you know the model is doing real computation inside one pass. If NLAs reveal "internal chain of thought," they should recover the intermediate arithmetic steps. They did not.
That is a meaningful negative result. It suggests NLAs may be picking up readable, high-level semantic content while missing the lower-level computational structure that drives specific outputs. Or that the activations sampled were from the wrong layers. Either way, it is a ceiling on what "reading a model's mind" currently means.
Miles Brundage's framing is the right one: NLAs complement probes and dictionary learning, they do not replace them. Anyone selling NLAs as full interpretability is overselling.
Why it matters for safety
Most interpretability work has the same dead end: even when you find a feature or a circuit, translating it into a sentence a policy team or auditor can act on takes another paper. NLAs flip the output format. The text comes out in English. An evaluator can read it.
That matters for monitoring deployed models. If you can run an NLA over activations during inference and surface a readable summary of what the model is internally tracking, you have a new auditing layer that does not depend on the model's self-report (which we already know is unreliable).
The catch is faithfulness. If the NLA decoder is wrong about what the activation means, you now have a credible-sounding sentence that lies about what the model is doing. The reconstruction-loss metric is what stands between this being a real safety tool and a generator of confident interpretability hallucinations.
Video: interpretability research overview
For background on how interpretability has evolved from probes to SAEs to NLAs, this overview covers the lineage:
The bigger pattern
NLAs landed in the same week as Goodfire's neural geometry agenda, which argues that networks "think in shapes" and that manifolds, not features, are the right primitive for steering model behavior. Goodfire's pitch is essentially that SAE-style feature shattering breaks down meaning that lives at the manifold level.
Two different bets:
- Anthropic NLAs: describe activations in the language humans already use.
- Goodfire manifolds: describe activations in the geometric structure the network actually uses.
Both are reactions to the same frustration: sparse autoencoders gave us a lot of features and not enough understanding. The next year of interpretability work is going to be these two threads tested against the same problems to see which produces actionable findings faster.
FAQ
Can I try NLAs myself?
Yes. Anthropic released open-model NLAs on Neuronpedia. You can browse decoded activations from public models.
Do NLAs work on Claude or only open models?
The public release is on open models. Anthropic's internal use on its own models is research-only.
Is this related to Claude's chain-of-thought?
Sort of. CoT is the model's output reasoning. NLAs look at internal activations whether or not the model is producing CoT. Greenblatt's test specifically asked whether NLAs could recover internal reasoning on single-pass problems where there is no visible CoT — and the answer was no, at least for the activations sampled.
Does this make Claude safer to use?
Not directly today. It is a research tool that may eventually feed into safer deployment, similar to how visibility into thinking depth only becomes useful once it is wired into product decisions.
Sources
- @AnthropicAI - Natural Language Autoencoders announcement
- @mlpowered - NLAs complement probing and dictionary learning; planning behavior and training-pipeline bug
- @RyanPGreenblatt - NLAs did not recover internal CoT on single-forward-pass math
- @GoodfireAI - Neural geometry research agenda (related)
- Neuronpedia - Open-model NLA explorer
Further Reading
Try all the models mentioned in this article
Admix gives you GPT-5, Claude, Gemini, and 350+ AI models in one app. Compare responses side by side. Free to start.
Start free on Admix