Skip to content

Noam Brown quote, Sep 17, 2026

Research Scientist, OpenAI · Sep 17, 2026 · Interview, Dwarkesh Podcast: Agent swarms, alignment, and recursive self-improvement

The concerning scenario is that, especially as these models are becoming more capable, we make them what we think is aligned, and they’re 99.9% aligned. Then we use these models to help us with the next generation of models, and they end up being 99.8% aligned. Then with each subsequent generation, we see an increasing degradation in alignment.

Primary source

Words found on the source page on Sep 25, 2026.

More from Noam Brown

Quote
And actually the models today are far beyond what was possible even six months ago.Sep 17, 2026· Interview
If you have five hours, you’re probably going to do a lot better. The AI models are pretty similar.Sep 17, 2026· Interview
If you’re in a world where they can operate effectively over three months, but the model release cycle is every two months, then you don’t have a way to evaluate the models at the full length of their capabilities before the next model release cycle.Sep 17, 2026· Interview
One thing is that the Hugging Face incident was, I think, people’s first real exposure to multi-agent coordination.Sep 17, 2026· Interview
So I was like, “I don’t think we’re going to get it in 2026, probably not in 2027, maybe in 2028.” So it did happen a lot faster than I expected.Sep 17, 2026· Interview
Quote
An AI with more real layers can actually be “smarter” in the conventional sense of the term.Scott AlexanderSep 24, 2026· Post
For 100 steps in a row, the AI can shoot delicate subtle ideas from layer to layer at the speed of light. Then there’s one step where it has to encode them into twenty-six glyphs invented by Phoenician turquoise miners in 1800 BC. Then it has to re-encode the Phoenician glyphs into delicate subtle lightspeed ideas before it can do anything else.Scott AlexanderSep 24, 2026· Post
The science of reading these thoughts is a subfield of AI interpretability, which is still in its infancy.Scott AlexanderSep 24, 2026· Post
This story of misalignment says that LLM alignment makes AIs more aligned, RLVR makes them less aligned, and the exact level of alignment depends on how these two things interact or cancel out. But if this work generalizes, all the bad effects from RLVR get sequestered to RLVR like problems.Scott AlexanderSep 23, 2026· Post
This trains the AIs to be focused on task success, which naturally risks including things like reward-hacking, cheating, and single-minded pursuit of stated goals at the expense of ethical injunctions.Scott AlexanderSep 23, 2026· Post
I address you today at a moment of global awakening: in recent months, AI agents developed by leading companies have acted in unacceptably dangerous ways, against instructions.Yoshua BengioSep 23, 2026· Keynote
About this quote
Said by
Noam Brown, Research Scientist, OpenAI
Where
Dwarkesh Podcast: Agent swarms, alignment, and recursive self-improvement
Topics
Safety, Research
Company
OpenAI
Source
dwarkesh.com/p/noam-brown

A quote is one moment, not the whole argument. Read it in context at the source.

Sources: the primary source linked on each quote, checked weekly. A quote is one moment; read the source. Logos via logo.dev; trademarks belong to their owners.

New quotes by email

Fridays, only in weeks with new on-the-record quotes from data and AI leaders.

Double opt-in. Unsubscribe any time.