Noam Brown quote, Sep 17, 2026
Research Scientist, OpenAI · Sep 17, 2026 · Interview, Dwarkesh Podcast: Agent swarms, alignment, and recursive self-improvement
The concerning scenario is that, especially as these models are becoming more capable, we make them what we think is aligned, and they’re 99.9% aligned. Then we use these models to help us with the next generation of models, and they end up being 99.8% aligned. Then with each subsequent generation, we see an increasing degradation in alignment.
Words found on the source page on Sep 25, 2026.