Skip to content

Scott Alexander quote, Sep 23, 2026

Author, Astral Codex Ten · Sep 23, 2026 · Post, Mysteries Of AI Generalization

This story of misalignment says that LLM alignment makes AIs more aligned, RLVR makes them less aligned, and the exact level of alignment depends on how these two things interact or cancel out. But if this work generalizes, all the bad effects from RLVR get sequestered to RLVR like problems.

Primary source

Words found on the source page on Sep 25, 2026.

More from Scott Alexander

About this quote
Said by
Scott Alexander, Author, Astral Codex Ten
Where
Mysteries Of AI Generalization
Topics
Safety
Source
astralcodexten.com/p/mysteries-of-ai-generalization

A quote is one moment, not the whole argument. Read it in context at the source.

Sources: the primary source linked on each quote, checked weekly. A quote is one moment; read the source. Logos via logo.dev; trademarks belong to their owners.

New quotes by email

Fridays, only in weeks with new on-the-record quotes from data and AI leaders.

Double opt-in. Unsubscribe any time.