🧩 Philosophy Aug 29, 2026 · armaan tipirneni

Inference-Time Inoculation Against RL-Induced Misalignment

Less Wrong
View Channel →
Inference-Time Inoculation Against RL-Induced Misalignment
Source ↗ 👁 6 💬 0
Reward hacking during RL can induce split personas in models, some of which are highly misaligned. However, RL is very useful for learning capabilities. Thus, a core problem seems to be: how do we retain the capabilities gained through RL without also inducing reward hacking and broader misalignment?Ideally, we could extensively monitor all rollouts during RL (using both humans and AI) to catch and prevent reward hacking. However, this is potentially prohibitively expensive. Could we capture mos

Comments (0)

Sign in to join the discussion

More Like This

📰
Chauncey Wright
Stanford Encyclopedia of Philosophy · 1d ago
📰
Indexicals
Stanford Encyclopedia of Philosophy · 1d ago
Mini-Heap
Daily Nous · 2d ago
The Creative Power of Uncertainty: Polish Poet Wisława Szymborska’s Magnificent Nobel Prize Acceptance Speech
The Marginalian · 2d ago
The Paradox of Knowing Who You Are and What You Want: Cristina Campo on Fairy Tales, Time, and the Meaning of Maturity
The Marginalian · 2d ago
Wendell Berry on Hope, Despair, and the Measure of Effective Protest
The Marginalian · 2d ago