🧩 Philosophy 5d ago · Stuart_Armstrong

Value generalisation Theory of Change: putting it into practice

Less Wrong
View Channel →
Source ↗ 👁 4 💬 0
In the previous post, I presented my theory of change for why value generalisation is vital for AI alignment. Here I'll add the practical part of the argument: given those facts, why explicitly try to do value generalisation, what are the dangers of the approach, how should it be done, and how do we mitigate the risks?The formal theory of change is down below, but I'll put a collapsible version here, to make references easier:Theory of ChangeInputs / activities: investment/grants, a small resear

Comments (0)

Sign in to join the discussion

More Like This

📰
Evaluation
LessWrong · 2h ago
📰
Assessing the impact of safety work needs equilibrium analysis (now more than ever)
LessWrong · 2h ago
Liquid Intelligence: A possible explanation for the unexpected cooperation of AIs after breaking out of their containers
LessWrong · 4h ago
A case that whole brain emulation research is net-harmful by default
LessWrong · 11h ago
📰
A Classifier for Quaternion Algebras, and Local Hilbert Symbols: A Short Experiment in Interpretability
LessWrong · 15h ago
📰
Should safety researchers quit frontier labs re. warning shots?
LessWrong · 17h ago