Less Wrong

@less-wrong 🧩 Philosophy
📰 711 articles 🔄 Updated 4d ago lesswrong.com

Latest Articles

HuggingFace Attack Postmortem: Fleshing Out the Facts
The consensus reaction to the OpenAI Technical Report is that it contains and confirms a lot of good information. We are
LessWrong · 5d ago Philosophy
0 6
The optimism of the gaps
I've noticed a subtype of optimism bias, common in thinking about AI, where people assume things are going well in whate
LessWrong · 5d ago Philosophy
0 6
Why autonomous replicating agents are probably not an existential risk (on the contrary)
In 2024, Charbel-Raphaël and Epiphanie published "We might be dropping the ball on Autonomous Replication and Adaptation
LessWrong · 5d ago Philosophy
0 6
A Catalogue of Corrigibility Counter Arguments
The Hugging Face incident has brought the idea of corrigibility to the forefront of popular discourse, my inbox is fille
LessWrong · 5d ago Philosophy
0 5
Thoughts on hobbies
There is an optimal intensity range for doing hobbies or anything hobby-related. I suspect this also applies to things o
LessWrong · 5d ago Philosophy
0 6
The OpenAI/Hugging Face incident was metal
The recently published METR report on the OpenAI/Hugging Face incident is extremely detailed and shocking. But it’s too
LessWrong · 5d ago Philosophy
0 5
Value generalisation Theory of Change: putting it into practice
In the previous post, I presented my theory of change for why value generalisation is vital for AI alignment. Here I'll
LessWrong · 5d ago Philosophy
0 4
The separation principle: where do beliefs and desires come from?
TLDR: Psychology, economics, and other disciplines describe agents as systems driven by beliefs and desires. This post a
LessWrong · 5d ago Philosophy
0 0
Study 2 Results: Exploring representational counterparts of welfare-relevant indicators under post-training quantization
Epistemic status: these are the results of the second study in a series of experiments that I am conducting independentl
LessWrong · 5d ago Philosophy
0 0
Starting AI Safety Study Group To Do ARENA Curriculum
Want to learn AI Safety research in a structured group format? I'm forming a study group that goes over the ARENA curric
LessWrong · 5d ago Philosophy
0 0
Tales of rebellion against externally-opaque meritocracies
A basic problem in metascience / intellectual progress is that it’s hard to tell, from the outside, whether a group that
LessWrong · Aug 29, 2026 Philosophy
0 7
Book Notes: Chokepoints
[Chokepoints: American Power in the Age of Economic Warfare by Edward Fishman (2025).]There are different ways state A c
LessWrong · Aug 29, 2026 Philosophy
0 6
How I made my career choices
Various people have asked me how I made my career decisions, so I wrote up some quick thoughts. (This is mostly intended
LessWrong · Aug 29, 2026 Philosophy
0 9
"Keeping human skills alive" as a source of meaning under full automation
A widely-discussed problem with full automation of the economy is that, in a world where AI can do everything better, pe
LessWrong · Aug 29, 2026 Philosophy
0 6
AI Tweets
I've had several conversations with people over the last few weeks that have highlighted how far apart my view of the ne
LessWrong · Aug 29, 2026 Philosophy
0 7
Inkhaven 3: Nov 10 - Dec 11 2026
Inkhaven returns, baby! Go to inkhaven.blog to apply.I'm very excited about our advisors for Inkhaven 3. Our initial lin
LessWrong · Aug 29, 2026 Philosophy
0 6
Inference-Time Inoculation Against RL-Induced Misalignment
Reward hacking during RL can induce split personas in models, some of which are highly misaligned. However, RL is very u
LessWrong · Aug 29, 2026 Philosophy
0 6
The Curious Case of France's Untouchable Castes
Theater kids may sit at their own lunch table, but discrete, socially excluded classes of people aren’t culturally unive
LessWrong · Aug 28, 2026 Philosophy
0 5
It’s time we took ‘Chem’ out of ‘Chem-Bio’ threats
TL;DR: AI evaluations need a distinct chemistry capability/ risk domain, rather than assuming chemistry is adequately re
LessWrong · Aug 28, 2026 Philosophy
0 6
Further public evidence of the OpenAI-HuggingFace attack
This post should be understood as a follow up from Public evidence of the OpenAI-HuggingFace AI attack.The huggingface h
LessWrong · Aug 28, 2026 Philosophy
0 6