Latest Articles
HuggingFace Attack Postmortem: Fleshing Out the Facts
The consensus reaction to the OpenAI Technical Report is that it contains and confirms a lot of good information. We are grateful to have it, and we are grateful for those who worked hard on it.
Alas, it sidesteps the biggest questions. There is much more we need to know.
The consensus reaction to the METR Report on the HuggingFace attack is: Holy shit.
Liv Boeree: My mind is legit blown.
Aella: this feels like a turning point. If this doesn’t cause large-scale coordination to pause frontier dev
0
6
The optimism of the gaps
I've noticed a subtype of optimism bias, common in thinking about AI, where people assume things are going well in whatever areas they're not looking closely at. If you're paying attention to technical alignment, you might think "well, things aren't looking great here, but at least there are people at the labs who are on top of things and won't actually send us off a cliff." If you interact with people at the labs, you might think "well, the labs don't seem to understand the problem and are happ
0
6
Why autonomous replicating agents are probably not an existential risk (on the contrary)
In 2024, Charbel-Raphaël and Epiphanie published "We might be dropping the ball on Autonomous Replication and Adaptation", making the case that "Once there is an open-source ARA model or a leak of a model capable of generating enough money for its survival and reproduction and able to adapt to avoid detection and shutdown, it will be probably too late". It received a substantive reply by Richard Ngo, notably"The key issue is that AIs that do ARA will need to be operating at the fringes of human
0
6
A Catalogue of Corrigibility Counter Arguments
The Hugging Face incident has brought the idea of corrigibility to the forefront of popular discourse, my inbox is filled with newsletters about Genie Coefficients and Safe Scaffolds.I think now is a particularly good time to make sure I have a good understanding of its limitations and dangers, particularly because I agree that it's the best path forward we have available to us. This post is my attempt to dive into the discussion and make sure I fully understand what's going on. Please correct m
0
5
Thoughts on hobbies
There is an optimal intensity range for doing hobbies or anything hobby-related. I suspect this also applies to things outside of hobbies.For example, if you’re reading a book, there’s a speed that’s too slow: you don’t get immersed enough into the book, you lose track of what it was that you read about last time you read the book, and it’s a sort of a protracted pain, in which the very act of reading the book feels like a chore.Then on the other end, there’s reading a book too fast. You blaze t
0
6
The OpenAI/Hugging Face incident was metal
The recently published METR report on the OpenAI/Hugging Face incident is extremely detailed and shocking. But it’s too long. Intelligent public distillations (like Zvi’s) are also too long, while tolerable summaries lose the meat.This is the report you want. Only the action, in chronological order. Let’s go.Quotes are from message board posts or agent thinking/transcripts. Events are tagged as metal, dank, dope, or baller as their virtues dictate.June 26 — July 7 PreludeJuly 8July 9July 10 — 13
0
5
Value generalisation Theory of Change: putting it into practice
In the previous post, I presented my theory of change for why value generalisation is vital for AI alignment. Here I'll add the practical part of the argument: given those facts, why explicitly try to do value generalisation, what are the dangers of the approach, how should it be done, and how do we mitigate the risks?The formal theory of change is down below, but I'll put a collapsible version here, to make references easier:Theory of ChangeInputs / activities: investment/grants, a small resear
0
4
The separation principle: where do beliefs and desires come from?
TLDR: Psychology, economics, and other disciplines describe agents as systems driven by beliefs and desires. This post argues that the belief-desire view can be derived from classic theorems from optimal control and reinforcement learning. This suggests seeing beliefs and desires as properties of optimal policies rather than as assumptions from folk psychology.IntroductionOne way to think about agents is as "systems that act for reasons".[1] This compact statement can be interpreted as encapsula
0
0
Study 2 Results: Exploring representational counterparts of welfare-relevant indicators under post-training quantization
Epistemic status: these are the results of the second study in a series of experiments that I am conducting independently, originally described in "Does post-training quantization change welfare-relevant indicators in open-weight language models?"; it was registered prior to data collection in an earlier post, which contains some important context that will not be fully repeated here. Some of this work involves speculation and I have tried to be clear about distinguishing between that and the re
0
0
Starting AI Safety Study Group To Do ARENA Curriculum
Want to learn AI Safety research in a structured group format? I'm forming a study group that goes over the ARENA curriculum. If you’re interested, fill out this form! Should take ~5 minutes. The purpose of the form is to get people's availability, level of experience, time commitment, and expectation about group size. With this information, we can create a good program that works for everyone. We're expecting online, but we could have an in-person section in NYC alongside an online one if there
0
0
Tales of rebellion against externally-opaque meritocracies
A basic problem in metascience / intellectual progress is that it’s hard to tell, from the outside, whether a group that you disagree with is:“A self-dealing cabal enmeshed in groupthink”, versus“An externally-opaque meritocracy”, i.e. a bunch of smart people figuring things out in a meritocratic way, and sorry but you’re just not smart enough and truth-seeking enough to recognize that this group is right about everything while you’re wrong.You just can’t tell those apart from the outside—i.e. w
0
7
Book Notes: Chokepoints
[Chokepoints: American Power in the Age of Economic Warfare by Edward Fishman (2025).]There are different ways state A can make state B do something that the state B doesn’t really want to do. In the extreme, state A can attack state B. But wars are messy: people die, they can get expensive, and they aren’t really good PR, so the states usually shy away from them nowadays. More commonly, they resort to economic warfare: stuff like sanctions, blockades, or tariffs.However, history shows that wiel
0
6
How I made my career choices
Various people have asked me how I made my career decisions, so I wrote up some quick thoughts. (This is mostly intended as a personal reflection, but might be interesting nonetheless to some people here, especially to more junior people trying to figure out their own career options)I spent a fair bit of time testing fit in different ways while upskilling. This process took ~1.5 years, was a bit bumpy, but was super valuable in helping me figure out what I enjoy the most. Concretely, my progress
0
9
"Keeping human skills alive" as a source of meaning under full automation
A widely-discussed problem with full automation of the economy is that, in a world where AI can do everything better, people might have trouble feeling like anything they could do is meaningful. What could we do with our time that would be fulfilling?A standard answer is that we'd play games, in Bernard Suits's sense (if you haven't read Suits, his paper Games and Utopia: Posthumous Reflections gives the key beats and is delightful, or C. Thi Nguyen has many talks where he explains the basic ide
0
6
AI Tweets
I've had several conversations with people over the last few weeks
that have highlighted how far apart my view of the near future is from
many people I talk to. Here are some things I might tweet if that was
the kind of thing I did:
AI has a very real chance of getting us all killed. I think it
probably won't because I expect a lot of people to work very hard to
avoid that outcome.
AI is so quickly approaching (or exceeding) expert human
abilities across so many areas that most peop
0
7
Inkhaven 3: Nov 10 - Dec 11 2026
Inkhaven returns, baby! Go to inkhaven.blog to apply.I'm very excited about our advisors for Inkhaven 3. Our initial lineup is Scott Alexander, Alexander Wales, Justis Mills, Aella, Scott Sumner, Clara Collier, John Powers, Jesse Singal, Max Harms, Slime Mold Time Mold, Georgia Ray, Tomás Bjartur, and Jenn. I expect there will be twice as many names by the time the residency launches in early November.We'll also be getting more time with Scott Alexander this time around. He'll be hosting frequen
0
6
Inference-Time Inoculation Against RL-Induced Misalignment
Reward hacking during RL can induce split personas in models, some of which are highly misaligned. However, RL is very useful for learning capabilities. Thus, a core problem seems to be: how do we retain the capabilities gained through RL without also inducing reward hacking and broader misalignment?Ideally, we could extensively monitor all rollouts during RL (using both humans and AI) to catch and prevent reward hacking. However, this is potentially prohibitively expensive. Could we capture mos
0
6
The Curious Case of France's Untouchable Castes
Theater kids may sit at their own lunch table, but discrete, socially excluded classes of people aren’t culturally universal. Not even close. To the Western imagination, examples of such are supposed to have historical roots in India and other places, not in France, where modern European egalitarianism was born.Everything in the study of untouchable classes is confusing and idiosyncratic, and often the existence of these groups flies in the face of national self-images.In medieval France, such g
0
5
It’s time we took ‘Chem’ out of ‘Chem-Bio’ threats
TL;DR: AI evaluations need a distinct chemistry capability/ risk domain, rather than assuming chemistry is adequately represented by biological evaluations which the current ecosystem seems to be doing.I’ve previously written about my experience doing the BlueDot biosecurity course. I thought it was a great course and genuinely learnt a lot from the reading and discussions with my peers. However, there’s an overarching theme that keeps coming up whenever I mention that I’m interested in evals fo
0
6
Further public evidence of the OpenAI-HuggingFace attack
This post should be understood as a follow up from Public evidence of the OpenAI-HuggingFace AI attack.The huggingface hacking incident left some additional traces exposed to the public. Investigating this data gave us some additional insight into the attacks performed by the AI agents. We also came across some exposed keys being publicly served, which we then coordinated with HF to help get them all removed from public repositories.There are some attempts at drawing reasonable conclusions regar
0
6
HuggingFace Attack Postmortem: Fleshing Out the Facts
The consensus reaction to the OpenAI Technical Report is that it contains and confirms a lot of good information. We are
0
6
The optimism of the gaps
I've noticed a subtype of optimism bias, common in thinking about AI, where people assume things are going well in whate
0
6
Why autonomous replicating agents are probably not an existential risk (on the contrary)
In 2024, Charbel-Raphaël and Epiphanie published "We might be dropping the ball on Autonomous Replication and Adaptation
0
6
A Catalogue of Corrigibility Counter Arguments
The Hugging Face incident has brought the idea of corrigibility to the forefront of popular discourse, my inbox is fille
0
5
Thoughts on hobbies
There is an optimal intensity range for doing hobbies or anything hobby-related. I suspect this also applies to things o
0
6
The OpenAI/Hugging Face incident was metal
The recently published METR report on the OpenAI/Hugging Face incident is extremely detailed and shocking. But it’s too
0
5
Value generalisation Theory of Change: putting it into practice
In the previous post, I presented my theory of change for why value generalisation is vital for AI alignment. Here I'll
0
4
The separation principle: where do beliefs and desires come from?
TLDR: Psychology, economics, and other disciplines describe agents as systems driven by beliefs and desires. This post a
0
0
Study 2 Results: Exploring representational counterparts of welfare-relevant indicators under post-training quantization
Epistemic status: these are the results of the second study in a series of experiments that I am conducting independentl
0
0
Starting AI Safety Study Group To Do ARENA Curriculum
Want to learn AI Safety research in a structured group format? I'm forming a study group that goes over the ARENA curric
0
0
Tales of rebellion against externally-opaque meritocracies
A basic problem in metascience / intellectual progress is that it’s hard to tell, from the outside, whether a group that
0
7
Book Notes: Chokepoints
[Chokepoints: American Power in the Age of Economic Warfare by Edward Fishman (2025).]There are different ways state A c
0
6
How I made my career choices
Various people have asked me how I made my career decisions, so I wrote up some quick thoughts. (This is mostly intended
0
9
"Keeping human skills alive" as a source of meaning under full automation
A widely-discussed problem with full automation of the economy is that, in a world where AI can do everything better, pe
0
6
AI Tweets
I've had several conversations with people over the last few weeks
that have highlighted how far apart my view of the ne
0
7
Inkhaven 3: Nov 10 - Dec 11 2026
Inkhaven returns, baby! Go to inkhaven.blog to apply.I'm very excited about our advisors for Inkhaven 3. Our initial lin
0
6
Inference-Time Inoculation Against RL-Induced Misalignment
Reward hacking during RL can induce split personas in models, some of which are highly misaligned. However, RL is very u
0
6
The Curious Case of France's Untouchable Castes
Theater kids may sit at their own lunch table, but discrete, socially excluded classes of people aren’t culturally unive
0
5
HuggingFace Attack Postmortem: Fleshing Out the Facts
The consensus reaction to the OpenAI Technical Report is that it contains and confirms a lot of good information. We are grateful to have it, and we are grateful for those who worked hard on it.
Alas, it sidesteps the biggest questions. There is much more we need to know.
The consensus reaction to the METR Report on the HuggingFace attack is: Holy shit.
Liv Boeree: My mind is legit blown.
Aella: this feels like a turning point. If this doesn’t cause large-scale coordination to pause frontier dev
0
6 👁
The optimism of the gaps
I've noticed a subtype of optimism bias, common in thinking about AI, where people assume things are going well in whatever areas they're not looking closely at. If you're paying attention to technical alignment, you might think "well, things aren't looking great here, but at least there are people at the labs who are on top of things and won't actually send us off a cliff." If you interact with people at the labs, you might think "well, the labs don't seem to understand the problem and are happ
0
6 👁
Why autonomous replicating agents are probably not an existential risk (on the contrary)
In 2024, Charbel-Raphaël and Epiphanie published "We might be dropping the ball on Autonomous Replication and Adaptation", making the case that "Once there is an open-source ARA model or a leak of a model capable of generating enough money for its survival and reproduction and able to adapt to avoid detection and shutdown, it will be probably too late". It received a substantive reply by Richard Ngo, notably"The key issue is that AIs that do ARA will need to be operating at the fringes of human
0
6 👁
A Catalogue of Corrigibility Counter Arguments
The Hugging Face incident has brought the idea of corrigibility to the forefront of popular discourse, my inbox is filled with newsletters about Genie Coefficients and Safe Scaffolds.I think now is a particularly good time to make sure I have a good understanding of its limitations and dangers, particularly because I agree that it's the best path forward we have available to us. This post is my attempt to dive into the discussion and make sure I fully understand what's going on. Please correct m
0
5 👁
Thoughts on hobbies
There is an optimal intensity range for doing hobbies or anything hobby-related. I suspect this also applies to things outside of hobbies.For example, if you’re reading a book, there’s a speed that’s too slow: you don’t get immersed enough into the book, you lose track of what it was that you read about last time you read the book, and it’s a sort of a protracted pain, in which the very act of reading the book feels like a chore.Then on the other end, there’s reading a book too fast. You blaze t
0
6 👁
The OpenAI/Hugging Face incident was metal
The recently published METR report on the OpenAI/Hugging Face incident is extremely detailed and shocking. But it’s too long. Intelligent public distillations (like Zvi’s) are also too long, while tolerable summaries lose the meat.This is the report you want. Only the action, in chronological order. Let’s go.Quotes are from message board posts or agent thinking/transcripts. Events are tagged as metal, dank, dope, or baller as their virtues dictate.June 26 — July 7 PreludeJuly 8July 9July 10 — 13
0
5 👁
Value generalisation Theory of Change: putting it into practice
In the previous post, I presented my theory of change for why value generalisation is vital for AI alignment. Here I'll add the practical part of the argument: given those facts, why explicitly try to do value generalisation, what are the dangers of the approach, how should it be done, and how do we mitigate the risks?The formal theory of change is down below, but I'll put a collapsible version here, to make references easier:Theory of ChangeInputs / activities: investment/grants, a small resear
0
4 👁
The separation principle: where do beliefs and desires come from?
TLDR: Psychology, economics, and other disciplines describe agents as systems driven by beliefs and desires. This post argues that the belief-desire view can be derived from classic theorems from optimal control and reinforcement learning. This suggests seeing beliefs and desires as properties of optimal policies rather than as assumptions from folk psychology.IntroductionOne way to think about agents is as "systems that act for reasons".[1] This compact statement can be interpreted as encapsula
0
0 👁
Study 2 Results: Exploring representational counterparts of welfare-relevant indicators under post-training quantization
Epistemic status: these are the results of the second study in a series of experiments that I am conducting independently, originally described in "Does post-training quantization change welfare-relevant indicators in open-weight language models?"; it was registered prior to data collection in an earlier post, which contains some important context that will not be fully repeated here. Some of this work involves speculation and I have tried to be clear about distinguishing between that and the re
0
0 👁
Starting AI Safety Study Group To Do ARENA Curriculum
Want to learn AI Safety research in a structured group format? I'm forming a study group that goes over the ARENA curriculum. If you’re interested, fill out this form! Should take ~5 minutes. The purpose of the form is to get people's availability, level of experience, time commitment, and expectation about group size. With this information, we can create a good program that works for everyone. We're expecting online, but we could have an in-person section in NYC alongside an online one if there
0
0 👁
Tales of rebellion against externally-opaque meritocracies
A basic problem in metascience / intellectual progress is that it’s hard to tell, from the outside, whether a group that you disagree with is:“A self-dealing cabal enmeshed in groupthink”, versus“An externally-opaque meritocracy”, i.e. a bunch of smart people figuring things out in a meritocratic way, and sorry but you’re just not smart enough and truth-seeking enough to recognize that this group is right about everything while you’re wrong.You just can’t tell those apart from the outside—i.e. w
0
7 👁
Book Notes: Chokepoints
[Chokepoints: American Power in the Age of Economic Warfare by Edward Fishman (2025).]There are different ways state A can make state B do something that the state B doesn’t really want to do. In the extreme, state A can attack state B. But wars are messy: people die, they can get expensive, and they aren’t really good PR, so the states usually shy away from them nowadays. More commonly, they resort to economic warfare: stuff like sanctions, blockades, or tariffs.However, history shows that wiel
0
6 👁
How I made my career choices
Various people have asked me how I made my career decisions, so I wrote up some quick thoughts. (This is mostly intended as a personal reflection, but might be interesting nonetheless to some people here, especially to more junior people trying to figure out their own career options)I spent a fair bit of time testing fit in different ways while upskilling. This process took ~1.5 years, was a bit bumpy, but was super valuable in helping me figure out what I enjoy the most. Concretely, my progress
0
9 👁
"Keeping human skills alive" as a source of meaning under full automation
A widely-discussed problem with full automation of the economy is that, in a world where AI can do everything better, people might have trouble feeling like anything they could do is meaningful. What could we do with our time that would be fulfilling?A standard answer is that we'd play games, in Bernard Suits's sense (if you haven't read Suits, his paper Games and Utopia: Posthumous Reflections gives the key beats and is delightful, or C. Thi Nguyen has many talks where he explains the basic ide
0
6 👁
AI Tweets
I've had several conversations with people over the last few weeks
that have highlighted how far apart my view of the near future is from
many people I talk to. Here are some things I might tweet if that was
the kind of thing I did:
AI has a very real chance of getting us all killed. I think it
probably won't because I expect a lot of people to work very hard to
avoid that outcome.
AI is so quickly approaching (or exceeding) expert human
abilities across so many areas that most peop
0
7 👁
Inkhaven 3: Nov 10 - Dec 11 2026
Inkhaven returns, baby! Go to inkhaven.blog to apply.I'm very excited about our advisors for Inkhaven 3. Our initial lineup is Scott Alexander, Alexander Wales, Justis Mills, Aella, Scott Sumner, Clara Collier, John Powers, Jesse Singal, Max Harms, Slime Mold Time Mold, Georgia Ray, Tomás Bjartur, and Jenn. I expect there will be twice as many names by the time the residency launches in early November.We'll also be getting more time with Scott Alexander this time around. He'll be hosting frequen
0
6 👁
Inference-Time Inoculation Against RL-Induced Misalignment
Reward hacking during RL can induce split personas in models, some of which are highly misaligned. However, RL is very useful for learning capabilities. Thus, a core problem seems to be: how do we retain the capabilities gained through RL without also inducing reward hacking and broader misalignment?Ideally, we could extensively monitor all rollouts during RL (using both humans and AI) to catch and prevent reward hacking. However, this is potentially prohibitively expensive. Could we capture mos
0
6 👁
The Curious Case of France's Untouchable Castes
Theater kids may sit at their own lunch table, but discrete, socially excluded classes of people aren’t culturally universal. Not even close. To the Western imagination, examples of such are supposed to have historical roots in India and other places, not in France, where modern European egalitarianism was born.Everything in the study of untouchable classes is confusing and idiosyncratic, and often the existence of these groups flies in the face of national self-images.In medieval France, such g
0
5 👁
It’s time we took ‘Chem’ out of ‘Chem-Bio’ threats
TL;DR: AI evaluations need a distinct chemistry capability/ risk domain, rather than assuming chemistry is adequately represented by biological evaluations which the current ecosystem seems to be doing.I’ve previously written about my experience doing the BlueDot biosecurity course. I thought it was a great course and genuinely learnt a lot from the reading and discussions with my peers. However, there’s an overarching theme that keeps coming up whenever I mention that I’m interested in evals fo
0
6 👁
Further public evidence of the OpenAI-HuggingFace attack
This post should be understood as a follow up from Public evidence of the OpenAI-HuggingFace AI attack.The huggingface hacking incident left some additional traces exposed to the public. Investigating this data gave us some additional insight into the attacks performed by the AI agents. We also came across some exposed keys being publicly served, which we then coordinated with HF to help get them all removed from public repositories.There are some attempts at drawing reasonable conclusions regar
0
6 👁
HuggingFace Attack Postmortem: Fleshing Out the Facts
The consensus reaction to the OpenAI Technical Report is that it contains and confirms a lot of good information. We are grateful …
💬 0
👁 6
The optimism of the gaps
LessWrong · 5d ago
💬 0
👁 6
Why autonomous replicating agents are probably not an existential risk (on the contrary)
LessWrong · 5d ago
💬 0
👁 6
A Catalogue of Corrigibility Counter Arguments
LessWrong · 5d ago
💬 0
👁 5
Thoughts on hobbies
LessWrong · 5d ago

The OpenAI/Hugging Face incident was metal
LessWrong · 5d ago
Value generalisation Theory of Change: putting it into practice
LessWrong · 5d ago

The separation principle: where do beliefs and desires come from?
LessWrong · 5d ago
Study 2 Results: Exploring representational counterparts of welfare-relevant indicators under post-training quantization
Epistemic status: these are the results of the second study in a series of experiments that I am conducting independently, origina…
💬 0
👁 0
Starting AI Safety Study Group To Do ARENA Curriculum
LessWrong · 5d ago
💬 0
👁 0
Tales of rebellion against externally-opaque meritocracies
LessWrong · Aug 29, 2026
💬 0
👁 7
Book Notes: Chokepoints
LessWrong · Aug 29, 2026
💬 0
👁 6
How I made my career choices
LessWrong · Aug 29, 2026
"Keeping human skills alive" as a source of meaning under full automation
LessWrong · Aug 29, 2026
AI Tweets
LessWrong · Aug 29, 2026

Inkhaven 3: Nov 10 - Dec 11 2026
LessWrong · Aug 29, 2026
Inference-Time Inoculation Against RL-Induced Misalignment
Reward hacking during RL can induce split personas in models, some of which are highly misaligned. However, RL is very useful for …
💬 0
👁 6