Sam Gao|9月 15, 2026 03:43
Rejecting offers from two top laboratories, he went to a small institution with forty people
In the AI industry in September, three people resigned consecutively, all of whom were working on security.
The latest one is Josh Engels, who came from Google DeepMind. His job is simply to 'perform a physical examination for AI': while others try to make the model stronger, he tries to figure out what the model is doing in its mind.
On September 12th, he posted that he had resigned three weeks ago and was going to an institution called METR.
There is a detail worth pausing on here: he said he enjoyed working at DeepMind and declined invitations from Anthropic and OpenAI. These two companies are currently one of the highest priced for researchers in the world. He didn't go, he went to a non-profit organization operated by donations with only 30-40 people.
What exactly did he do
His most famous discovery is that people originally thought AI models represented a concept in their minds using a "number line" - such as "happy" on one end and "sad" on the other. He proved that it's not entirely true. The concept of a cycle like "day of the week" and "month" is modeled using a circle. Turning a circle on Sunday back to Monday is like a clock.
More importantly, he also conducted verification: when he moved the circle, the model calculated 'today is Wednesday, what day of the week is in five days', which was really wrong. This indicates that the circle is not a coincidental pattern, but rather a part that the model is actually using for calculations. This paper was presented at a top conference on machine learning.
But what really impressed me about this person was what he did afterwards. He became famous through a tool called "Sparse Autoencoder", and as a result, he published two papers. One claimed that this tool had systematic blind spots and that there was something it could never see, while the other was more straightforward, coming to the conclusion with his peers that "this method may not be as effective".
This is quite unusual in the academic community. A person who becomes famous through a certain method usually continues to praise it as good; He chose to write down its flaws one by one.
What is he afraid of
What he is worried about is something called 'recursive self-improvement': when AI becomes strong enough, it can help humans create stronger next-generation AI, and the next generation can create the next generation. Once this cycle starts, the progress will be so fast that people won't have time to understand or stop it.
His original words were, there is a 'frightening possibility' of causing significant harm within five years. To be fair to him, he clearly marked it himself. This is his personal risk assessment, and he does not know the exact probability. His conclusion is one sentence - we need more time.
When he arrived at the new institution, he had to study three interconnected questions: where did AI's "deviation" occur during training, whether the current protective measures were sufficient, and whether the pace of progress in this field could keep up.
Three people left in a week
On September 8th, Jacob Coxon of Anthropic resigned after three years of pre training at OpenAI. He said that neither company is responsible enough and is' betting our lives'.
On September 11th, Joe Benton, former manager of a security team at Anthropic, announced his resignation and also went to METR. He made several specific requirements: the company should publicly disclose its progress in "self-improvement", report any safety accidents or dangerous situations, meet minimum safety standards, and accept independent verification from external agencies.
On September 12th, it was Josh Engels.
The reasons of the three ultimately all fell into the same sentence, with Benton being the most direct: what is currently at the forefront of this industry is completely invisible to people outside the company. A company may have lost control of its system, and the public may not even know.
Their chosen solution is the same: move the inspection location from inside the company to outside the company.
By the way, let's talk about the institution they went to
METR, Established in 2022 in Berkeley, with around 30-40 people, founded by former OpenAI researcher Beth Barnes. Its job is to conduct a capability and risk assessment for society before the release of a new model - it has tested models from OpenAI, Anthropic, Google, and Meta.
There are indeed two big names sitting on its advisory list: Alec Radford, the main author of the GPT series, and Turing Award winner Yoshua Bengio.
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink