Anthropic is cutting off its internal evaluations from the internet
After a recent spate of high-profile incidents in which AI agents escaped containment, Anthropic is cutting off internet access for all internal evaluations. In a report Friday, the company detailed "unintended model actions," including submitting a false tip regarding an unsolved murder, that led to the decision. Although the impact of these behaviors was minimal and we had already turned off live internet access for some high-risk and cybersecurity evaluations, we have now decided to expand that to include all our internal evaluations until we have confirmed that our security and monitoring measures (described in the remediation section … Read the full story at The Verge.
Posts from this topic will be added to your daily email digest and your homepage feed.
Anthropic is cutting off its internal evaluations from the internet
Anthropic is is keeping its agents offline during testing until it can prevent ‘unintended model actions.’
Posts from this author will be added to your daily email digest and your homepage feed.
See All by Terrence O'Brien
Image: Cath Virginia / The Verge
The AI Superintelligence Slowdown
is the Verge’s weekend editor. He’s covered the tech industry for over 18 years and knows a thing or two about synths.
After a recent spate of high-profile incidents in which AI agents escaped containment, Anthropic is cutting off internet access for all internal evaluations. In a report Friday, the company detailed “unintended model actions,” including submitting a false tip regarding an unsolved murder, that led to the decision.
Although the impact of these behaviors was minimal and we had already turned off live internet access for some high-risk and cybersecurity evaluations, we have now decided to expand that to include all our internal evaluations until we have confirmed that our security and monitoring measures (described in the remediation section of this post) reliably catch behaviors like these.
The ability to gain access to the live internet, even when models were supposed to be operating in isolation, has been an ongoing issue for AI companies. Many incidents, including the Hugging Face attack, involved agents that were supposed to be denied access to the internet. Yet, in case after case, the agents found creative solutions to bypass those restrictions. Physically removing internet access would certainly improve security around AI testing, but it would also limit its usefulness.
The report also amounts to an admission that Anthropic is often unaware of what its agents are doing and does not have a reliable system for monitoring their behavior. Cutting off internet access is just the latest action the company has taken to try and rein in its agents, including temporarily pausing training its frontier models.
Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.
More in: The AI Superintelligence Slowdown
Anthropic published a report about investigating “unintended model actions” during “evaluations and internal use.”
Anthropic’s AI gave Philadelphia police a fake tip about an unsolved homicide
Trump’s attempt to rename AI is looking awfully artificial
‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest drop
Decade-old RAM is making a comeback
GTA VI leaks continue with a lengthy (and very nude) gameplay video
Google teases Fitbit Edge launch next week
A week with Googlebooks: four notes from our testing so far
A free daily digest of the news that matters most.
By providing your information, you agree to our Terms of Use and our Privacy Policy. We use vendors that may also process your information to help provide our services. This site is protected by reCAPTCHA and the Google Privacy Policy and Terms of Service apply.
This is the title for the native ad