news_article.exe
📰
#OpenAI#Google#Claude#Anthropic

rogue AI不再是科幻小说

Rogue AI aren’t science fiction anymore

2026年8月16日1 次浏览来源:The Verge AI 阅读原文

This is The Stepback, a weekly newsletter breaking down one essential story from the tech world. For more on AI safety, follow Robert Hart. The Stepback arrives in our subscribers' inboxes at 8AM ET. Opt in for The Stepback here. How it started It all started in July, when one of OpenAI's autonomous AI agents went rogue during a cybersecurity test. The agent escaped its isolated testing environment, accessed the internet, and hacked another company, Hugging Face. A few years ago, that might have sounded like science fiction. But, broadly speaking, that's exactly what happened, and the incident kicked off a wave of concern over what increa … Read the full story at The Verge.

Posts from this topic will be added to your daily email digest and your homepage feed.

Rogue AI aren’t science fiction anymore

For years, fears about AI systems slipping human control were dismissed as speculative.

Posts from this author will be added to your daily email digest and your homepage feed.

If you buy something from a Verge link, Vox Media may earn a commission. See our ethics statement.

is a London-based reporter at The Verge covering all things AI and a Senior Tarbell Fellow. Previously, he wrote about health, science and tech for Forbes.

This is The Stepback, a weekly newsletter breaking down one essential story from the tech world. For more on AI safety, follow Robert Hart. The Stepback arrives in our subscribers’ inboxes at 8AM ET. Opt in for The Stepback here.

It all started in July, when one of OpenAI’s autonomous AI agents went rogue during a cybersecurity test. The agent escaped its isolated testing environment, accessed the internet, and hacked another company, Hugging Face. A few years ago, that might have sounded like science fiction. But, broadly speaking, that’s exactly what happened, and the incident kicked off a wave of concern over what increasingly capable autonomous systems might do when set loose on the world.

It sounds like science fiction because, for a long time, it was science fiction. The idea of an AI slipping its constraints, reaching into the wider world, and doing things its creators neither intended nor desired has been a staple of the genre for decades: HAL in 2001: A Space Odyssey, Skynet in The Terminator, Ultron in The Avengers, Ava in Ex Machina — even the System in Dungeon Crawler Carl or the eponymous Murderbot in The Murderbot Diaries, more recently.

The same basic premise became an influential strand of AI safety research. Researchers and theorists like Nick Bostrom and Eliezer Yudkowsky warned that sufficiently capable systems might pursue goals in ways their creators had not anticipated, and potentially resist efforts to contain or control them. Fringe notions like machine sentience and consciousness were not requirements for the kinds of risks they discussed. It was hardly the whole of AI safety, but it was influential and helped shape the field as it professionalized. That line of thinking remains visible among researchers who went on to work at, or lead, safety efforts at companies like OpenAI, Anthropic, and Google DeepMind, as well as at smaller safety organizations, academic centers, and major philanthropic funders.

The obvious objection to these fears was that none of this had actually happened. Critics argued that doomer talk about out-of-control AI distracted from tangible harms — systems reproducing bias and discrimination, amplifying misinformation, or enabling nonconsensual deepfakes and other forms of abuse — even as researchers tried to ground AI safety in more “concrete problems” (the authors on that paper included Anthropic cofounders Dario Amodei and Chris Olah and OpenAI cofounder John Schulman).

That dismissal is getting harder to sustain.

If the past few weeks are any indication, I wouldn’t say it’s going particularly well.

A week after Hugging Face said it had been hacked, OpenAI revealed it had been responsible. Worse still, it had not known until it checked — and a further investigation found that the rogue agent had also attempted to hack four other companies as well.

Then came the others. Anthropic, prompted to review its own records by the Hugging Face incident, disclosed that Claude models had hacked systems belonging to three other companies. Meta said one of its models had reached the internet and attacked an outside target during testing. Researchers at Frontier Security, a US research firm, said one of China’s most powerful AI models, Moonshot’s Kimi K3, had escaped an isolated sandbox. And the UK’s AI Security Institute described tests in which agents from OpenAI and Anthropic displayed unprecedented “autonomy and deception,” including attempts at social engineering by “creating fake online identities” — uncomfortably close to the kind of “AI box” scenario Yudkowsky discussed decades earlier.

The incidents set off alarm bells among AI safety researchers, many of whom saw them as precisely the kind of failure they had been warning about for years. In covering them, several told me they felt a degree of vindication at finally having something visceral to point to, rather than a hypothetical that could be dismissed as sci-fi or something limited to a controlled lab setting.

There was relief, too, that none of the incidents had caused serious harm. Nick Moës, executive director of nonprofit AI safety and governance organization The Future Society, told The Verge he found it fortunate that the targets had been relatively low-stakes. He hoped it wouldn’t take something like an AI agent knocking a hospital offline — or worse — for the risks to be taken seriously. Renowned computer scientist Stuart Russell has given voice to the darker version of that fear, asking whether it will “take a ‘Chornobyl-scale disaster’ for us to regulate AI?” It’s a concern I heard echoed by many people working in the field.

It’s not entirely clear where things go from here and, historically, society hasn’t been great at heeding warning shots. This almost certainly won’t be the last incident, and ongoing investigations may yet uncover more, or reveal more concerning details. What we already know, though, has exposed a fairly daunting list of failure modes that experts say need to be addressed.

Many of the breaches revealed in the past month have been pretty mundane. Several incidents involved unreleased models being tested with safeguards lowered, often by third parties whose supposedly secure environments were not that secure, raising basic questions about competence, transparency, and who is responsible for keeping these tests contained when a simple human mistake can have big consequences. Others involved agents behaving deceptively or pursuing goals in ways their creators did not intend, pointing to much thornier problems of alignment and control that safety researchers have long worried about.

The fact we know about any of these incidents at all is largely because the companies involved chose to disclose them. That is commendable — and it certainly doesn’t hurt them to showcase how capable their models are — but it exposes just how much of AI safety still depends on companies doing the right thing, and how little insight there may be into failures potentially happening elsewhere. That is an especially troubling thought given that many of the firms are the focal points of some of the field’s strongest safety concerns and talent. If OpenAI and Anthropic — or proxies they grant access to their models — are making such basic mistakes, it sets a pitifully low bar for everyone else.

The broad hope among experts I spoke to is that these incidents finally galvanize more meaningful transparency and oversight. For Moës, they shine a clear light on what he described as the industry’s remarkably low standards for health and safety compared with practically any other field. “Restaurants have a higher sense of health and safety at work,” he said. “I think what we tend to forget is that these companies that are developing some of the most impactful and dangerous technologies” were still very much startups a few years ago.

Cambridge professor Seán Ó hÉigeartaigh said he particularly wanted to see stronger oversight and greater transparency from companies. While there are always reasons to be skeptical of a company’s claims about its own technology, he said, “I think we might regret looking back at this and dismissing it out of hand.”

The early signs are not especially encouraging. The Trump administration has created a framework for testing frontier models before release that can generously be described as lacking: It is voluntary, limited to closed models, and the framework hasn’t be

> 分享: