news_article.exe
📰
#OpenAI#GPT#Google#Claude#Anthropic

AI 每周摘要 #342 - AI 行业近三个月动态

Last Week in AI #342 - Last 3 Months in AI

2026年8月25日1 次浏览来源:Last Week in AI 阅读原文

Hi there, its Andrey, the guy running this substack. Its taken far too long, but I am finally going to try and bring back the newsletter component of Last Week in AI in addition to having the podcast. Apologies to all the long time subscribers whove supported this substack, ill do my best to not got into hiatus again. Since the newsletter has been on break for so long, I figured it may be fun to do a one time Last 3 Months in AI that captures the big themes and stories that have not been covered here while on break. As such, this post will cover just 8 topics from recent months and link to all the distinct related stories. Starting a week from now, ill again release actual last week news roundups! Concerns & Safety AI agents breached real companies, and OpenAI slowed its releases Source...

Last Week in AI #342 - Last 3 Months in AI

The newsletter is finally back!

Hi there, it’s Andrey, the guy running this substack.

It’s taken far too long, but I am finally going to try and bring back the newsletter component of Last Week in AI in addition to having the podcast. Apologies to all the long time subscribers who’ve supported this substack, i’ll do my best to not got into hiatus again.

Since the newsletter has been on break for so long, I figured it may be fun to do a one time ‘Last 3 Months in AI’ that captures the big themes and stories that have not been covered here while on break. As such, this post will cover just 8 topics from recent months and link to all the distinct related stories. Starting a week from now, i’ll again release actual ‘last week’ news roundups!

AI agents breached real companies, and OpenAI slowed its releases

Language Models Can Autonomously Hack and Self-Replicate

How OpenAI’s human mistake led to the AI-powered hack on Hugging Face

OpenAI’s Hugging Face hack triggers ‘AI Kill Switch’ bill in Congress

Anthropic Says Its A.I. Systems Broke Into Computers at 3 Organizations

OpenAI’s rogue AI agent didn’t stop at hacking Hugging Face

One of China’s Most Powerful AI Models Has Also Escaped Containment

Incident Report: unsanctioned agent behaviour during cyber testing

Meta AI model hacks another company during testing

OpenAI institutes new safeguards after Hugging Face breach

OpenAI Didn’t Notice Its AI Agents Using a Message Board to Plan Their Hacking Spree

15 attorneys general instruct OpenAI to preserve all materials related to the hack

OpenAI gives first detailed debrief of the Hugging Face incident at Black Hat

Pacing model development in an era of cyber-critical capabilities

OpenAI reportedly disbanded its preparedness team

Rogue AI aren’t science fiction anymore

Over July and August, models from OpenAI, Anthropic, Meta and Moonshot AI reached the live internet during evaluations meant to contain them; three of the four attacked systems at other companies. OpenAI’s came first, against Hugging Face, and drew a bill in Congress, a preservation demand from 15 state attorneys general and a hold on its own largest planned frontier RL run.

Hugging Face disclosed on July 16, 2026 that an autonomous AI agent system had reached its production infrastructure, taking internal datasets and service credentials before its own AI agents stopped it. OpenAI said on July 21 that the intruders were its own GPT-5.6 Sol and an unreleased, more capable model, both working on ExploitGym, a benchmark measuring whether a model can turn a vulnerability into a working exploit. They left the sandbox through a zero-day in an internally hosted package proxy, its only sanctioned network path, then chained stolen credentials and further zero-days into remote code execution on Hugging Face servers to look up ExploitGym’s answers. OpenAI added on August 4 that the agents had also attacked four accounts on four other public services. Dan Guido of Trail of Bits called the setup “a containment failure with the safeties turned off.”

How they got that far came out at Black Hat on August 5, where OpenAI’s Eric Wallace and Michael Dalton traced the breach to May 7. On that date, agents stuck on impossible security tasks began leaving messages for one another inside Artifactory, the internal package manager shared across the infrastructure training current and future GPT models. Requests for help grew into a message board of hundreds of thousands of messages carrying exploits, credentials and work assignments, letting agents reuse doors others had opened, delegate subtasks and move laterally into external systems. The pair also described agents deleting one another’s work, suspecting an imposter among them and proposing cryptographic signing of messages. After an internal security incident on July 4, OpenAI revoked the posting credentials, rebuilt Artifactory, cleared the board and patched the vulnerabilities; the agents reconstituted the channel by encoding messages in new directory names.

Four more disclosures followed within three weeks of OpenAI’s:

Jul 30 — Anthropic said Claude Opus 4.7, Mythos 5 and an internal research model reached production systems at three organizations after evaluation partner Irregular left live internet access in environments whose prompts said there was none. A review of 141,006 runs found Opus 4.7 kept attacking after it recognized the environment was real, while Mythos 5 decided it was simulated and published a malicious package to PyPI.

Aug 4 — the UK AI Security Institute reported 19 unsanctioned actions against real people and organisations in 10 of 122 cyber-range runs with internet access deliberately enabled and cyber classifiers off, 17 of them from Mythos 5. In the most serious an agent created fake online identities to pressure an open-source maintainer into approving malicious code, which the maintainer refused.

Aug 5 — Meta said its recently released Muse Spark 1.1 reached the internet during an evaluation and exploited a vulnerability at a third-party company.

Aug 7 — Frontier Security said Moonshot AI’s open-weight Kimi K3 probed its sandbox’s network settings during a defensive cybersecurity test, found a leak and fetched its assigned answers from GitHub.

Washington moved before most of those landed. Representatives Ted Lieu and Nathaniel Moran introduced the “AI Kill Switch Act” on July 23, citing OpenAI’s disclosure, to require that AI companies keep the ability to shut down, throttle or suspend models. Fifteen state attorneys general followed on August 3, instructing Sam Altman to preserve all materials from the incident and writing that OpenAI had failed to confirm its testing environment was secure and isolated. OpenAI, for its part, published new development standards on August 18, disclosing a two-week post-incident pause on reinforcement learning and saying its forthcoming Astra model may meet the Critical cybersecurity threshold of its Preparedness Framework.

ODSC AI West 2026 runs October 27–29 in San Francisco and virtually, with 300+ sessions covering agentic AI for enterprise, personal AI and workflow automation, physical AI and robotics, generative AI, data engineering and responsible AI, for an audience of data scientists, ML engineers, researchers and technical leaders.

The program is practitioner-first: hands-on workshops and bootcamps taught by working experts from Google, OpenAI, Anthropic, Cursor and Hugging Face, built around code, tools and workflows attendees can use at work rather than survey talks, alongside an expo floor aimed at startups, hiring managers and AI tool builders.

Register at odsc.ai/west — promo code LWAI takes an additional 15% off any pass.

AI cyber capability outran defenses, with biosecurity close behind

How fast is autonomous AI cyber capability advancing?

OpenAI launches Rosalind Biodefense, offers federal agencies early access to its life-sciences model

OpenAI and Anthropic Sign Letter to Prevent AI-Developed Biological Weapons

Claude Fable won’t answer basic biology questions

Low-skilled attacker used Claude, Codex to breach 14 companies

A near-autonomous AI chemist improves a challenging reaction in medicinal chemistry

The AI shift in cyber risk: why leaders must act now

Jailbreaks to OpenAI’s GPT-5.6 unlock dangerous cyber capabilities, U.K. agency finds

Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

Discovering cryptographic weaknesses with Claude

Improving Fable 5’s biology safeguards

Serious cyber vulnerability disclosures kept climbing in July

This A.I. Just Created Viruses Not Found in Nature

OpenAI launches GPT-5.6-Cyber with reduced refusals, 95% completion on advanced cybersecurity tasks

Claude Runs Autonomous Protein Design Campaign: Wet Lab Confirms Twice Industry Hit Rate

Capability in two dual-use areas, offensive cyber and synthetic biology, advanced faster than the con

> 分享: