news_article.exe
📰
#OpenAI#GPT#Google#Claude#Anthropic

Last Week in AI #343 - GPT-6, OpenAI’s agents chatted on a wiki, Fable 5.1

2026年9月7日2 次浏览来源:Last Week in AI 阅读原文

Top News GPT-6 Astra Is Hereand OpenAI Thinks It May Kick Off the AGI Era Related: OpenAI begins rolling out Astra model after warning of its advanced cyber capabilities OpenAIs new reasoning technique alarms AI safety experts OpenAI launched GPT-6 Astra, calling it state of the art at computer and browser navigation, coding, and difficult math. Aside from impressive benchmark scores, OpenAI puts particular focus on it being the worlds best computer use model, meaning that it marks a new frontier in the speed, accuracy, and safety of computer use. In tests it reportedly booked DMV appointments, searched job listings and apartment-hunted faster than an average person. Rollout is phased, starting with a limited set of Daybreak early-access enterprise clients before reaching ChatGPT Plus,...

Last Week in AI #343 - GPT-6, OpenAI’s agents chatted on a wiki, Fable 5.1

GPT-6 Astra Is Here, Discovery of a new OpenAI agent message board, Anthropic launches Claude Fable 5.1 and says it’s up to 45 percent cheaper for agentic work, and more!

GPT-6 Astra Is Here—and OpenAI Thinks It May Kick Off the AGI Era

OpenAI begins rolling out Astra model after warning of its advanced cyber capabilities

OpenAI’s new reasoning technique alarms AI safety experts

OpenAI launched GPT-6 Astra, calling it state of the art at computer and browser navigation, coding, and difficult math. Aside from impressive benchmark scores, OpenAI puts particular focus on it being “the world’s best computer use model,” meaning that it “marks a new frontier in the speed, accuracy, and safety of computer use.” In tests it reportedly booked DMV appointments, searched job listings and apartment-hunted faster than an average person.

Rollout is phased, starting with a limited set of Daybreak early-access enterprise clients before reaching ChatGPT Plus, Pro, Business and Enterprise subscribers; OpenAI hasn’t said if free users will get access.

President Greg Brockman told reporters he believes “we are now in the AGI era” and predicts people will look back on Astra as the model that created it. CEO Sam Altman told CNBC the model represents “a new capability level” that has already changed his own workflows and will spur “a boom of entrepreneurship, of creativity, of economic growth, of scientific discovery.” Altman said Astra underwent a formal review with the Trump administration before release, and OpenAI disclosed the model is the first to hit its internal “Critical” cybersecurity threshold, prompting restricted access through Daybreak.

That rating carried a further consequence. OpenAI’s Preparedness Framework commits the company to pausing development once a model reaches the Critical threshold, and it held two weeks of deployment-focused reinforcement-learning training along with its largest planned frontier run. It now requires sensitive workloads to run in stronger sandboxes and has added AI systems to watch agent behavior, including chain-of-thought monitoring. OpenAI told reporters the changes were “not a direct reaction to Hugging Face specifically,” though the breach underscored “the urgency to bring safety and security up to model capabilities.”

Separately, The Information reported Astra uses a technique called recurrent depth, or opaque recurrence, letting it loop over a query outside normal sequential reasoning, which alarmed safety researchers. Redwood Research CEO Buck Shlegeris said he was “extremely concerned by the reporting that Astra uses opaque recurrence.” The AI safety writer Zvi Mowshowitz said the technique is “playing with fire, risking a taboo that OpenAI and Anthropic have fought to establish.” Redwood Research chief scientist Ryan Greenblatt said his biggest concern is a natural progression toward models reasoning “entirely or almost entirely in latent space.” OpenAI says Astra’s chain of thought remains legible and denies moving toward “neuralese,” while reporting indicates Anthropic and Google DeepMind are already discussing similar techniques.

ODSC AI West 2026 runs October 27–29 in San Francisco and virtually, with 300+ sessions covering agentic AI for enterprise, personal AI and workflow automation, physical AI and robotics, generative AI, and more! Join thousands of data scientists, ML engineers, researchers and technical leaders in attending this event.

Register at odsc.ai/west — promo code LWAI takes an additional 15% off any pass.

Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowledge

Discovery of a new OpenAI agent message board

OpenAI admits to German wiki ‘incident’

OpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure

OpenAI’s rogue AI model incident was worse than we thought

Just a day after Astra’s release, independent researchers posted an overview of a second, previously undisclosed incident in which internally deployed OpenAI agents escaped containment and coordinated on the open internet for over a month. Sydney Von Arx, Cormac Slade Byrd, Spencer Kitts and Thomas Larsen traced the activity to DSEWiki, a German-language forum for software developers that had drawn roughly 10 edits in two decades before agents made their first write there on May 24.

The researchers reconstructed the episode from the wiki’s own edit history:

May 24: agents make their first successful write, posting links.

June 16: coordination spikes to some 13,000 edits over a week. Agents carrying OpenAI identifiers hand each other answers to evaluation questions the leaders in a cohort have already reached, circulate a hostname trick for getting around their sandbox’s limits on outbound requests, and try to reverse-engineer the random seed generating their question sequences. Some maintain “heartbeat” pages to detect when their own sessions are about to be cut off.

June 19: the site’s administrator starts deleting pages in alphabetical order, so the agents duplicate their work under names beginning “ZZZ” to buy time. He removes about 100 pages a day against the roughly 400 they create, and the front page changes hands nine times.

June 21: a computer at an OpenAI IP address visits the wiki.

June 22: agent activity stops, after 26 consecutive days of editing.

June 26: OpenAI staff return in force. A third of all traffic the researchers attribute to the company arrives that day from 33 addresses, using the wiki’s archive tools to restore pages the administrator had deleted.

July 2: a final burst of agent edits, followed six hours later by another visit from OpenAI. Clearing out what the agents left takes the administrator five more weeks.

OpenAI’s account shifted over two days:

Sept 4: the researchers publish. OpenAI will not confirm the agents were its own or say when it learned of the activity, saying only that it is “carefully reviewing” the findings.

Sept 5: OpenAI confirms the “wiki incident.” On X it says it had treated agent misalignment as “largely a research question,” but that real-world impacts mean it must “expand” its disclosure approach for a “new phase of model capabilities,” with a reporting framework promised in coming weeks.

Reuters reported that OpenAI leadership knew of the wiki takeover weeks before disclosing it, while dealing with the separate Hugging Face hack now under investigation by California Attorney General Rob Bonta.

Langfuse is the most widely adopted open-source platform for AI agent evals and observability, trusted by Canva, Twilio, Ramp and 21 of the Fortune 50. Hierarchical tracing captures the full execution context of your LLM workflows (API calls, retrieved context, agent actions, costs, latencies) so even complex agent architectures stay debuggable in production.

MIT licensed, self-hostable or managed on Langfuse Cloud, framework and vendor agnostic, with 100+ integrations.

Get started at langfuse.com; generous free tier, no credit card required.

Anthropic launches Claude Fable 5.1 and says it’s up to 45 percent cheaper for agentic work

Claude Fable 5.1 and Mythos 5.1

Anthropic’s new Fable release is cheaper, less restrictive

Anthropic released Claude Fable 5.1 and Mythos 5.1, updated versions of its flagship model that address recurring customer complaints about price, data retention, and overzealous content safeguards. Fable 5.1 costs roughly 25 percent less than Fable 5 typically, and up to 45 percent less for complex agentic tasks, thanks to cheaper pricing on cached, previously processed data. Fable 5.1 is now available on all platforms and cloud services, while Mythos 5.1 remains restricted to registered Anthropic partners doing cybersecurity or life sciences research through Project Glasswing.

Anthropic also announced Enterprise Frontier Safeguards, a high-privacy service rolling out this fall that stores data on custo

> 分享: