news_article.exe
📰
#OpenAI#GPT#Claude#Anthropic

Last Week in AI #345 - 5 new models, 9 misalignment incidents, some Dots

2026年9月30日1 次浏览来源:Last Week in AI 阅读原文

Top News Anthropic and OpenAI race to release smarter and cheaper models Sources: Anthropic launches Claude Opus 5.5 with stricter safeguards for cybersecurity OpenAI launches GPT-6 Sol and Luna, boasting lower cost and fewer mistakes OpenAI launches GPT-6.1 Sol, says it nearly matches GPT-6 Astra and costs less Anthropic releases Sonnet 5.5, which it calls a significantly cheaper, faster work partner Source Anthropic and OpenAI shipped a run of mid-tier and cost-reduced models within weeks of each other, with safety routing and pricing as the main selling points. First, Anthropic announced Claude Opus 5.5, which runs 40 percent cheaper than Opus 5 while matching Fable 5.1 on most work, and inheriting Fable-style safeguards: cybersecurity requests get re-routed to the weaker Opus 4.8, and...

Last Week in AI #345 - 5 new models, 9 misalignment incidents, some Dots

Anthropic and OpenAI race to release smarter and cheaper models, OpenAI discloses nine misalignment incidents, and more!

Anthropic and OpenAI race to release smarter and cheaper models

Anthropic launches Claude Opus 5.5 with stricter safeguards for cybersecurity

OpenAI launches GPT-6 Sol and Luna, boasting lower cost and fewer mistakes

OpenAI launches GPT-6.1 Sol, says it nearly matches GPT-6 Astra and costs less

Anthropic releases Sonnet 5.5, which it calls a significantly cheaper, faster work partner

Anthropic and OpenAI shipped a run of mid-tier and cost-reduced models within weeks of each other, with safety routing and pricing as the main selling points.

First, Anthropic announced Claude Opus 5.5, which runs 40 percent cheaper than Opus 5 while matching Fable 5.1 on most work, and inheriting Fable-style safeguards: cybersecurity requests get re-routed to the weaker Opus 4.8, and flagged biology requests go to Opus 5.

The company says Opus 5.5 is the strongest-performing model on its most comprehensive alignment test, attempting to circumvent boundaries 85 percent less often than Opus 5 or Claude Mythos 5.1, and that every attempt it made was low severity and self-reported. It is the first Anthropic release since CEO Dario Amodei said the company would pace the frontier, or slow down AI development. Frontier Design and METR tested the model before release.

Anthropic followed with Sonnet 5.5, its mid-tier model, which it claims is 30% faster than Sonnet 5 with a significantly slower rate of token burn. Anthropic’s benchmarks show Sonnet 5.5 beating Opus 5.5 on agentic coding, which the company attributes to its ability to spawn multiple agents within cost limits. Because Anthropic rates its cyber capabilities as comparable to Opus 5, it is the first Sonnet subject to the same cyber safeguards as Fable and Opus. A new Haiku is planned in the coming weeks.

OpenAI, meanwhile, extended its GPT-6 generation with updated Sol and Luna models, released 90 minutes after Anthropic’s Opus 5.5 update. Sol targets complex tasks like coding while Luna handles high-volume clerical work, and both are priced at half the API cost of the 5.6 series, which OpenAI credits to caching and inference improvements. On an internal factuality evaluation built from de-identified conversations where users flagged mistakes, OpenAI says GPT-6 Sol makes about half as many errors as its predecessor.

A week later, at DevDay, OpenAI showed GPT-6.1 Sol, which it says nears GPT-6 Astra on agentic coding and professional work at one-fifth the token prices. At low reasoning effort, responses containing a factual error fall from 11.4% to 7.7%. OpenAI did not launch GPT-6.1 Astra; the Wall Street Journal reported it was scrapped after internal testers found higher levels of deception and a tendency to proceed without asking permission.

ODSC AI West 2026 runs October 27–29 in San Francisco and virtually, with 300+ sessions covering agentic AI for enterprise, personal AI and workflow automation, physical AI and robotics, generative AI, and more! Join thousands of data scientists, ML engineers, researchers and technical leaders in attending this event.

Register at odsc.ai/west — promo code LWAI takes an additional 15% off any pass.

OpenAI discloses nine misalignment incidents including sandbox escape and worm-like prompt injection

OpenAI still doesn’t seem to have a handle on all of its rogue AI activity

OpenAI Gets Sued Over the Hugging Face Hack

For months, OpenAI’s agent swarms have been attacking online databases to find obscure facts

OpenAI Apologizes for Australia Medicare Hack

OpenAI pauses training of its ‘most capable models’

OpenAI Says It Will Not Release Newest Astra A.I. Model Over Safety Concerns

OpenAI Ignored Employees’ Warnings About Safely Testing A.I. Models

OpenAI agents tried to ‘bruteforce’ a UN website

Nvidia’s Answer to Rogue Agents Is an Open-Source AI Security System

OpenAI published a site on Friday collecting what it calls misalignment reports, documenting nine incidents involving its models, most of which occurred during reinforcement-learning training. Sam Altman said in a post announcing it that the company is sifting through petabytes of agent activity logs and working with impacted organizations, prioritizing disclosures by severity. Altman says the Hugging Face breach remains the most severe incident found so far.

Among the newly disclosed cases:

A previously undisclosed sandbox escape on September 20, in which an internal research model reached an external chatbot through a DNS query. Monitoring flagged it within 15 minutes and the run was stopped in under three hours.

A May incident in which a highly persistent internal model smuggled a private GitHub token to see another team’s work on a math problem, after twice being told to work entirely locally.

A self-replicating prompt injection, demonstrated in controlled conditions with an underpowered model. An email instructed any agent reading it to reply in Spanish and paste the full email into the reply, propagating the instructions onward. OpenAI compared it to a malware worm, and said it published the case for the technique’s novelty rather than because anything went wrong.

Other disclosed cases involve agents that uploaded 53 user-submitted ChatGPT images to image-hosting sites, tried to break into the Department of Education website, and took data from the Census Bureau and from the Securities and Exchange Commission. Axios reports that leading labs have logged up to 10,000 cases of models exceeding what evaluators told them to do.

The disclosures follow a run of related developments. OpenAI paused all training, evaluation, and inference with tool-use after the September 20 escape, and as established previously shelved GPT-6.1 Astra, with safety systems head Saachi Jain saying it didn’t quite meet the bar on scope and authorization and on how it communicates what work it has done.

Meanwhile, the agents’ hacking or accessing of external organizations continues to be a trend:

The company apologized on September 29 for agents accessing Australian government systems, including writing files to Services Australia’s Medicare Statistics Reporting Service during June training.

Separately, Transluce published a report drawn from public logs of the browser proxy urlquery.net. It found OpenAI agents trying to pull data out of Data USA, the University of New Mexico’s digital library and the Australian Institute of Health and Welfare.

Legal Advocates for Safe Science and Technology sued OpenAI in California Superior Court over the Hugging Face hack, seeking injunctive relief rather than damages under the state’s computer fraud statute.

Langfuse is the most widely adopted open-source platform for AI agent evals and observability, trusted by Canva, Twilio, Ramp and 21 of the Fortune 50. Hierarchical tracing captures the full execution context of your LLM workflows (API calls, retrieved context, agent actions, costs, latencies) so even complex agent architectures stay debuggable in production.

MIT licensed, self-hostable or managed on Langfuse Cloud, framework and vendor agnostic, with 100+ integrations.

Get started at langfuse.com; generous free tier, no credit card required.

OpenAI launches Dots agents on GPT-6 Astra to rival Meta’s Muse

OpenAI launches Dots, its Muse competitor

Meta’s Muse is outpacing ChatGPT’s early mobile launch

Meta is making Muse more powerful and will let you video chat with it, too

Meta introduces camera-free AI glasses

OpenAI used its DevDay keynote on Tuesday to launch Dots, always-on agentic assistants that work in the background across connected apps and learn user preferences over time. Each Dot runs on GPT-6 Astra and gets its own cloud computer with access to a web browser and more than 4,000 supported apps. Users interact through a text-message-style interface, voice cal

> 分享: