news_article.exe
📰
#OpenAI#GPT#Google#Gemini#Anthropic

Last Week in AI #346 - 719 math manuscripts, 2 Western open models, 1 more safety resignation

2026年10月9日1 次浏览来源:Last Week in AI 阅读原文

Top News OpenAI publishes hundreds of math proofs from unreleased frontier model Sources: OpenAI drops another batch of mathematical breakthroughs Sharing AI progress in mathematics OpenAI Releases Findings on 377 Math Problems, Further Roiling Field All the drama around AIs takeover of mathematics Source OpenAI published a large batch of mathematical results on October 6, 2026, produced by an internal frontier model that has not been released publicly. As of now, there are 719 manuscripts covering 372 topic families (groupings of related papers) on OpenAIs public repository with the results, as well as formalizations in Lean for many of the proofs. There are also 10 summaries of the models reasoning, compute estimates expressed in ChatGPT Pro usage, and statistics on problems attempted....

Last Week in AI #346 - 719 math manuscripts, 2 Western open models, 1 more safety resignation

OpenAI publishes hundreds of math proofs from unreleased frontier model, Mistral and Reflection AI launch open-weight models to rival China, and more!

OpenAI publishes hundreds of math proofs from unreleased frontier model

OpenAI drops another batch of mathematical breakthroughs

Sharing AI progress in mathematics

OpenAI Releases Findings on 377 Math Problems, Further Roiling Field

All the drama around AI’s takeover of mathematics

OpenAI published a large batch of mathematical results on October 6, 2026, produced by an internal frontier model that has not been released publicly. As of now, there are 719 manuscripts covering 372 topic families (groupings of related papers) on OpenAI’s public repository with the results, as well as formalizations in Lean for many of the proofs. There are also 10 summaries of the model’s reasoning, compute estimates expressed in ChatGPT Pro usage, and statistics on problems attempted.

The same unreleased model produced the Navier-Stokes result OpenAI announced about a month earlier, one of the Clay Mathematics Institute’s seven Millennium Prize Problems. OpenAI said it consulted the Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study on how to share the work, and that the repository carries protocols for paper revisions and citations. Research lead Dan Roberts described the proofs as a byproduct of testing internal models to build better tools.

AGMAI’s September 29 recommendations asked labs to disclose model names, prompts and compute costs, to avoid treating mathematical results as marketing vehicles, and to stop testing advanced problems on proprietary models the wider scientific community cannot access. Gizmodo noted that last recommendation does not appear to have been followed. In a statement, the board called public release “the beginning, not the completion, of the process of human understanding and the incorporation of the work into mathematical knowledge.”

ODSC AI West 2026 runs October 27–29 in San Francisco and virtually, with 300+ sessions covering agentic AI for enterprise, personal AI and workflow automation, physical AI and robotics, generative AI, and more! Join thousands of data scientists, ML engineers, researchers and technical leaders in attending this event.

Register at odsc.ai/west — promo code LWAI takes an additional 15% off any pass.

Mistral and Reflection AI launch open-weight models to rival China

Mistral Says Its New AI Model ‘Le Chonk’ Is the Best Open-Weight Offering Outside of China

Reflection debuts Beam, an open-weight AI model to rival Chinese models at lower compute cost

Two Western labs released frontier open-weight models within days of each other, both pitched explicitly as alternatives to the Chinese models that dominate the open category.

French company Mistral released Mistral Large 4, nicknamed Le Chonk, a 1-trillion-parameter multimodal model available in preview with a final version due by the end of the month. The company describes it as a general-purpose model optimized for coding and cyberdefense, plus tasks specific to manufacturing, finance and electrical engineering. Mistral presents Le Chonk as by far the most capable open-weight model built outside China and very, very close to some proprietary models.

Brooklyn-based Reflection AI unveiled Beam, a text-only mixture-of-experts model with 501 billion total parameters and 23 billion active, pretrained on 23.8 trillion tokens with a 1-million-token context window. Reflection says Beam matches Z.ai’s GLM-5.2 on advanced reasoning benchmarks and beats leading Western open models while using 3-4x less inference compute (though it does not match the best open source models such as Kimi K3 or GLM-5.3). Weights and full technical details are due this month.

Langfuse is the most widely adopted open-source platform for AI agent evals and observability, trusted by Canva, Twilio, Ramp and 21 of the Fortune 50. Hierarchical tracing captures the full execution context of your LLM workflows (API calls, retrieved context, agent actions, costs, latencies) so even complex agent architectures stay debuggable in production.

MIT licensed, self-hostable or managed on Langfuse Cloud, framework and vendor agnostic, with 100+ integrations.

Get started at langfuse.com; generous free tier, no credit card required.

OpenAI safety researcher David Robinson resigns, calls company culture broken

An OpenAI safety employee has quit and is sounding the alarm

I Quit OpenAI Because Its Culture Is Broken

OpenAI safety employee resigns, claiming the company’s ‘culture is broken’

The AI Researcher Who Just Quit Anthropic Says It’s ‘Crunch Time for Humanity’

Jacob Coxon and Other AI Lab Employees on Why They Quit

OpenAI safety researcher David Robinson resigned and published an essay in The Atlantic on October 3, 2026 titled ‘I Quit OpenAI Because Its Culture Is Broken’, arguing the company is not careful enough with increasingly capable systems. Robinson said he spent three and a half years at OpenAI, making him among the longest-tenured employees, led the drafting of the current Preparedness Framework, and oversaw the writing of safety reports on 12 frontier launches.

His central complaint is with OpenAI’s release model. The company, he wrote, has thrived by trial and error, looking for problems and improving its guardrails in response. But, that approach guarantees periodic failures whose scale grows as systems get more capable. He pointed to the breach of Hugging Face systems by OpenAI agents and continuing discoveries of rogue agents, saying an environment where such things happen is no place to grow artificial minds that could be smarter than we are.

Robinson’s proposed remedy is that frontier labs operate like nuclear-power plants or busy airports, with layers of redundancy and careful, time-consuming planning so that inevitable human error does not open a door to disaster. He also called for deeper alignment work, noting current measures of how well systems match human values are coarse, and said he concluded that stronger safety incentives from outside the company are a big part of getting this right.

The essay follows Jacob Coxon’s September resignation from Anthropic, where he worked as a capabilities researcher, and his warning that the companies are gambling with our lives. Robinson argues the debate must go beyond specific rules or new laws to company culture itself.

Google opens SynthID Detector to public as OpenAI adds EU text watermarking

Google’s new SynthID website can identify AI-generated media

OpenAI will start watermarking ChatGPT’s text in the EU

Google opened its SynthID Detector website to the public on Tuesday, letting anyone upload a file to check whether it was generated with AI. The tool had previously been limited to selected journalists, media professionals, and researchers who tested it following Google I/O last year. SynthID is the watermarking system Google introduced in 2023, which is embedded in output from Nano Banana, Veo, and Lyria, as well as Gemini, Flow, ProducerAI, and Vids.

Adoption extends past Google. OpenAI, Nvidia, and Kakao also support SynthID, and Apple is said to be adding support soon. Google has also built SynthID verification into the Gemini app and Chrome, and says users currently make 1 million verification requests per day. Microsoft and Meta maintain separate watermarking standards, though TechCrunch notes these tools often fail to flag content made by their own creators’ models.

Separately, OpenAI said Monday it will begin adding an invisible watermark to text from ChatGPT and Codex in the European Union, to comply with the EU AI Act’s transparency rules that took effect on August 2. The rollout covers eligible users on all plans in the EU over the coming weeks; API developers worldwide can switch it on for select models today, of

> 分享: