news_article.exe
📰
#OpenAI

Priorities and principles for effective third party assessments

2026年9月22日1 次浏览来源:OpenAI Blog 阅读原文

OpenAI outlines priorities and principles for rigorous, secure, and independent third-party AI safety assessments of frontier models and safeguards.

Priorities and principles for effective third party assessments

Principles for effective assessments

Frontier AI labs carry an immense responsibility in training, evaluating, and deploying models safely. Third party assessments are a critical part of balancing that responsibility, expanding opportunities for input on AI safety, keeping the world informed, and keeping labs accountable to clear and independently supported safety claims.

As part of our efforts to pace the frontier, OpenAI is committed to supporting independent assessments with deep levels of access across training, evaluation, and deployment. That access should enable assessors to challenge our assumptions, identify risks we may have missed, and reach their own conclusions about the effectiveness of our safeguards.

We have long worked with third party assessors at various stages of the model development and deployment process. We have also incorporated third party assessments into our Preparedness Framework practices and supported organizations and legislation that advocate for a more rigorous and accountable process. Throughout these engagements, we have provided deep forms of access, including information about our technical safeguards, visible chain of thought access, and unprecedented levels of confidential data and internal deployment access for incident response and monitor red teaming. The priorities and principles shared here focus on our engagement with independent assessment organizations in the private and non-profit sector on technical safety assessments. They complement our work with governments on testing and evaluation, where distinct roles and responsibilities may call for different approaches.

Making these assessments effective requires strong independence mechanisms, scientific rigor, robust security practices, and clear responsibilities. Labs have a responsibility to enable meaningful scrutiny while protecting sensitive information. Labs and independent assessors share the responsibility for getting this right, and should be operating with shared international standards for safety and security practices. Below, we propose four priority areas for deeper assessment, alongside principles for rigorous, secure, and independent work.

Third party assessments are most useful when they address specific, consequential questions: Does the evidence support a lab’s safety case and safety claims? Do evaluations adequately test the risks they are intended to measure? Do safeguards work under realistic conditions?

The assessments described here are intended to take different forms depending on the safety questions being examined. We expect to support multiple assessments in parallel and over different periods of time, with some lasting weeks and others several months. While third-party assessments can also be part of pre-deployment work and may inform deployment decisions, the work described here is generally longer-term and launch-agnostic—focused on examining particular safety claims in depth over time.

Throughout our priority areas and principles, we refer to safety claims and safety cases. What we mean by these terms is the following:

Safety claim: A specific assertion about a model or system’s capabilities, behavior, or safeguards that bears on its safety and can be assessed against evidence. A claim should identify the risks and conditions it addresses, along with relevant assumptions and limitations.

Safety case: A structured argument, supported by evidence, explaining why a model or system’s risks are adequately managed for a specified activity, such as training, evaluation, or deployment. A safety case connects individual safety claims to the evidence supporting them and makes explicit the assumptions, uncertainties, and remaining risks that could affect its conclusions.

We propose four priority areas for deeper assessment, alongside principles for rigorous, secure, and independent work.

Independent assessment of safety cases, spanning training, evaluation, internal deployment and external deployment.

Assessment of safety cases requires expertise in alignment, control methods such as monitoring, cybersecurity, biological and chemical misuse and red teaming. Safety cases consist of claims including training, capability evaluations, and safeguards, which can be assessed as a whole or in parts (see priorities 2 and 3 below). Multiple assessors will likely need to examine different parts of the cases, drawing on their respective expertise. Together, their assessments should answer questions such as:

Is the evidence for safety cases for training, evaluation, and deployments substantiated? Were the conditions of the safety case followed during training, evaluation, and deployment?

Do our safety cases cover the most urgent risks identified in the course of an assessment? Do the safety claims support the overall safety case? Are there any gaps or areas for improvement?

Are effective methods used to identify and reduce incentives in training that could reward deception, reward hacking, destructive actions, or circumventing restrictions?

Assessment of critical safeguards, across internal and external deployments

Our safeguard stack is always evolving to meet the changing capabilities and landscape. Our safeguards span training, internal deployment, and external deployment. They currently include model-level safeguards, enforcement safeguards, and security safeguards, as well as misalignment monitors—covering a wide variety of risks (such as loss of control, and misuse in cyber, biological and chemical risks). Technical partnerships can identify weaknesses in safeguards now while improving assessment methods and accelerating standards development, providing a stronger technical basis for future public policies to pace the frontier—particularly for internal deployments, where safety and security standards are still nascent.

Independent assessments should examine how well these safeguards work and where they may fall short, such as:

Using “grey box” access, are safeguards robust to adversarial testing (jailbreaks) and do they sufficiently protect against capability uplift in high risk domains (e.g, cyber, bio)? Does our adversarial testing and red teaming cover the most important risks?

In authorized testing under realistic operating conditions, how do agents interact with cyber defenses such as access controls, sandboxing, and detection and response systems? Which defenses prevent, detect, or contain harmful actions, and where do they fail?

Do our misalignment monitors have any critical gaps that could lead to loss of control or severe misalignment, for both internal and external deployments? How reliable is chain-of-thought monitoring as a source of evidence for safety or alignment as model capabilities improve?

Is appropriate monitoring implemented across all relevant training, evaluations, and deployment, in a way that cannot easily be disabled?

Are safeguards implemented commensurate with the capabilities?

Assessment of capability evaluations that cover Preparedness risk categories (Chemical and Biological Risks, Cybersecurity, AI Self-Improvement) and alignment evaluations for misalignment risks

Our Preparedness Framework requires evaluations to assess key frontier risk areas. As thresholds are surpassed and evaluations saturate, it is important to consistently refresh and ensure coverage and quality of evaluations that assess capabilities in Preparedness risk areas, and alignment evaluations that seek to assess severe misalignment risks.

Do evaluations that assess Preparedness risks adequately cover the Preparedness risk threshold definition? Are thresholds set correctly for these evaluations?

Are evaluations updated when models consistently achieve the highest scores, and do the new tests meaningfully measure more advanced capabilities?

Do our alignment evaluations adequately cover severe misalignment risks, and what important behaviors or condition

> 分享: