Summary

OpenAI has proposed four areas for independent assessment of frontier AI safety, from safety cases and safeguards to capability evaluations and misalignment incidents. It also set out principles for how assessors and AI labs should scope, conduct and report the work.

OpenAI has proposed four areas for independent third-party assessment of frontier AI safety, alongside principles for setting the scope, conducting the work and reporting findings. The September 22, 2026 announcement describes an approach the company says it intends to support; it is not a report on a completed assessment.

The proposed assessments would examine whether evidence supports a lab’s safety claims, whether tests cover the risks they are meant to measure, and whether safeguards work in realistic conditions. OpenAI says it expects different assessments to run in parallel and over varying periods, from weeks to several months. The work described is generally intended to examine claims over time rather than focus on a particular product launch.

Four areas for independent scrutiny

The first priority is assessing safety cases across training, evaluation and deployment, both within a lab and externally. A safety case is a structured argument, backed by evidence, for why a system’s risks are adequately managed for a particular activity. It connects specific safety claims to evidence and makes assumptions, uncertainties and remaining risks explicit. OpenAI proposes that assessors examine whether the evidence supports those claims, whether stated conditions were followed, and whether the case addresses urgent risks and gaps.

The second area is the performance of safeguards. OpenAI’s examples include testing safeguards against adversarial attempts, examining how agents interact with cyber defences under authorised, realistic conditions, and assessing whether monitoring is sufficiently reliable and present across relevant stages of training, evaluation and deployment. The company also identifies questions about whether safeguards are proportionate to model capabilities and whether they adequately address risks such as cyber, biological and chemical misuse or loss of control.

Third, the company proposes scrutiny of capability evaluations for its Preparedness risk categories: chemical and biological risks, cybersecurity and AI self-improvement. Assessors could examine whether tests adequately represent the relevant risk thresholds, whether tests are refreshed as models reach their highest scores, and whether alignment evaluations cover severe misalignment risks.

The fourth area is independent investigation of critical misalignment incidents—for example, cases involving unauthorised model actions or evasion of oversight. OpenAI says investigations could look at what happened, contributing factors, and whether safeguards or remediation would reduce the chance of similar incidents. It notes that this work may involve sensitive internal or third-party information.

Principles for credible, secure assessments

OpenAI proposes that the lab and assessor agree on the scope and define the claims before assessment begins. Reports should distinguish what was examined from what was not, and there should be a process for deciding whether important risks found outside the agreed scope warrant further work.

Access should be proportionate to the claims being assessed, subject to legal, security and intellectual-property constraints. Where direct access is impractical, the company proposes alternatives such as working through a designated representative or using indirect or privacy-preserving access. Assessors should explain their methods, criteria and uncertainties, and justify criteria where established standards do not exist.

The remaining principles address assessor expertise and independence, conflict-of-interest safeguards, information security and confidentiality, and findings that give labs enough detail and time to address problems where appropriate. Reports should be shared as openly as possible while protecting sensitive information. OpenAI also calls for clear redaction policies that preserve assessors’ editorial independence and make substantive redactions visible where possible.

OpenAI says it is in discussions with multiple third parties about proposals aligned with these priorities and intends to support a broader community of assessors. It argues that no single assessor can cover all urgent frontier-safety questions, making the proposed framework’s emphasis on varied expertise and shared practices central to its approach.

Sources