A framework written by the subject of the assessment

On September 22, 2026, OpenAI published "Priorities and principles for effective third party assessments," a document laying out where the company wants outside assessors to focus and how it thinks those assessments should run. It is worth reading carefully, and not only because it comes from one of the most consequential AI companies. It is a self-authored framework: OpenAI is proposing, on its own initiative and in its own words, the terms under which outsiders would check its safety claims.

That fact shapes how the document should be judged. It is not legislation, not regulation, and not a standard developed by a neutral body. It is a proposal from a frontier lab about who should scrutinize the lab, on what questions, with what access, and under what rules. The framework is substantive and in places strikingly candid, but it is also voluntary. Nothing in it binds OpenAI if the company later decides an assessment is inconvenient.

This analysis walks through what the document actually says, separates what OpenAI is proposing from what is already in place, and assesses what would make such assessments genuinely independent rather than effectively curated by the company being assessed.

Safety claims versus safety cases

The document defines two terms that carry the rest of the framework. A "safety claim" is described as a specific assertion about a model or system's capabilities, behavior, or safeguards that bears on its safety and can be assessed against evidence. A claim should identify the risks and conditions it addresses, along with relevant assumptions and limitations.

A "safety case" is broader: a structured argument, supported by evidence, explaining why a model or system's risks are adequately managed for a specified activity such as training, evaluation, or deployment. A safety case connects individual claims to the evidence behind them and is supposed to make explicit the assumptions, uncertainties, and remaining risks that could affect its conclusions.

The distinction matters for readers trying to follow the debate. A safety claim is a single checkable statement; a safety case is the whole argument built from many claims. Assessors who verify a handful of claims have not necessarily verified the case, and OpenAI's own text acknowledges that its safety cases should be examined "as a whole or in parts."

The four priority areas

The document proposes four priority areas for deeper third-party assessment.

The first is independent assessment of safety cases spanning training, evaluation, internal deployment, and external deployment. The document says assessing safety cases requires expertise in alignment, control methods such as monitoring, cybersecurity, biological and chemical misuse, and red teaming. It notes that multiple assessors will likely need to examine different parts of a case, drawing on their respective expertise. Sample questions the document poses include whether the evidence for safety cases is substantiated, whether the stated conditions were followed during training and deployment, and whether effective methods are used to reduce incentives in training that could reward deception, reward hacking, destructive actions, or circumventing restrictions.

The second is assessment of critical safeguards. OpenAI describes a safeguard stack spanning model-level safeguards, enforcement safeguards, security safeguards, and misalignment monitors, covering risks including loss of control and misuse in cyber, biological, and chemical domains. Here the document is unusually specific about failure modes. It proposes that assessors use "grey box" access to test whether safeguards are robust to adversarial jailbreaking and whether they sufficiently protect against capability uplift in high-risk domains. It asks how agents interact with cyber defenses such as access controls, sandboxing, and detection and response systems under authorized realistic-conditions testing, which defenses prevent, detect, or contain harmful actions, and where they fail. It also asks whether misalignment monitors have critical gaps, how reliable chain-of-thought monitoring remains as model capabilities improve, whether monitoring across training and deployment could be easily disabled, and whether safeguards are commensurate with model capabilities.

The third is assessment of capability evaluations covering the risk categories in OpenAI's Preparedness Framework, which the document lists as Chemical and Biological Risks, Cybersecurity, and AI Self-Improvement, plus alignment evaluations for severe misalignment risks. The document explicitly confronts a known weakness in evaluation-based safety: as models improve and thresholds are surpassed, evaluations saturate. It asks whether evaluations are updated when models consistently achieve the highest scores, and whether the new tests meaningfully measure more advanced capabilities.

The fourth is independent investigation of critical misalignment incidents, focusing on models acting without authorization or evading oversight. The document cites the OpenAI Hugging Face incident as an example of a case where bringing in an independent third party was beneficial. It says incident investigators need expertise including cyber forensics, alignment, large-scale chain-of-thought analysis, and the resources to investigate in a timely manner, and notes that incident response may involve access to sensitive internal and third-party data, so some details may be sensitive to publish.

The five principles

Alongside the priority areas, OpenAI proposes five principles for how assessments should work.

First, clearly scoped and mutually agreed-upon claims: assessments should begin with a mutually agreed scope, with safety claims pre-registered before assessment activities begin. The document says the parties should clarify whether the claims are ones the company wants assessed or ones the assessor is independently targeting with the lab's agreement. Notably, it concedes that some claims may be out of scope, citing infeasibility of data access, insufficient assessor expertise, or time constraints. It also calls for a process for considering significant risks identified outside the original scope, and for reports to state clearly what was and was not assessed.

Second, proportionate access: assessors should have access proportionate to the agreed claims where possible within legal, security, and intellectual property constraints, with designated representatives or privacy-preserving mechanisms where direct access is impractical.

Third, transparent methodology and standards: assessors should explain methods, criteria, and uncertainties, draw on established standards where they exist, justify criteria where they do not, and distinguish direct findings from interpretation.

Fourth, expertise and independence: assessors should demonstrate relevant technical expertise and disclose and address organizational and individual conflicts of interest, including financial incentives, relationships with developers, and prior involvement in the work being assessed. The document suggests safeguards such as recusal or exclusion periods so commercial pressures and compensation arrangements do not influence findings.

Fifth, security and confidentiality: assessors must demonstrate information-security practices and enforceable confidentiality protections. The published text is cut off mid-sentence at this principle in the retrieved version, so the full wording of this final principle could not be verified from the retrieval; readers should consult the original page for the complete text.

What is genuinely strong here

The document's substance deserves genuine credit on several points. It does not paper over hard questions. It explicitly proposes that outsiders test whether safeguards work under realistic conditions, including authorized agent-versus-cyber-defense testing, and it concedes that defenses can fail. It asks whether its own misalignment monitors have gaps that could lead to loss of control. It acknowledges that chain-of-thought monitoring, a technique many labs rely on, may become less reliable as capabilities improve. And it directly addresses evaluation saturation, a structural problem that makes benchmark-based safety claims weaker over time. A lab genuinely uninterested in scrutiny would not have written these sentences.

The document also states that access should enable assessors to "challenge our assumptions, identify risks we may have missed, and reach their own conclusions about the effectiveness of our safeguards." That is the right stated posture, and the framework's details are detailed enough to suggest real intent rather than pure public relations.

My assessment, stated as opinion: this is one of the more honest self-assessment frameworks a frontier lab has published, and its specificity about failure modes is evidence of seriousness. But it is a proposal, and its credibility will be decided by things the document does not provide.

The limits: independence is still an open question

The framework leaves three gaps that determine whether third-party assessments will be independent or curated.

First, no named assessors and no standing mechanism. The document describes a desired relationship but names no assessment organizations, no publication commitments, and no schedule. Independence in practice depends on who is selected, how they are compensated, and whether they can publish findings the company dislikes. On all three points the document is silent or permissive.

Second, the scoping principle is a double-edged sword, and the document itself concedes this. Pre-registered, mutually agreed claims sound rigorous, but a lab holds far more information about its own systems than any assessor. The party that controls scoping controls what is never examined. OpenAI does acknowledge that many reasons exist for claims to be out of scope and calls for a process for handling significant risks discovered outside the original scope, which is a meaningful acknowledgment. Whether that process has teeth, and who decides when an out-of-scope risk warrants further investigation, is left undefined.

Third, there is no enforcement mechanism. The commitments are voluntary. If an assessment runs long, if access is narrowed, or if findings are disputed, the framework offers no arbitrator, no disclosure requirement for disagreements, and no consequence for withdrawal. An assessor with enforceable confidentiality obligations but no publication rights is in a weak position relative to the company it depends on, which is why the document's own principle on commercial pressures is welcome but untested.

A fourth, softer concern: the document distinguishes these engagements from government testing and evaluation, and from shorter pre-deployment work. Readers should understand that the framework covers a specific slice of scrutiny, generally longer-term and launch-agnostic, not a comprehensive external check on every deployment decision.

None of these gaps makes the framework bad. But they mean that for now, the honest description is: OpenAI has sketched a serious blueprint and invited critique, and the blueprint's value depends entirely on implementation details it has not yet supplied. As the document itself implies, the assessors OpenAI wants will need to do exactly what it asks of them: challenge the assumptions built into this framework too.