OpenAI wants outside safety evaluators involved before its models are finished, not just shortly before launch.
The company said Tuesday that it plans to support independent technical safety assessments during training, evaluation, and deployment, expanding beyond the pre-launch reviews it has relied on more heavily in the past.
Lama Ahmad, who leads OpenAI’s work with outside safety experts, told Bloomberg that the company previously focused on bringing third parties in shortly before launch. “As the stakes get higher, we want to make sure we’re also looking at things like training and evaluation, which do have high stakes, in addition to our deployments,” Ahmad said.
OpenAI said the assessments should have strong independence mechanisms, scientific rigor, robust security practices, and clear responsibilities.
The company said it is already talking with multiple potential assessors, including AI research groups METR and Redwood Research. Both groups previously investigated an incident involving OpenAI models gaining unauthorized access to Hugging Face systems.
What outside groups will examine
OpenAI outlined four areas where it wants deeper independent scrutiny.
Assessors could examine whether the evidence behind the company’s safety cases supports its claims, including safeguards used during training, evaluation and deployment. They could also test critical safeguards against jailbreaks and other adversarial attacks and examine defenses against high-risk capabilities in areas such as cybersecurity and biological misuse.
Other work could focus on whether capability evaluations adequately measure risks covered by OpenAI’s Preparedness Framework, as well as whether alignment evaluations can detect serious misalignment.
The company also wants independent investigations into selected incidents involving models acting without authorization or attempting to evade oversight. OpenAI said some assessments could last weeks while others could run for several months, meaning the work is intended to examine safety claims over time rather than serve solely as a final pre-launch check.
The hard part is independence
OpenAI’s proposal comes as AI companies face growing pressure to prove that their safety claims can withstand scrutiny from organizations outside the companies building the models.
The company says assessors should agree on clearly defined claims before testing begins, receive access proportionate to what they are evaluating, and disclose conflicts of interest. It also says sensitive information may require assessments to happen on company-managed devices or at company facilities.
That creates a practical tension. Meaningful safety testing may require access to sensitive model data, internal systems, or security controls, but tight restrictions can also limit how independently outside researchers can validate a company’s claims.
OpenAI says it wants findings to be as transparent as possible while still protecting confidential information, proprietary technology, and security-sensitive details.
What it means for AI development
Conducting assessments during training allows teams to identify vulnerabilities when adjustments to model architecture or safety protocols remain viable, avoiding late-stage discoveries right around release.
However, earlier evaluation windows do not automatically guarantee robust oversight. The true value of this independent scrutiny rests on evaluator qualifications, the depth of system access granted, and the precise scope of claims cleared for testing.
Key operational details also remain unconfirmed. OpenAI has yet to name official assessment partners or outline public access guidelines. In contrast, Anthropic has taken a more concrete step, announcing last week that it would embed evaluators from Accenture to test its frontier models.
What it means for users
While end users of ChatGPT are unlikely to notice immediate changes, the primary value of this initiative unfolds behind the scenes. Independent evaluations could allow OpenAI to detect alignment, security, and misuse risks earlier in the development lifecycle, preventing issues before systems are deployed.
Ultimately, users stand to gain stronger safety protections if these third-party assessments yield actionable insights that OpenAI implements. However, full transparency may be limited, as certain disclosures could be withheld to safeguard proprietary systems or sensitive vulnerabilities.
OpenAI says it plans to expand the independent assessment ecosystem and work toward shared international standards for technical safety reviews as frontier models become more capable. The larger test will be whether outside evaluation becomes a meaningful check on frontier-model development or simply another layer of company-managed assurance.
Let us teach you How to Talk to AI for free! Try our six-minute course at The Neuron Academy and learn a few simple ways to write better prompts and get more useful results from AI, or browse our other AI course for free for seven days. Check out all the lessons here →
Read the full article here