OpenAI Expands Outside Safety Reviews Into Model Training
OpenAI wants outside safety evaluators involved before its models are finished, not just shortly before launch.
TechRepublic
Publisher
Sep 23, 2026 at 4:33 PM UTC · Updated hace 5 horas · 3 min de lectura

OpenAI wants outside safety evaluators involved before its models are finished, not just shortly before launch.
The company said Tuesday that it plans to support independent technical safety assessments during training, evaluation, and deployment, expanding beyond the pre-launch reviews it has relied on more heavily in the past.
Lama Ahmad, who leads OpenAI’s work with outside safety experts, told Bloomberg that the company previously focused on bringing third parties in shortly before launch. “As the stakes get higher, we want to make sure we’re also looking at things like training and evaluation, which do have high stakes, in addition to our deployments,” Ahmad said.
OpenAI said the assessments should have strong independence mechanisms, scientific rigor, robust security practices, and clear responsibilities.
The company said it is already talking with multiple potential assessors, including AI research groups METR and Redwood Research. Both groups previously investigated an incident involving OpenAI models gaining unauthorized access to Hugging Face systems.
What outside groups will examine
OpenAI outlined four areas where it wants deeper independent scrutiny.
Assessors could examine whether the evidence behind the company’s safety cases supports its claims, including safeguards used during training, evaluation and deployment. They could also test critical safeguards against jailbreaks and other adversarial attacks and examine defenses against high-risk capabilities in areas such as cybersecurity and biological misuse.
Article Intelligence
Topics
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
