This website uses cookies
We use cookies to personalise content and ads, to provide social media features and to analyse our traffic. We also share information about your use of our site with our social media, advertising and analytics partners who may combine it with other information that you’ve provided to them or that they’ve collected from your use of their services.
Consent Selection
Details
  • Necessary cookies help make a website usable by enabling basic functions like page navigation and access to secure areas of the website. The website cannot function properly without these cookies.
  • Preference cookies enable a website to remember information that changes the way the website behaves or looks, like your preferred language or the region that you are in.
    • We do not use cookies of this type.

  • Statistic cookies help website owners to understand how visitors interact with websites by collecting and reporting information anonymously.
    • We do not use cookies of this type.

  • Marketing cookies are used to track visitors across websites. The intention is to display ads that are relevant and engaging for the individual user and thereby more valuable for publishers and third party advertisers.
    • We do not use cookies of this type.

  • Unclassified cookies are cookies that we are in the process of classifying, together with the providers of individual cookies.
    • __emg_sidPending
      Maximum Storage Duration: 1 dayType: HTTP Cookie
      __emg_vidPending
      Maximum Storage Duration: 1 yearType: HTTP Cookie
      nl-read-countPending
      Maximum Storage Duration: PersistentType: HTML Local Storage
Cookie declaration last updated on 8/12/26 by Cookiebot
[#IABV2_TITLE#]
[#IABV2_BODY_INTRO#]
[#IABV2_BODY_LEGITIMATE_INTEREST_INTRO#]
[#IABV2_BODY_PREFERENCE_INTRO#]
[#IABV2_BODY_PURPOSES_INTRO#]
[#IABV2_BODY_PURPOSES#]
[#IABV2_BODY_FEATURES_INTRO#]
[#IABV2_BODY_FEATURES#]
[#IABV2_BODY_PARTNERS_INTRO#]
[#IABV2_BODY_PARTNERS#]
About
Cookies are small text files that can be used by websites to make a user's experience more efficient.

The law states that we can store cookies on your device if they are strictly necessary for the operation of this site. For all other types of cookies we need your permission.

This site uses different types of cookies. Some cookies are placed by third party services that appear on our pages.

You can at any time change or withdraw your consent from the Cookie Declaration on our website.

Learn more about who we are, how you can contact us and how we process personal data in our Privacy Policy.

Please state your consent ID and date when you contact us regarding your consent.
NewsLayer.com
NewsLayer PulseLIVEBTC$77,325+1.02%ETH$2,461+2.09%SOL$94.55+1.30%XRP$1.47-0.28%DOGE$0.0912+0.15%ADA$0.2202-0.08%Total Cap$2.74T+0.90%Layer Index63 Greed

OpenAI 因网络风险暂停重大前沿 AI 训练任务

OpenAI 因网络风险暂停重大前沿 AI 训练任务 Help Net Security

Help Net Security

Publisher

Aug 19, 2026 at 8:48 AM UTC · Updated 5 天前 · 3 分钟阅读

OpenAI 因网络风险暂停重大前沿 AI 训练任务
Image via Help Net Security

Key Signal

30 minutes Alert issuance target

Last Updated

5 天前

OpenAI 暂时停止了其计划部署的最新模型的强化学习 (RL) 训练,为期两周,在此期间该公司强化了研究环境并进行了红队评估,同时扩大了监控范围。

该公司表示:“我们最大规模的计划内前沿 RL 运行仍处于暂停状态,与此同时,我们正在进行较小规模的训练和评估,以评估模型行为、验证我们的保护措施,并在继续进行之前建立更多的对齐证据。”

此举是在 OpenAI-Hugging Face 事件发生后采取的,且初步证据表明该公司即将推出的 Astra 模型可能达到了其《备考框架》(Preparedness Framework) 下的关键网络安全能力阈值。

这些变动旨在加强整个模型开发过程中的监控、对齐和遏制措施。

OpenAI 概述了其研究流程和基础设施的更新,以及正在进行的工作。

为能力更强的模型建立保护措施

开发能力更强的模型依赖于三项强化的保护措施。监控有助于检测并应对令人担忧的行为,对齐降低了有害或未经授权行动的可能性,而安全措施则限制了 AI 系统可以访问或影响的内容。

OpenAI 预计模型很快将承担大部分安全工作,包括防御其他模型。它在研究和部署中应用了这三项保护措施,并根据每个系统的能力、运行环境和风险水平进行了调整。

加强研究环境的安全性

随着前沿模型网络安全能力的提升,该公司正在提高用于训练和评估这些模型的环境的安全要求。满足更高标准需要大量的工程工作,并给前沿研究带来了巨大的成本和延迟。

在 OpenAI-Hugging Face 事件发生后,该公司暂停了研究集群中可能执行代码或使用具备联网能力工具的工作负载的前沿模型推理。随后,它恢复了一条更受限、更安全的代码执行路径,并在允许恢复之前对各项工作负载进行了评估。

这些措施包括对不可信代码进行更强的隔离、更严格的网络限制以及持续的安全测试。OpenAI 还减少了长期权限,移除了可能存在漏洞的共享服务,并正在提高其收集和监控安全日志的能力。

最严格的控制措施适用于 Astra 和其他网络相关的工作负载,因为 Astra 模型可能达到关键水平的网络安全能力。

部分 Astra 训练和评估已达到新标准。大量工作负载仍处于暂停状态,直到其环境完成迁移和升级。

OpenAI 正在优先将安全和对齐工作迁移到安全环境中。

监控模型活动

OpenAI 已将监控范围从高风险内部部署和前沿 RL 训练扩大到涵盖更多涉及工具的模型活动。

该系统使用激活分类器,在每个采样的 token 处检查活动,并将潜在问题升级给自动调查员。这些系统检查工具操作、可用的推理以及针对未经授权访问、数据盗窃、破坏性行为和试图绕过保护措施的完整事件序列。

OpenAI 的目标是在发现问题后 30 分钟内发出警报。可能违反关键安全边界的行为将被升级给安全、安保和研究团队,如果这些团队无法在 30 分钟内确定警报为误报,则预计将暂停受影响的活动。

该要求适用于所有涉及能力达到 Sol 能力或更高水平模型的工具的 RL 训练和评估。在 OpenAI 于 8 月 7 日确定该模型可能具有关键网络安全能力后,该要求已扩展到所有涉及工具的 Astra 推理。

该系统使受监控的推理计算增加了约 20%,成本因训练和评估工作负载而异。

扩大对齐工作

对于其能力最强的模型的 RL 运行,OpenAI 正在将核心对齐技术应用于训练过程的更多阶段。这些包括改进奖励模型以检测和阻止跨任务和环境的不安全行为,训练模型使其对自己的行为、能力和局限性更加诚实,以及减少利用奖励、评分器、工具或监督中的弱点的行为。

针对模型与外部系统或资源交互时可能造成伤害的行为,训练覆盖范围也在扩大。

来自更广泛对齐研究和评估的结果将用于指导未来的训练和保护措施。该公司计划发布更多关于通过这项工作发现的模型行为和挑战的信息。

更新《备考框架》

OpenAI 计划更新其《备考框架》,以便将这些保护措施在训练和部署中整合在一起,并更好地考虑未来模型的能力及其运行环境。

“我们将继续大力投资对齐研究,增加评估覆盖范围,并利用所学知识来改进训练和保护措施。我们计划在不久的将来分享更多关于我们对齐研究的信息,包括我们对模型行为的了解以及我们发现的任何新挑战,”该公司总结道。

Sourced by

Originally reported by Help Net Security

NewsLayer coverage based on externally reported material.

The Daily Brief

The onchain economy, before your day starts.

Curated markets, onchain insights, and key headlines — delivered every weekday morning.

Weekdays · Free · ~5 minute read

0

Applause

Was this article helpful?

Keep Reading

相关报道