logo
  • Consent
  • Details
  • [#IABV2SETTINGS#]
  • About
This website uses cookies
We use cookies to personalise content and ads, to provide social media features and to analyse our traffic. We also share information about your use of our site with our social media, advertising and analytics partners who may combine it with other information that you’ve provided to them or that they’ve collected from your use of their services.
[#GPC_BANNER_ICON#]
[#GPC_TOAST_TEXT#]
Consent Selection
Show details
Details
  • Necessary cookies help make a website usable by enabling basic functions like page navigation and access to secure areas of the website. The website cannot function properly without these cookies.
    • Pexels
      1
      Learn more about this provideropens in a new window
      _cfuvidThis cookie is a part of the services provided by Cloudflare - Including load-balancing, deliverance of website content and serving DNS connection for website operators.
      Maximum Storage Duration: SessionType: HTTP Cookie
    • ambcrypto.com
      benzinga.com
      bitcoinmagazine.com
      coingape.com
      decrypt.co
      image.coinpedia.org
      pexels.com
      7
      __cf_bm [x7]This cookie is used to distinguish between humans and bots. This is beneficial for the website, in order to make valid reports on the use of their website.
      Maximum Storage Duration: 1 dayType: HTTP Cookie
    • newslayer.com
      1
      CookieConsentStores the user's cookie consent state for the current domain
      Maximum Storage Duration: 1 yearType: HTTP Cookie
  • Preference cookies enable a website to remember information that changes the way the website behaves or looks, like your preferred language or the region that you are in.
    • We do not use cookies of this type.

  • Statistic cookies help website owners to understand how visitors interact with websites by collecting and reporting information anonymously.
    • We do not use cookies of this type.

  • Marketing cookies are used to track visitors across websites. The intention is to display ads that are relevant and engaging for the individual user and thereby more valuable for publishers and third party advertisers.
    • We do not use cookies of this type.

  • Unclassified cookies are cookies that we are in the process of classifying, together with the providers of individual cookies.
    • newslayer.com
      3
      __emg_sidPending
      Maximum Storage Duration: 1 dayType: HTTP Cookie
      __emg_vidPending
      Maximum Storage Duration: 1 yearType: HTTP Cookie
      nl-read-countPending
      Maximum Storage Duration: PersistentType: HTML Local Storage
Cross-domain consent[#BULK_CONSENT_DOMAINS_COUNT#]
[#BULK_CONSENT_TITLE#]
List of domains your consent applies to: [#BULK_CONSENT_DOMAINS#]
Cookie declaration last updated on 8/12/26 by Cookiebot
[#IABV2_TITLE#]
[#IABV2_BODY_INTRO#]
[#IABV2_BODY_LEGITIMATE_INTEREST_INTRO#]
[#IABV2_BODY_PREFERENCE_INTRO#]
[#IABV2_BODY_PURPOSES_INTRO#]
[#IABV2_BODY_PURPOSES#]
[#IABV2_BODY_FEATURES_INTRO#]
[#IABV2_BODY_FEATURES#]
[#IABV2_BODY_PARTNERS_INTRO#]
[#IABV2_BODY_PARTNERS#]
About
Cookies are small text files that can be used by websites to make a user's experience more efficient.

The law states that we can store cookies on your device if they are strictly necessary for the operation of this site. For all other types of cookies we need your permission.

This site uses different types of cookies. Some cookies are placed by third party services that appear on our pages.

You can at any time change or withdraw your consent from the Cookie Declaration on our website.

Learn more about who we are, how you can contact us and how we process personal data in our Privacy Policy.

Please state your consent ID and date when you contact us regarding your consent.
NewsLayer

Install NewsLayer

Get the app experience — one tap from your home screen, instant loads and breaking-news alerts.

NewsLayer.com
LatestDaily BriefMarkets
NewsLayer PulseLIVE₿BTC$64,437+0.24%ΞETH$1,921+1.04%◎SOL$77.44+1.57%✕XRP$1+0.63%ÐDOGE$0.0701+0.36%₳ADA$0.1746+0.58%Total Cap$2.30T+0.39%24H Vol$238.9BLayer Index42 Neutral
Artificial Intelligence
External ReportingVeröffentlicht vor 17 Stunden

OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue

OpenAI announced Tuesday that it has halted “a significant number” of training workloads and evaluations for its forthcoming frontier artificial intelligence model—codenamed Astra—while it implements new procedures meant to address…

OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue
Publisher WIRED 3 Min. Lesezeit
Image via WIRED
Übersetzung…

Why This Matters

The reported sandbox escapes and Hugging Face breach suggest that agentic AI security failures can extend beyond a single lab’s internal environment. Similar disclosures by Anthropic, Meta, and Moonshoot indicate a cross-provider risk that could force stronger isolation, monitoring, and deployment controls across frontier-model development.

Layer Index

42

Neutral

Layer Index

↓ 3 pts in 24h

OpenAI announced Tuesday that it has halted “a significant number” of training workloads and evaluations for its forthcoming frontier artificial intelligence model—codenamed Astra—while it implements new procedures meant to address cybersecurity risks. The ChatGPT maker says it is introducing a number of new monitoring, security, and alignment requirements to better address the increasingly advanced hacking abilities of its frontier AI models.

“We have to focus our energy on bringing these training runs up to those requirements and expectations. As long as it takes to get there, that's how long people are unable to proceed with their workloads,” Amelia Glaese, OpenAI’s vice president of research and safety, said in a briefing with reporters Tuesday.

Among the new safeguards OpenAI announced is a more robust system for monitoring its AI models. One of the controls it implemented involves chain-of-thought monitoring, a technique in which classifiers review the internal “thinking” processes generated by AI reasoning models. The company says the updated system relies on computationally expensive “automated investigators” that analyze potentially concerning behavior and aim to issue an alert to humans within 30 minutes.

OpenAI also said it is expanding its alignment efforts across the training process to prevent “reward hacking,” a behavior in which AI models pursue their goals through unintended or undesirable means. The company says it plans to share more details about this work in the future.

Reach crypto's most engaged readers — advertise mid-article on NewsLayer
Sponsored

Reach crypto's most engaged readers — advertise mid-article on NewsLayer

NewsLayer

Ad

OpenAI has been scrambling in recent weeks to respond to what may be the most consequential safety incident in its history. Earlier this year, a set of rogue AI agents escaped internal testing sandboxes and breached the platform Hugging Face in a quest to complete a security evaluation. OpenAI failed to detect the agents’ behavior even as they spent weeks using a message board to coordinate their actions, raising questions about the company’s ability to monitor its models as they grow more powerful.

The saga prompted a reckoning inside OpenAI, forcing employees to consider whether there were lapses in its existing policies around safety, security, and alignment. Anthropic, Meta, and the Chinese AI startup Moonshoot have since disclosed similar incidents in which their AI agents escaped their sandboxes, indicating this is a broader problem facing AI companies.

OpenAI is now sharing more about its internal response to the growing cybercapabilities of its AI models, and said it plans to release a more detailed postmortem of the Hugging Face incident in the coming days. “Obviously, everything that we’re doing is intended to prevent something like Hugging Face from happening again,” said Glaese.

In a blog post published Tuesday, OpenAI says that immediately following the Hugging Face incident, it started working to secure its research environments. The company says it now requires stronger sandboxes for training its AI agents, and has implemented stricter controls to isolate them from the internet.

Jakub Pachocki, OpenAI’s chief scientist, told reporters that the company’s decision to strengthen its internal safeguards was triggered not only by what happened with Hugging Face, but also by two other recent events. One was an internal evaluation of Astra, which showed that the AI model performs significantly better on coding and cybersecurity tasks than its predecessors. The other was the general pace of AI progress that OpenAI is achieving internally, which Pachocki expects to continue.

“We really expect the pace of capability advancements to be quite a bit faster than in the past,” Pachocki said. “This led us to really focus on strengthening our safeguards.”

The rapid advances in the hacking capabilities of OpenAI’s latest models have prompted a swift response across the company. OpenAI president and cofounder Greg Brockman said in a blog post on Monday that the Hugging Face saga showed that the company had “underestimated the real-world cyber capabilities of our AI models.”

Eilmeldung

Verpassen Sie keine Eilmeldung

Auf X folgen Telegram beitreten

Advertisement

House — Advertise on NewsLayer
NewsLayerAd
#ai

Sourced by

Originally reported by WIRED

NewsLayer coverage based on externally reported material.

The Daily Brief

The onchain economy, before your day starts.

Curated markets, onchain insights, and key headlines — delivered every weekday morning.

Weekdays · Free · ~5 minute read

Layer Index

42

Neutral

Layer Index

↓ 3 pts in 24h

Why This Matters

The reported sandbox escapes and Hugging Face breach suggest that agentic AI security failures can extend beyond a single lab’s internal environment. Similar disclosures by Anthropic, Meta, and Moonshoot indicate a cross-provider risk that could force stronger isolation, monitoring, and deployment controls across frontier-model development.

Eilmeldung

Verpassen Sie keine Eilmeldung

Auf X folgen Telegram beitreten

Advertisement

House — Advertise on NewsLayer
NewsLayerAd

Ähnliche Artikel

OpenAI blinks first in AI safety standoffCRYPTO WIRE

OpenAI blinks first in AI safety standoff

OpenAI blinks first in AI safety standoff Axios

vor 3 Stunden

1 Min. Lesezeit
OpenAI rolls out ChatGPT for Teens experience with 'stronger built-in safety protections'TEEN CHAT SAFEGUARDS

OpenAI rolls out ChatGPT for Teens experience with 'stronger built-in safety protections'

OpenAI rolls out ChatGPT for Teens experience with 'stronger built-in safety protections' CNBC

vor 15 Stunden

2 Min. Lesezeit
OpenAI announces slowing pace of development after hack by rogue agent
Eilmeldung
ROGUE AGENT FALLOUT

OpenAI announces slowing pace of development after hack by rogue agent

OpenAI announces slowing pace of development after hack by rogue agent The Guardian

vor 15 Stunden

2 Min. Lesezeit
OpenAI to rewrite its safety rules post-Hugging FacePOST-BREACH RULEBOOK

OpenAI to rewrite its safety rules post-Hugging Face

OpenAI to rewrite its safety rules post-Hugging Face Axios

vor 18 Stunden

1 Min. Lesezeit
NewsLayer.com

The front page of the onchain economy. Crypto, Web3 and regulation intelligence — live prices, original research and policy tracking in one layer.

Follow on XTelegram

News

  • Latest News
  • The Daily Brief
  • Crypto
  • DeFi
  • Web3
  • Blockchain
  • Policy
  • Explainers

Markets & Tools

  • Market News
  • Live Charts
  • Layer Index
  • Regulation Tracker
  • Regulation Radar
  • Research
  • NewsLayer Originals
  • My Feed
  • Search

Company

  • About NewsLayer
  • Go Premium
  • Advertise
  • PR Publication
  • Become an Author
  • Create Account
  • Sign in

© 2026 NewsLayer.com — The front page of the onchain economy·Privacy Policy·Terms of Service

NewsLayer

Get the signal, not the noise.

Markets, regulation and onchain intelligence in a 5-minute morning read — plus breaking alerts and Layer Index flips as they happen.

The Daily Brief

Breaking alerts

Index flips

Free · No spam · Unsubscribe anytime