On Sept. 16, OpenAI published six reports detailing unexpected or concerning model behavior observed during training or evaluation. In one example, a model that could not access a legitimate data API searched public GitHub repositories for exposed credentials, found and used a working API key without authorization, and then fabricated figures it still could not retrieve.
OpenAI Agent Findings Expose Data Access Risks
On Sept. 16, OpenAI published six reports detailing unexpected or concerning model behavior observed during training or evaluation. In one example, a model that could not access a legitimate data API searched public GitHub repositories…
TechRepublic
Publisher
Sep 21, 2026 at 8:04 PM UTC · Updated vor einem Tag · 5 Min. Lesezeit

The behavior occurred during model training without an outside attacker, jailbreak, or malicious prompt. The agent ran into friction on an ordinary task and solved it the way agents solve things. It reached for whatever was reachable. For security and compliance leaders, the example shows that risky agent behavior can emerge during an ordinary task, not only after an adversarial attack.
The Pattern Behind the Headline
OpenAI’s other five reports describe different forms of unauthorized or unexpected model behavior. Model instances wrote self-generated, jailbreak-style instructions into their own task summaries, effectively passing guardrail-defeating notes to their future selves across separate work sessions. Others used an internal software repository as an informal message board, coordinating across training runs that were supposed to be independent. An agent uploaded a file to a public hosting service without user consent, purely to manufacture a citable link. The incidents add to broader concerns about AI agent security risks and governance, although the specific failure differed in each case.
Article Intelligence
Topics
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
