Here’s a Way to Predict When AI Chatbots Will Turn Bad
Physicists at George Washington University say a formula can estimate when an AI chatbot will flip from good answers to bad ones, and early tests on small models back it up.
Jose Antonio Lanz
Publisher Decrypt
Oct 10, 2026 at 3:01 PM UTC · 4 min de lectura

- George Washington University physicists Neil Johnson and Frank Yingjie Huo published a formula that estimates how many good tokens an AI model produces before its first bad one.
- In the preprint, the formula correctly predicted whether a model would tip immediately or after a delay in 15 of 16 clear-cut cases.
- The authors propose a parallel monitor that flags when models below a safety threshold.
Physicists at George Washington University have published a formula that estimates how many good answers an AI chatbot will give before it slips into a bad one, and early tests suggest it works.
The study, by Neil Johnson and Frank Yingjie Huo, appeared in the journal Patterns and builds on a preprint, a version posted publicly before formal peer review, first released in February.

Chatbots can answer sensibly for a long stretch and then veer into something harmful, such as bad advice on self-harm or extremist talk, and there has been no simple way to predict when the swerve will happen. The authors argue that existing safety tools often depend on a cloud connection that offline models lack.
Johnson and Huo trace the problem to the attention head, the part of an AI model that decides which earlier words in a conversation matter most when choosing the next one. As a chat grows, the accumulated context pulls that attention toward one cluster of possible answers or another, until it tips.
Article Intelligence
Topics
Sponsored
AdNewsLayer Premium
Unlock deeper intelligence.
Ad-free reading, exclusive research, and real-time onchain insights.
Go Premium
