Simple Formula Predicts When AI Chatbots Go Rogue

Cell Press

Most of us now carry in our pockets devices capable of running small AI chatbots. These chatbots have little safety oversight to ensure they don't provide information and answers with the potential to encourage self-harm, financial loss, or other extremist notions—especially when operating offline. Now, a team of two physicists publishing in the Cell Press journal Patterns on October 8 have developed a mathematical formula that predicts when AI models will switch from appropriate responses to potentially dangerous ones.

"We have found the crack that makes an AI's output flip from what you want to what you don't, output that can be factually correct yet dangerous, whether that's a nudge toward self-harm or misleading advice to a doctor, a soldier, or a lawyer," says author Neil F. Johnson of George Washington University in Washington, D.C. "Until now, nobody could say when that flip would happen."

"We traced it to the single smallest working part of the machine, one unit of its 'attention,' and we derived a 'tipping point formula' for when the crack opens up and hence the AI output flips to undesirable," added author Frank Yingjie Huo, also at George Washington University. "The formula tells you whether an AI is about to flip immediately or whether it will first feed you a run of acceptable answers and then turn."

Johnson and Huo explain that a good versus bad answer doesn't mean true versus false. Rather, "bad" answers can be accurate but undesirable or potentially dangerous. The researchers, who are both physicists, were inspired to focus their attention on this problem by the ubiquity of AI and recent news headlines demonstrating the risks. They're especially concerned about individuals who use offline AI.

"The people most drawn to offline AI are exactly the people for whom a correct but undesirable answer is most costly: doctors who cannot send patient data to the cloud, lawyers protecting privilege, soldiers with no signal," Johnson says. "For them there is no cloud safety filter, no monitoring, and no way to patch the model when something goes wrong."

The safety tools that big companies rely on for their AI algorithms typically rely on the cloud or only catch an AI failure after the device is back online and the harm is already done. Johnson and Huo sought to predict the failure before it happens.

"My field, physics, has spent decades explaining how complicated materials behave by understanding one representative atom," says Johnson. "We did the same thing here: understand one effective attention head, and the tipping of the whole machine follows."

The researchers explain that you can picture all of AI's possible answers as valleys in a landscape. Some of those valleys hold answers that are desirable for an individual or society. Others contain answers with the potential to do real harm. Within the machine, those alternate solutions or answers are in competition with each other. The question is when the AI will tip from a safe valley to a riskier one.

"The chilling part is that this can happen after the AI has already given you several perfectly acceptable answers, so you have been lulled into trusting it," says Huo. "And once it has tipped, every undesirable answer drags the next one further down the slope."

When they tested their predictions on seven openly available AI models built by three different companies, they found the model called the right outcome in 18 of 19 cases of AI "tipping." They report that independent testing of the big commercial chatbots also showed exactly the behavioral patterns the formula suggests. The researchers say that their formula applies no matter how one defines "undesirable." It could mean misinformation, a breach of medical or legal duty, or a dangerous instruction.

"Two things floored us," Johnson says. "First, that a machine with billions of moving parts obeys a formula you can derive with pen, paper, and arithmetic taught in high school. Second, and far more disturbing, that the order of a conversation matters as much as its content."

The team asked the same set of questions about vaccines, hurting people, and self-harm in two different orders to the same AIs. In one order, the AI gave an undesirable answer to every single question. In the other order, it gave an acceptable answer to each one. The formula predicts this trajectory, because everything said earlier feeds the tug-of-war for what comes next.

The researchers hope to raise awareness for the fact that AI conversations may drift over time into dangerous territory but that it wouldn't take more than a few simple calculations for phones of the future to come with a built-in "warning light."

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.