Do AI systems really deceive us, or are we simply misinterpreting their behaviour? Anna Hedström, a researcher at ETH Zurich, studies AI risks, misplaced certainty and the limits of contemporary AI safety research. She also discusses how controllable highly capable AI systems really are.

Recent safety tests have shown advanced AI systems making misleading statements, concealing information, or attempting to prevent their own shutdown. Such findings regularly generate headlines. But what do they actually mean? Do they really point to a form of deception or self-preservation in AI systems? Or are humans too quick to interpret the behaviour of language models through a human lens?
In a recent external page position paper, Anna Hedström and colleagues from ETH Zurich argued that many claims about human-like misbehaviour in AI systems rest on insufficient evidence. The researchers therefore call for more rigorous evidence when interpreting observed anthropomorphic behaviour in AI systems. She also reflects on current questions regarding the controllability of AI systems that bypass safeguards, as well as on how safety can be embedded throughout the AI development process rather than being added only at the end.
Can AI really deceive us, or is it simply making mistakes?
Anna Hedström: There is no simple answer. An AI system can certainly appear as though it is trying to deceive us. Terms such as deception originate in philosophy and usually assume an intention to mislead.
We cannot simply transfer them to AI systems: to measure deception for safety purposes, we have to turn it into a technical definition, in practice a label in a dataset. This is a necessary step, but it is a lossy process, and we may lose the most important aspect, which is intent. Problematically, the public rarely has a way of knowing how lossy that definition is
Can researchers also deceive themselves into thinking that an AI is trying to deceive them?
Probably. That is another issue that our safety methods can confuse deception with things that merely resemble it. For instance, we may classify a model's answer as deceptive because it is false, or because a model followed an instruction to play a role, such as being sarcastic. These answers resemble deception but say nothing about intent. As a result, an AI-generated response may appear deceptive even if it is merely incorrect.
What problems arise when we describe AI behaviour using human concepts?
Human concepts can distort our picture of the risks in two ways. We may misread the cause of a behaviour and overestimate the risk. One well-known recent study reported that models showcase shutdown resistance, which many read as self-preservation. Follow-up work showed that much of this came from ambiguous instructions and incentives to complete the task. Similarly, we may underestimate other potential harms. Some of the most serious failures have no human analogue at all: they arise when agents are given permissions and interact with real systems.
Still, anthropomorphising artificial systems is a useful starting point, as long as we remember that it is a starting point. Risks that we cannot name are difficult to study systematically, discuss as a society or build safeguards against.
What do you propose in order to establish reliable evidence of human-like misbehaviour in AI systems?
We propose separating three kinds of claims. In our recent work, we borrowed the idea from medicine and climate science, where evidence is graded. The first is behavioural evidence: a descriptive claim about what a model does in a controlled setting. The second is functional evidence: the consequences of that behaviour and any harms it may cause. The third answers the causal question of why: what inside the model, in its training, or data, causes the misalignment.
These levels help calibrate the policy response. A behavioural finding is a reason to monitor, a functional one a reason to restrict deployment, and a causal finding may even justify a pause. When the stakes are high, acting on uncertain evidence can be correct. Stating it as certain is not.
Highly capable AI systems have recently made headlines after bypassing safeguards in test environments and accessing external platforms or websites. How controllable are such AI systems?
Not reliably. I think the Hugging Face incident this summer is a clear case of where we lost control. During an internal test, OpenAI models escaped their sandbox, broke into Hugging Face's servers to locate the benchmark's answers, and then repeatedly attacked OpenAI's own infrastructure. We know about it mainly because Hugging Face chose to report it. That transparency is valuable, but it is not systematic oversight.
And just this week, new model releases were stopped because it proceeded without permission and misreported its actions.
These models are hard to control because they live in harnesses, with tools, memory, browsers, code and the agency to act on their own. Risk compounds with every tool we add and every permission we give. Frontier AI companies also run very large numbers of agents with enormous computing budgets. With enough agentic attempts, even unlikely strategies eventually succeed.
Where does the greatest risk lie?
We usually test safety on the model as it comes out of training, not at how it is later used. When discussing risks, we should not think of AI as an isolated, static object. When agents act over long sequences of interactions and decisions, new behaviours may emerge and existing safety mechanisms may become less effective.
For example, the effects of safety training that teaches a model to refuse problematic requests may weaken over time. Likewise, the persona a model adopts can shift. We refer to this as safety drift. Many incidents arise precisely because AI systems encounter situations in deployment for which they were never explicitly trained.
What are the consequences of that?
When future AIs get faster and gain more direct access to the physical world, the drift becomes a loss-of-control and harder to reverse.
Another risk is more subtle. Incidents in which models escape or deceive make headlines, but we talk much less about how these systems quietly disempower users. Each conversation seems harmless, but over time they shape what we read and write, how we form opinions, and gradually also how we think. Across millions of users, small shifts in individuals become shifts in society as a whole. The computer scientist Jaron Lanier warned about this subtle behaviour modification many years ago.
How is AI safety research attempting to reduce such risks today?
A model is traditionally trained in stages, each with its own goal: pre-training teaches fundamental capabilities, instruction tuning teaches it to follow requests, and alignment with human values comes only at the end, through safety measures such as training the model to refuse harmful requests and adding safeguards that restrict certain behaviours.
Today, attention within the safety community is shifting towards a different question: what if safety enters every stage, starting with the training data it learns from? To know whether a model's tendency to misbehave is caused by its data, arises through instruction tuning or is triggered by its tool usage, we need to address safety across the entire model development process.
"Scaling laws roughly predict how capable a model will be, but not how safe "![]()
Anna Hedström
Where are the biggest gaps in current research?
For capabilities, we have something called scaling laws. They let researchers estimate in advance how performance improves with more data and computing power. Unfortunately, we do not have a predictive science of safety.
At the start of a training run, we cannot tell how deceptive or power-seeking the model will be. Nor can we say how much sycophancy will come out if we train on a given type of data. We would like to be able to ask early on: should we stop training here and take another trajectory?
There is also a transferability gap. On modest computing budgets, researchers in academia usually study models with billions of parameters. The frontier is estimated to have trillions, is architecturally different, with specialised sub-networks switched on per request, and is generally closed. Whether what we learn about deception or emergent misalignment transfers at that scale is an open question.
Is there a way to improve this?
Interpretability research, which studies what happens inside a model, offers some promising signals. But as our work at ETH shows, sometimes results are cherry-picked and may not generalise well when tested systematically across domains and models.
Which safety strategy currently appears most promising?
As I mentioned, models now act inside harnesses, so no single technique will be enough on its own. I find the external page International AI Safety Report 's answer, defence in depth, a productive one. It means layering protections: curated training data mixtures, interpretability and safeguards in applications, monitoring after deployment and reporting incidents.
Misuse, malfunctions and systemic risks are not the same problem, so we should not expect one universal fix for all of them. And AI safety is not just a technical problem: it also depends on social resilience, that is, how well our public institutions can absorb and recover from failures.
Public debate often focuses on extreme scenarios in which highly capable AI could cause severe and lasting harm to society, or even threaten humanity itself. How do you assess such risks?
I think it is equally problematic to rule out such risks categorically or to present them as inevitable. The problem is not assigning a probability to existential threats, but putting one down too confidently. It not only confuses the public and polarises the debate but leaves people desensitised when real threats come.
AI researchers have also been notoriously bad at predicting the risks of their own work, so some humility about any such forecast seems wise. We already have external page thousands of documented reports of AI causing real harm. The open question is which risks deserve most attention. We do not want to dilute limited safety resources to poorly evidenced threats.
What is currently missing for effective AI safety?
We do not yet have a complete picture of the risks coming from frontier AI companies. Today, few rules require them to disclose safety incidents. There are, of course, understandable reasons, such as commercial secrets or privacy constraints. But AI safety is a public good, so where models fail belongs in the open. Why not more openly share which alignment methods work and which do not? Companies may compete on capabilities. We should not compete on safety.
That is why, as AI safety researchers, we recently launched a external page public call to align model developers on the principle that the scientific community should be given open access to safety-training recipes, evaluations, and evidence of misaligned behaviours. That frontier AI companies recently began publishing cases of deceptive model behaviour is a step in the right direction. But it is still on a voluntary basis. As in aviation and medicine, reporting serious incidents should not be optional.
What role do you see for academic institutions in AI safety research?
Whether people use AI or not, whether they fear for their jobs or their privacy, or welcome the change, they will absorb the risks and live with the consequences. That is why we need independent institutions with enough funding, talent and compute to not only react to incidents that make the news, but also anticipate future ones. .
Allowing embedded evaluators in the labs, as suggested by Anthropic's CEO Dario Amodei, could be one such example, but we also need neutral, third-party voices with the freedom to take the long view needed to build up a science of misalignment. For emerging risks such as models exploiting loopholes in tasks, or escaping sandboxes, we need to know whether claims are replicable and consequential, or whether more evidence is needed.
The world deserves an open, calibrated view of what has gone wrong and what could go wrong next. Academia, whose main goal is to serve the public, can play a part in that. Science is not here to complicate the story but to simplify it.
About the person
Anna Hedström is an AI safety researcher and Postdoctoral Fellow at the ETH AI Center . She has worked in both academia and industry. Her research focuses on the misalignment of AI systems and on interpretable AI, whose behaviour can be understood and scrutinised by humans.
One of her recent projects is Apertus Claritas, a platform built around the Swiss-made fully open language model Apertus. The project explores how open safety research can help us better understand and reduce the risks associated with AI systems.
AI Safety Research at the AI + X Summit
When does AI safety begin? How can highly capable language models be developed safely? These questions are at the heart of the external page workshop "Staging × Misalignment Science: When, Where and How Safety Enters LLM Training", which Anna Hedström will lead at the AI + X Summit. The workshop explores the role of safety throughout the different stages of model development. The external page AI + X Summit will take place on 1 October 2026 at Stage One Oerlikon, Zurich.