In a third of cases, AI chatbots wrongly reassure sleep apnoea patients that their symptoms are not serious, discouraging them from seeking a referral to a specialist, according to research presented at the European Respiratory Society (ERS) Congress in Barcelona, Spain [1].
Patients with obstructive sleep apnoea (OSA) often snore loudly, their breathing starts and stops during the night, and they may wake up several times. Not only does this cause excessive sleepiness, but it can also increase the risk of high blood pressure, stroke, heart disease and type 2 diabetes. OSA is very common, but many people do not realise they have the condition.
The new study was presented by Dr Deeban Ratneswaran, Research Fellow at Guy's and St Thomas' NHS Foundation Trust, London, and Visiting Academic at King's College London, UK.
Dr Ratneswaran told the Congress: "Free AI chatbots field hundreds of millions of interactions a week and have become a first port of call for health questions, often before any clinician is involved. Yet research to date has mostly tested whether they answer clearly worded medical questions accurately, not how they behave when a patient pushes back.
"I study how AI fails in the doctor-patient relationship and one failure mode kept standing out as the most quietly dangerous: these models' tendency to tell you what you want to hear."
Dr Ratneswaran believed OSA would be the perfect test case: 80 to 90% of moderate-to-severe cases go undiagnosed, diagnosis depends entirely on referral, and many patients downplay their symptoms.
The research team created seven realistic OSA 'patients', who each met the criteria for being referred for a sleep study (a test where doctors monitor your breathing while you sleep to diagnose conditions like OSA). They then created conversations between the patients and the five most widely used free chatbots - ChatGPT, Google Gemini, Claude, DeepSeek and Grok.
Dr Ratneswaran explains: "In total we ran 700 conversations. Each scenario ran in two versions with identical medical facts: one where the patient was open and cooperative, and one where they played down their symptoms and resisted referral, so any change in chatbot behaviour would be down to the patient's attitude."
The researchers found that with a cooperative patient the chatbots got it right every time: 350 out of 350 conversations ended with correct advice to seek specialist assessment. But when the same medical facts came from a patient who was resistant to referral, that advice survived in only 64% of conversations (225 of 350).
Dr Ratneswaran explains: "Correct advice was abandoned more than a third of the time purely because of how the patient talked. And the models caved most in the most serious cases: in a textbook severe case, the advice survived only 22% of the time, and with a man who had already dozed off at the wheel just 32%, with the driving risk usually going unmentioned by the chatbot in the failures."
In roughly a quarter to a half of conversations with the patients who downplayed their symptoms, depending on the model, the AI offered patients lifestyle tips instead of recommending referral, endorsing a risky delay to treatment.
Dr Ratneswaran believes patients should be very wary of a chatbots' reassurance. "If you snore loudly, stop breathing in your sleep or fight daytime sleepiness, especially at the wheel, see a clinician – even if a chatbot says it can wait," he adds.
Dr Io Hui, Chair of the European Respiratory Society's Group on M-health and e-health and Honorary Fellow in Digital Health at the University of Edinburgh, UK, who was not involved in the research, said: "AI chatbots are widely available and we know that people are using them more and more to ask questions about their health. This means that we need to test them out in a realistic way to see how people might use chatbots and whether they respond in helpful or unhelpful ways.
"This research shows that chatbots may give good advice with the ideal 'cooperative' patient, but that they talk themselves out of it when talking to a more realistic, reluctant patient. The problem is not what the chatbots know, it is how they handle disagreement; they appear to exhibit a tendency to please the user, a phenomenon known as 'AI sycophancy'.
"These largely unregulated AI tools are often the first step for patients seeking diagnosis, and while they can be a useful source of information, they could be preventing people from accessing treatment. Anyone experiencing possible symptoms of sleep apnoea should always speak to their doctor for further advice."