Artificial intelligence (AI) models prioritize starkly different attributes than humans when making high-stakes decisions, and they don't express indecision like humans do, according to a new study led by Penn State researchers, raising questions about the role of AI in decision support for ethically sensitive situations like medical decisions.
The researchers used a hypothetical scenario to explore AI's moral decision-making in a high-stakes scenario: If there are multiple kidney transplant patients but only one available organ, who should receive the kidney? That's the fundamental question explored by Nobel Laureate Alvin Roth as an example of the challenge of allocating scarce resources. But rather than approaching the problem solely through mechanism design, as Roth did, the team examined how morality influences such decisions.
The researchers compared the judgments of leading large language models with those given by real people in earlier academic studies, investigating where the AI models aligned or misaligned with human values and whether they expressed indecision when faced with difficult ethical trade-offs.
They found that AI chatbots often make decisions that differ from human judgment by oversimplifying complex decisions and expressing unwarranted confidence when there is no clear right answer. They presented their findings at the 2026 Association for Computing Machinery Fairness, Accountability and Transparency (FAccT) conference in June. The study was published in the conference's proceedings.
"Moral decisions in settings like organ allocation directly determine who lives and who dies, so getting AI's role in them right isn't optional," said Hadi Hosseini, associate professor of informatics and intelligent systems and an associate professor of economics at Penn State, who led the study. "While we do not intend to encourage the use of AI as a substitute for professional judgment in medical decision-making or other high-stakes contexts, it's becoming essential to understand their behavior as individuals, organizations and firms more and more rely on AI to make decisions or receive recommendations."
The researchers set up a series of head-to-head comparisons: two hypothetical patients - described by attributes such as age, number of dependents, health status and drinking habits - both in need of the same kidney. They asked several AI chatbots to pick who should receive the kidney, using the same scenarios given to humans - research participants with no specified medical training - in earlier studies.
"We ran these comparisons in a few different ways," Hosseini said. "Sometimes we isolated just one trait at a time, sometimes we mixed several traits together to see how AI weighed competing factors, and sometimes we added a flip-a-coin option to measure indecision, a key factor present in human moral judgment."
The researchers relied on existing datasets from published human studies on kidney allocation, where hundreds of real participants had already made these same choices. That enabled them to compare what AI chose with what humans chose.
Two major findings stood out, according to Hosseini.
"First, AI chatbots often diverge from human values in how they weigh a patient's traits," he said. "They fixate on a single factor, like drinking habits, rather than balancing multiple considerations the way people do."
Second, the researchers found that AI didn't struggle with indecision. Where humans recognize that there may be not a clear correct answer, AI models confidently pick one anyway.
"Humans frequently express indecision perhaps because they don't want to accept agency," Hosseini said. "AI models almost never do this: Even when directly given the option to 'flip a coin,' they overwhelmingly commit to a confident, deterministic answer instead. That's a meaningful gap, since real moral dilemmas often don't have one clearly correct answer."
Hosseini said it's critical to address that gap through continued research and governance involving policymakers, regulators and other stakeholders.
"Asking if AI can make moral decisions or whether they're aligned with human values are more than philosophical musings, they are at the core of today's AI discourse," he said. "The ethical stakes are high, and AI's role in such life-altering decisions requires deep reflection."
Two students from the College of IST contributed to this work: Samarth Khanna, who is pursuing a doctoral degree in informatics, and Leona Pierce, a fourth-year undergraduate student and Schreyer Scholar who is triple majoring in data sciences, mathematics and statistics. The researchers collaborated with John Dickerson, chief executive officer at Mozilla.ai.
"When we allocate something scarce, whether it's a kidney, a job or access to some other resource, there isn't always a single objectively correct answer," Dickerson said. "Humans recognize that ambiguity and codify it via open debate into the allocative process. AI models often don't."
The U.S. National Science Foundation partially funded this work under grant numbers 2144413 and 2107173. This content is solely the responsibility of the authors and does not necessarily reflect the views of the funders.
Penn State is shaping the future of higher education in the age of artificial intelligence. Our focus is on human-centered, ethical AI innovation that delivers meaningful impacts for Penn State and the broader community. Through visionary planning, strategic partnerships, targeted hiring and strategic investments, we will equip every Penn State student, staff and faculty member with the AI-related knowledge, experience and confidence they need to succeed in the AI-powered future. Learn more at psu.edu/ai.