Can Large Language Models Capture Human Risk Preferences?

Tsinghua University Press

Large language models (LLMs) are increasingly used as agents to simulate human behavior, yet their fidelity in complex decision-making under uncertainty remains insufficiently understood. To address this gap, we develop a comparative framework that benchmarks LLM-simulated risk preferences against empirical human behavior. Using demographic profiles from surveys conducted in Sydney, Hong Kong, and Nanjing, we construct role-playing prompts and evaluate three LLMs on abstract lottery-choice tasks. We adopt the classical Constant Relative Risk Aversion (CRRA) framework as a domain-neutral "standard ruler" to compare risk attitudes. The analysis yields three main findings. First, off-the-shelf LLMs do not exhibit a universal risk profile: the two GPT models are more risk-averse than human benchmarks, whereas Gemini is more risk-seeking. Second, prompt language systematically affects simulated risk attitudes, with English-to-Chinese switching inducing a more conservative shift in most cases. Third, LLMs do not reliably reproduce the empirical heterogeneity of human risk preferences, tending either to generate overly concentrated distributions or unrealistically large dispersion. Taken together, these findings show that off-the-shelf LLMs remain vulnerable to model-family-specific miscalibration, language-sensitive distortions, and failures in distributional fidelity. Rigorous empirical calibration is therefore necessary before off-the-shelf LLMs can be reliably deployed in computational social science and choice modeling.

The team published their study in Communications in Transportation Research (https://doi.org/10.26599/COMMTR.2026.9640025 ) .

Our findings highlight an important limitation of using off-the-shelf LLMs as tools for behavioral prediction. Since different model families exhibit different baseline calibration biases, the choice of model can materially affect the inferred pattern of public risk preferences. In practice, one model family may overstate conservatism, whereas another may overstate willingness to accept risk. Without empirical calibration, such biases can distort inference and lead researchers to draw policy conclusions that do not accurately reflect observed human behavior.

This limitation is particularly consequential in transportation research, where risk perception is central to decision-making under uncertainty. Travel behavior frequently involves probabilistic trade-offs, including route choice under unreliable travel times, mode switching during service disruptions, and the adoption of emerging mobility technologies under safety and performance uncertainty. If the underlying risk preference parameters are systematically miscalibrated, demand forecasts, welfare evaluation, and policy design may all be biased. For example, using uncalibrated LLM-generated data to infer willingness to adopt safety-critical systems such as autonomous vehicles or low-altitude mobility services could yield either overly conservative or overly optimistic projections, depending on the model family used.

The multilingual results add a further layer of caution. In linguistically diverse settings, prompt language is not a neutral implementation choice: it can systematically perturb the behavioral calibration of the model. This is especially relevant for transportation systems serving multilingual populations, where researchers may be tempted to use native-language prompting as a straightforward way to improve realism. Our results suggest that such an assumption is unwarranted unless the model has first been validated against human benchmarks in the relevant linguistic context.

Overall, the implications of this study are methodological as much as substantive. Off-the-shelf LLMs should not be treated as direct substitutes for human respondents in risk-sensitive behavioral applications. Instead, they should be regarded as tools whose outputs require domain-specific and language-sensitive calibration. Rigorous validation against human ground truth remains a necessary prerequisite for deploying these models in transportation and other cross-cultural social science settings.

DOI Link:

https://doi.org/10.26599/COMMTR.2026.9640025

About Communications in Transportation Research

Communications in Transportation Research was launched in 2021, with academic support provided by Tsinghua University and China Intelligent Transportation Systems Association. The Editors-in-Chief are Professor Xiaobo Qu, a member of the Academia Europaea from Tsinghua University, and Professor Xiaopeng (Shaw) Li from University of Wisconsin–Madison. The journal mainly publishes high-quality, original research and review articles that are of significant importance to emerging transportation systems, aiming to serve as an international platform for showcasing and exchanging innovative achievements in transportation and related fields, fostering academic exchange and development between China and the global community.

It has been indexed in SCIE, SSCI, Ei Compendex, Scopus, CSTPCD, CSCD, OAJ, DOAJ, TRID and other databases. It was selected as Q1 Top Journal in the Engineering and Technology category of the Chinese Academy of Sciences (CAS) Journal Ranking List. In 2022, it was selected as a High-Starting-Point new journal project of the "China Science and Technology Journal Excellence Action Plan". In 2024, it was selected as the Support the Development Project of "High-Level International Scientific and Technological Journals". The same year, it was also chosen as an English Journal Tier Project of the "China Science and Technology Journal Excellence Action Plan Phase Ⅱ". In 2024, it received the first impact factor (2023 IF) of 12.5, ranking Top1 (1/58, Q1) among all journals in "TRANSPORTATION" category. In 2026, its 2025 IF was announced as 12.7, maintaining the Top1 position (1/66, Q1) in the same category.

From Volume 6 (2026), Communications in Transportation Research will be published by Tsinghua University Press on the SciOpen platform with the official journal website at https://www.sciopen.com/journal/2097-5023 . We kindly request that all new manuscript submissions be made through the journal's submission system at https://mc03.manuscriptcentral.com/commtr

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.