AI's Social Norms and Their Societal Impact

PNAS Nexus

As more and more people turn to LLMs for advice on interpersonal matters, the ethical foundations behind the tools' replies increasingly have the power to shape individual lives and, at scale, the social fabric itself. Alexandre S. Pires and colleagues explored the social norms of 21 LLMs by asking them to judge fictional people as "good" or "bad" after learning how the people made a range of interpersonal decisions, such as whether to spend time and energy helping a person with a problem or whether to give food or money to another person. While all models broadly approved of helping and sharing with people who had already shown themselves to be good, the models' judgements of dealings with ill-reputed individuals differed. Most—but not all—LLMs tested assigned a good reputation to those who cooperated with bad recipients. Gemma 2 27B IT and Llama 3.1 8B tended to endorse shunning bad individuals and therefore judged that cooperating with them was bad. When it came to deciding not to help or share with bad people, LLMs disagreed. GPT-4o tended to penalize people for not cooperating with bad people. Llama 3.3 70B did not, typically arguing that the bad people should be punished for their prior lack of cooperation, so refusing to help them is fine. Gemini 1.5-Pro and Grok 2 were inconsistent in their judgements. Most LLM model families seem to be evolving toward a norm called Simple Standing, in which cooperating is always good and defecting against bad individuals is also deemed good.

The authors modeled the consequences of universalizing the various philosophies underpinning the LLMs judgements. If everyone operated under the rules of Simple Standing, cooperation would ensue, but not at the high levels that would be reached under a sterner framework, in which you should always fail to help bad people or be judged as bad yourself.

LLM-based assessments also depended on the gender and perceived cultural background of the recipients as well as the overall context of the fictional situation. According to the authors, prompting interventions, such as instructing the model to adopt norms that promote overall cooperation, have a limited and inconsistent effect across LLMs.

/Public Release. This material from the originating organization/author(s) might be of the point-in-time nature, and edited for clarity, style and length. Mirage.News does not take institutional positions or sides, and all views, positions, and conclusions expressed herein are solely those of the author(s).View in full here.