Workshop Paper · Published

Do Large Language Models Understand Morality Across Cultures?

Hadi Mohammadi, Yasmeen F. S. S. Meijer, Efthymia Papadopoulou, Robert A. Bagheri

Utrecht University, The Netherlands

Proceedings of the 2nd LUHME Workshop, ECAI 2025, pp. 30–39

Abstract

Recent advancements in large language models (LLMs) have established them as powerful tools across numerous domains. However, persistent concerns about embedded biases, such as gender, racial, and cultural biases arising from their training data, raise significant questions about the ethical use and societal consequences of these technologies.

This study investigates the extent to which LLMs capture cross-cultural differences and similarities in moral perspectives. Specifically, we examine whether LLM outputs align with patterns observed in international survey data on moral attitudes. To this end, we employ three complementary methods: (1) comparing variances in moral scores produced by models versus those reported in surveys, (2) conducting cluster alignment analyses to assess correspondence between country groupings derived from LLM outputs and survey data, and (3) directly probing models with comparative prompts using systematically chosen token pairs.

Our results reveal that current LLMs often fail to reproduce the full spectrum of cross-cultural moral variation, tending to compress differences and exhibit low alignment with empirical survey patterns. These findings highlight a pressing need for more robust approaches to mitigate biases and improve cultural representativeness in LLMs. We conclude by discussing the implications for the responsible development and global deployment of LLMs, emphasizing fairness and ethical alignment.

Key Contributions

Cross-Cultural Moral Evaluation

A systematic investigation of the extent to which LLMs capture cross-cultural differences and similarities in moral perspectives, tested against international survey data on moral attitudes.

Three Complementary Methods

Comparison of variances in moral scores produced by models versus surveys, cluster alignment analyses of country groupings, and direct probing with comparative prompts using systematically chosen token pairs.

Evidence of Compressed Variation

Results showing that current LLMs often fail to reproduce the full spectrum of cross-cultural moral variation, tending to compress differences and exhibit low alignment with empirical survey patterns.

Implications for Responsible LLMs

A discussion of the pressing need for more robust bias mitigation and improved cultural representativeness, with implications for the responsible development and global deployment of LLMs emphasizing fairness and ethical alignment.

Citation

@inproceedings{mohammadi2025morality,
  title={Do Large Language Models Understand Morality Across Cultures?},
  author={Mohammadi, Hadi and Meijer, Yasmeen F. S. S. and Papadopoulou, Efthymia and Bagheri, Robert A.},
  booktitle={Proceedings of the 2nd LUHME Workshop},
  year={2025},
  pages={30--39},
  url={https://aclanthology.org/2025.luhme-1.3/}
}