EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
IOPS Winter Conference 2025, Utrecht University, December 2025. Best Oral Presentation Award.
I work on explainable NLP, cultural fairness in large language models, and human-AI collaboration. PhD from Utrecht University.
I build NLP models whose decisions domain experts can inspect and question. This covers post-hoc interpretability methods, token-level explanations, and transparent pipelines for sensitive classification tasks in the social sciences. All of this work follows open-science practice: code, data, and paper materials are released so results can be reproduced and extended.
Large language models encode moral judgments, but not everyone’s. I study how LLMs represent cross-cultural variation in moral judgments and societal norms, and how to evaluate their alignment across cultures, including EvalMORAAL, an interpretable chain-of-thought and LLM-as-judge framework for moral alignment.
When can we trust an LLM’s annotations and explanations, and when do we still need a human? I compare model outputs against human rationalizations, measure demographic bias in annotation, and study how explanations should be written to match human expectations. A related line uses reinforcement learning to extract decision rules from human evaluations.
IOPS Winter Conference 2025, Utrecht University, December 2025. Best Oral Presentation Award.
6th NLTP Content Meeting Session, Utrecht, February 2025.
BNAIC/BeNeLearn 2024, Utrecht, November 2024. Presented as Communications Chair of the conference.
33rd IOPS Winter Conference, Amsterdam, December 2023.
AI-lab Users Meeting (ASReview), Utrecht University, November 2023.
33rd Meeting of Computational Linguistics in the Netherlands (CLIN33), Antwerp, September 2023.
Human Data Science Content-Oriented Meeting, Utrecht University, October 2023.
33rd Meeting of Computational Linguistics in the Netherlands (CLIN33), Antwerp, September 2023.
Skills Lab, M&S Research Hay Day 2023, Utrecht University, June 2023. With Dr. Huyen Nguyen and Daniel Anadria.
I am open to research collaborations in explainable NLP, cultural fairness in LLMs, annotation reliability, and applied machine learning for the social sciences. If our interests overlap, email me at hadi.mohammadi@outlook.com.