EvalMORAAL: Interpretable Chain-of-Thought and LLM-as-Judge Evaluation for Moral Alignment in Large Language Models
IOPS Winter Conference 2025, Utrecht University — December 2025. Best Oral Presentation Award.
I work on explainable NLP, cultural fairness in large language models, and human-AI collaboration. PhD from Utrecht University.
I build NLP models whose decisions domain experts can inspect and question. This covers post-hoc interpretability methods, token-level explanations, and transparent pipelines for sensitive classification tasks in the social sciences. All of this work follows open-science practice: code, data, and paper materials are released so results can be reproduced and extended.
Large language models encode moral judgments, but not everyone's. I study how LLMs represent cross-cultural variation in moral judgments and societal norms, and how to evaluate their alignment across cultures — including EvalMORAAL, an interpretable chain-of-thought and LLM-as-judge framework for moral alignment.
When can we trust an LLM's annotations and explanations, and when do we still need a human? I compare model outputs against human rationalizations, measure demographic bias in annotation, and study how explanations should be written to match human expectations. A related line uses reinforcement learning to extract decision rules from human evaluations.
IOPS Winter Conference 2025, Utrecht University — December 2025. Best Oral Presentation Award.
6th NLTP Content Meeting Session, Utrecht — February 2025.
BNAIC/BeNeLearn 2024, Utrecht — November 2024. Presented as Communications Chair of the conference.
39th IOPS Winter Conference, Amsterdam — December 2023.
AI-lab Users Meeting (ASReview), Utrecht University — November 2023.
33rd Meeting of Computational Linguistics in the Netherlands (CLIN33), Antwerp — October 2023.
Human Data Science Content-Oriented Meeting, Utrecht University — October 2023.
33rd Meeting of Computational Linguistics in the Netherlands (CLIN33), Antwerp — September 2023.
Skills Lab, M&S Research Hay Day 2023, Utrecht University — June 2023. With Dr. Huyen Nguyen and Daniel Anadria.
I am open to research collaborations in explainable NLP, cultural fairness in LLMs, annotation reliability, and applied machine learning for the social sciences. If our interests overlap, email me at hadi.mohammadi@outlook.com.