Cover of the thesis Let Me Explain!
Doctoral Thesis · Utrecht University

Let Me Explain!

Explainable NLP for Understanding Large Language Models

Laat me uitleggen!
Verklaarbare NLP voor het begrijpen van grote taalmodellen
(met een samenvatting in het Nederlands)

by Hadi Mohammadi

Public defense · 9 September 2026, 14:15 · Utrecht University

Read the thesis

Cover — open the flipbook
Private preview · public after the defense

Flip through the full thesis online

Password-protected for now. It opens to everyone on 9 September 2026.

Open the flipbook

Have the password? Enter it to read the thesis now.

About this thesis

This thesis presents six connected studies on explainable Natural Language Processing (NLP) for understanding large language models (LLMs). As LLMs grow more capable and widely deployed, methods to understand, verify, and evaluate their behavior across different contexts and cultures have lagged behind their raw performance.

The work addresses three dimensions of this gap. The first is methodological: explanation methods that travel well from generic feature attribution to the linguistic and cultural complexity of text. The second is empirical: the reliability of LLM-produced annotations under demographic bias and prompt variation. The third is normative: cross-cultural moral alignment, asking whether large models reflect a narrow set of values when deployed globally, and how to evaluate that systematically.

Across eight chapters and six core studies, the thesis combines a survey of the explainable NLP literature with empirical work on sexism detection, AI-text detection, annotator demographic bias, and moral alignment evaluation against the World Values Survey and PEW Global Attitudes Survey.

Defense information

University
Utrecht University, The Netherlands
Department
Methodology and Statistics
Promotors
Dr. Robert A. Bagheri
Prof. dr. Daniel L. Oberski
Co-promotor
Dr. Anastasia Giachanou
Defense
Wednesday 9 September 2026, 14:15
Layman’s talk
14:00
Venue
Academiegebouw, Domplein 29, 3512 JE Utrecht
Assessment committee
  • Prof. dr. Antal van den Bosch · Utrecht University
  • Prof. dr. Mehdi Dastani · Utrecht University
  • Prof. dr. Els Lefever · Ghent University
  • Prof. dr. Rens van de Schoot · Utrecht University
  • Prof. dr. Suzan Verberne · Leiden University
Also in the doctoral examination committee
  • Prof. dr. Peter van der Heijden · Utrecht University

Chapter overview

  • 1
    Introduction
    The opacity challenge of LLMs, the case for explainability across high-stakes domains, and an overview of the thesis contributions.
  • 2
    A Survey of Explainable NLP Applications
    A domain-specific review of XNLP across healthcare, finance, content moderation, customer relationship management, and beyond — with a critical take on evaluation, real-world applicability, and the role of human interaction.
  • 3
    A Transparent Pipeline for Sexism Detection
    An explainable detection pipeline that combines a CustomBERT ensemble with LIME and SHAP explanations to identify sexism in English and Spanish social-media posts.
  • 4
    Explainability-Based Token Replacement on LLM-Generated Text
    Explainability-guided token replacement reveals brittle features in AI-text detectors, and shows how explanation-targeted edits flip detector outputs while preserving meaning.
  • 5
    Reliability of Explainable LLM Annotations under Demographic Bias
    A study of how explainability tools and demographic personas affect LLM annotation reliability, with mixed-effects models that quantify variability sources.
  • 6
    Cultural Variations in LLM Moral Judgments
    Probing whether LLMs mirror cross-cultural moral attitudes from the World Values Survey, comparing log-probability scoring against country-level ground truth.
  • 7
    EvalMORAAL: Interpretable Moral Alignment in LLMs
    A transparent CoT framework evaluating moral alignment in 20 LLMs against the WVS and PEW surveys, with model-as-judge peer review and a persistent Western / non-Western alignment gap.
  • 8
    Discussion and Conclusion
    Cross-chapter synthesis, limitations, and directions for future work in explainable NLP and culturally aware language models.

Underlying papers

The empirical chapters are based on six peer-reviewed and preprint papers, listed in chapter order. The full publication list, including pre-PhD work, lives on the publications page.

Equal contribution.

How to cite

The full reference will be updated when the thesis is officially deposited with Utrecht University after the defense. For now, please cite as:

@phdthesis{mohammadi2026letmeexplain,
  author = {Mohammadi, Hadi},
  title  = {Let Me Explain! Explainable NLP for Understanding
            Large Language Models},
  school = {Utrecht University},
  year   = {2026},
  type   = {Doctoral thesis},
  note   = {To be defended 9 September 2026}
}