Let Me Explain!

Making large language models say why.

Hadi Mohammadi Academiegebouw, Utrecht

Download the slides PDF · 20 slides · 1.5 MB

Slides

  1. Title slide with the Utrecht University logo. Let Me Explain! Explainable NLP for Understanding Large Language Models. Hadi Mohammadi, Methodology, Statistics & Data Science, Utrecht University. Public defence · 9 September 2026.
    01Let Me Explain!Explainable NLP for Understanding Large Language Models
  2. Pale yellow slide with the words: Making large language models say why.
    02Making large language models say why.
  3. How fast this happened: It took two months. Bar chart, time to reach 100 million users: Spotify 11 years, Netflix 10 years, Airbnb 8 years, Twitter 5 years, Facebook 4.5 years, Dropbox 4 years, WhatsApp 3.5 years, Instagram 2.5 years, TikTok 9 months, ChatGPT 2 months, highlighted in red. Source: World of Statistics.
    03How fast this happenedIt took two monthsSource: World of Statistics
  4. Before my PhD: You will not always be told why. Three boxes joined by arrows: Your application (income, job, address), then The model, a black box with a question mark, then DENIED (no reason given). Below: You apply for a loan. The answer is no. Nobody can tell you why.
    04Before my PhDYou will not always be told whyYou apply for a loan. The answer is no. Nobody can tell you why.
  5. The problem: What is inside the box? Left, Traditional approach: a magnifying glass over a black box shows a decision tree. Income above €30,000? No leads to NO. Yes leads to Steady job?, where yes leads to YES and no leads to NO. Caption: a rule you can read. An arrow points to Neural network: the magnifying glass shows columns of numbers such as -1.01, +0.22 and -1.47. Caption: billions of numbers · no rule to read.
    05The problemWhat is inside the box?
  6. And it got worse: How large is a large language model? About 1,600 times bigger in two years. A tiny blue dot, labelled (that dot), is BERT, 2018, 110 million parameters. Next to it a large red square: GPT-3, 175 billion, two years later. A bigger dashed outline: GPT-4, larger still.
    06And it got worseHow large is a large language model?About 1,600 times bigger in two years
  7. Screenshot of a BBC News web page with the headline 'A predator in your home': Mothers say chatbots encouraged their sons to kill themselves, dated 8 November 2025. Below: Nobody could say why it answered the way it did. BBC News, 8 November 2025.
    07The problemNobody could say why it answered the way it did.BBC News, 8 November 2025
  8. Yellow slide: Why did it say that? Three years, six studies, one question.
    08Why did it say that?Three years, six studies, one question.
  9. How we look inside: Ask which words did the work. An example post: Women who study engineering should stay in the kitchen anyway. Five words are highlighted in yellow, darkest on kitchen, then Women, stay and should, lightest on anyway. An arrow points to: the darker the highlight, the more that word drove the decision.
    09How we look insideAsk which words did the work
  10. Chapter 3: Online harm is not all or nothing. What a yes/no label sees: two boxes, FINE and HARMFUL. What actually happens online: a colour bar running from pale yellow to deep red, labelled a clumsy joke, a stereotype, a put-down, open abuse. Proposition 5 · Online harm is not all or nothing, so a simple yes-or-no label misses most of the picture.
    10Chapter 3Online harm is not all or nothingProposition 5 · Online harm is not all or nothing, so a simple yes-or-no label misses most of the picture.
  11. Chapter 4: We used our own tool to fool our own detector. 1. Swap the give-away words. Written by a machine: We delve into the data and demonstrate a significant gain, with delve, demonstrate and significant highlighted. Three words swapped: We look into the data and show a clear gain. 2. Can the detector still tell? Bar chart, before and after the swap: one detector on its own, 0.81 before, 0.25 after. The ensemble (several detectors together), 0.83 before, 0.64 after. Proposition 4 · The same explanation that makes us trust a model can also show someone how to trick it.
    11Chapter 4We used our own tool to fool our own detectorProposition 4 · The same explanation that makes us trust a model can also show someone how to trick it.
  12. Chapter 5: A persona does not make a model more human. Why people disagree: donut chart, 92% what the tweet says, the other 8% is who they are. Does a persona help? A boxed quote: “You are a female individual, aged 18–22, who identifies as White, has a Bachelor’s degree, and currently resides in Europe.” It did not help. Bar chart, accuracy, English, just ask and with the persona: GPT-4o 0.76 and 0.75, smallest model 0.50 and 0.47. Four models tested, none got clearly better. Proposition 6 · Asking a model to act like a certain group does not make it more human. Often it just makes it worse.
    12Chapter 5A persona does not make a model more humanProposition 6 · Asking a model to act like a certain group does not make it more human. Often it just makes it worse.
  13. Chapter 7: They answer with a Western accent. Bar chart, how well the model matches what people really said: Western countries 0.82, Everyone else 0.61, a gap marked 21 points. 20 models · 64 countries · 23 moral topics. Proposition 9 · Large language models are not neutral judges of right and wrong. They answer with a Western accent.
    13Chapter 7They answer with a Western accentProposition 9 · Large language models are not neutral judges of right and wrong. They answer with a Western accent.
  14. The whole PhD, in one picture. A large circle labelled everything humans know. Below: Imagine everything humans know.
    14The whole PhD, in one pictureImagine everything humans know.
  15. Close-up of the top edge of the circle, with a red arrow inside it pointing up at the edge. Text: You pick one tiny spot. And you push.
    15The whole PhD, in one pictureYou pick one tiny spot. And you push.
  16. Close-up of the circle edge, now with a small bump pushed out past the original curve, which is shown dashed. A red dot at the tip of the bump is labelled This is my PhD.
    16The whole PhD, in one pictureThis is my PhD.
  17. The whole circle again. A tiny red mark on its top edge is ringed by a dashed red circle, with an arrow and the label that dent is mine. Below: Now zoom back out.
    17The whole PhD, in one pictureNow zoom back out.
  18. Why it still matters: But the dent is the point. Three cards. Builders can catch a model that is right for the wrong reason. Rule-makers can ask for the reasons, not only the results. All of us can stop treating it like magic. Below: If you take one thing home: ask it why. If it cannot tell you, do not let it decide anything that matters.
    18Why it still mattersBut the dent is the pointIf you take one thing home: ask it why. If it cannot tell you, do not let it decide anything that matters.
  19. At 14:15: In a few minutes, the conversation continues. Chair: Prof. dr. Marian Jongmans. Supervisors: Dr. Robert A. Bagheri, Prof. dr. Daniel L. Oberski, Dr. Anastasia Giachanou. Assessment committee: Prof. dr. Antal van den Bosch, Prof. dr. Mehdi Dastani, Prof. dr. Els Lefever, Prof. dr. Rens van de Schoot, Prof. dr. Suzan Verberne. Also in the doctoral examination committee: Prof. dr. Peter van der Heijden.
    19At 14:15In a few minutes, the conversation continues
  20. Closing slide on yellow with the Utrecht University logo. “These propositions were written by a human, even if my own research has made that hard to prove.” Proposition 14. Thank you.
    20Thank youProposition 14 · “These propositions were written by a human, even if my own research has made that hard to prove.”

1 / 20

The left and right arrow keys move between slides. Home and End go to the first and last slide. F turns full screen on and off.

All 20 slides

    About the talk

    Date
    Time
    14:00, before the public defence at 14:15
    Venue
    Academiegebouw, Domplein 29, Utrecht
    Slides
    20, in English · PDF (1.5 MB)