Summer school · proposed for summer 2027

Opening the Black Box

A Hands-On Summer School in Explainable AI for Large Language Models

Five days on interpreting, evaluating, and trusting language models. Mornings are lectures; afternoons put the instruments in your hands: feature attribution, attention analysis, probing, chain-of-thought faithfulness tests, and mechanistic interpretability, all on runnable labs that fit a free Colab GPU. By Friday, teams audit a model of their own choosing and present what they found.

5 days · 10 lectures · 13 labs · Utrecht, the Netherlands (host in discussion) · 25 to 40 participants

13 / 13

Labs that actually run

Every afternoon runs on published, version-pinned notebooks that execute end to end on a free-tier Colab T4 in twenty minutes or less, re-verified weekly by CI.

€0

Compute for participants

A laptop, a browser, and a Google account are the whole setup. No installs, no cluster, no licenses.

1 : 15

Hands-on support

One teaching assistant per fifteen participants in every afternoon session, plus a daily debrief with the instructor.


The five days

Day 1

Why models need explaining

Lectures: shortcut learning and the Clever Hans problem; a working map of the explanation method families.

Lab: train two 95%-accurate classifiers and watch one collapse under distribution shift; explain one prediction five different ways and see how little they agree.

Day 2

The instruments

Lectures: feature attribution done right (LIME, SHAP, Integrated Gradients); attention analysis and its limits.

Lab: both instrument families on one shared specimen classifier, ending with the measurement of how far attribution and attention rankings diverge.

Day 3

Build it, break it

Lectures: probing, counterfactuals, and concept directions; explanation-guided attacks.

Lab: the red-team challenge: teams build a transparent detector, then attack it through its own explanations. Leaderboard metric: fewest edits to flip a verdict.

Day 4

Trust and evaluation

Lectures: evaluating explanations (random floors, deletion tests, judge audits); whether chain-of-thought reasoning is faithful.

Lab: a faithfulness harness with bootstrap confidence intervals, and a test bench that catches models silently adopting planted hints.

Day 5

Inside the model, into the world

Lectures: mechanistic interpretability (patching, induction heads, sparse autoencoders, steering); explainability in production and the EU AI Act.

Capstone: teams audit a model or prompt of their choice with the week’s instruments and present their findings.

Format

Daily rhythm

  • 09:00 to 10:30 · lecture I
  • 11:00 to 12:30 · lecture II
  • 13:30 to 15:30 · guided lab
  • 16:00 to 17:00 · exercises or challenge
  • 17:00 to 17:15 · debrief

Who it is for

Graduate students, PhD candidates, and practitioners who use large language models and need to interpret or evaluate them. Working Python and basic machine learning assumed; no interpretability background required. A five-day school maps to 1.5 to 2 EC with the capstone as assessment.

Instructor

Hadi Mohammadi researches explainable NLP at Utrecht University and is a Senior AI & Data Science Expert at AcademicTransfer. The school is built on his book Hands-On Explainable AI: Interpreting, Evaluating, and Trusting Large Language Models; the afternoon labs are its citable companion notebooks (DOI 10.5281/zenodo.21804605). A free, self-paced edition of the course and a two-minute playground are both public.


Course materials

The full package (syllabus, ten lecture decks, lab run sheets, the challenge and capstone briefs, and organizer documents) is available to participants and reviewers. Enter the access code to download it.

AES-256-GCM · decrypted only in your browser · nothing is sent anywhere

Download the materials (zip)

If the download does not start by itself, use the button above. The archive contains a participant folder and an organizer folder.

© 2026 Hadi Mohammadi · all materials all rights reserved · private page, please don’t share the link