Weekend Learning

Weekend Learning is my LinkedIn newsletter on how LLM evaluation breaks. Every two weeks it takes one finding, recomputes every number in it from the data, and ends with an hour you can spend running the same check yourself.

I'm Hadi Mohammadi. I did my PhD in natural language processing and explainable AI at Utrecht University, and I work on AI and data science in the Netherlands.

Subscribe on LinkedIn to get each new issue on the Saturday it comes out. Below is every issue so far, newest first, with the command its hour starts from. The packages are MIT licensed and archived with a DOI, and they're all on the software page of my site.

Issues

  1. Issue 6 · 19 September 2026

    533 of 4,873 real experiments had a significant winner. 134 survived being counted.

    Of 4,873 real A/B tests, 533 produced a significant winner. Chance alone would have produced about 244 of them.

    pip install abkit
    abkit demo --out results
  2. Issue 5 · 12 September 2026

    106,134 arena votes rank 55 models. They separate one neighbouring pair out of 54.

    Closing the median gap on this board would take roughly 8.9 million votes at the current design, against the 106 thousand collected.

    pip install arenakit
    arenakit demo --out results
  3. Issue 4 · 5 September 2026

    Three models rated their own confidence. A constant beat all three.

    One of my models said 95 on 878 of its 1,000 answers and was right on 79% of them.

    pip install calikit
    curl -O https://mohammadi.cv/assets/data/blog/calibration/llama3.1-8b.jsonl
    calikit audit llama3.1-8b.jsonl --rescale 0,100
  4. Issue 3 · 29 August 2026

    HumanEval catches a real 3-point gap about one time in four

    Even at a 90% baseline, its floor is six points: the smallest true gap it catches four times in five.

    pip install abeval
    abeval power --baseline 0.75 --delta 0.03
  5. Issue 2 · 22 August 2026

    27 of the 28 emotions in GoEmotions fall below the reliability floor

    Given the per-rater reliability actually observed, you would need roughly 13 raters per item to produce aggregated labels that are 0.8 reliable. The dataset averages 3.6, which projects to 0.52.

    pip install raterkit
    raterkit report ratings.jsonl --out audit
  6. Issue 1 · 16 August 2026

    My LLM judge picked whichever answer came first, 91% of the time

    The part I didn't expect: two of the three never once called a tie between two identical answers.

    pip install judgekit
    judgekit report judge.jsonl --out audit

The Persian sister is for beginners

«یادگیری آخر هفته» is the Persian sister of this newsletter, written for people starting out. Every two weeks, on Fridays, it explains one basic idea from machine learning and language models in plain Persian, with an exercise that needs nothing installed. Many of its readers are in Iran on slow and expensive connections, so every lesson also works offline once it's downloaded.

The two take turns, so each weekend brings one of them. The Persian lessons have their own page, and the newsletter is on LinkedIn. The rest of my writing is on my site.