Weekend Learning
Weekend Learning is my LinkedIn newsletter on how LLM evaluation breaks. Every two weeks it takes one finding, recomputes every number in it from the data, and ends with an hour you can spend running the same check yourself.
I'm Hadi Mohammadi. I did my PhD in natural language processing and explainable AI at Utrecht University, and I work on AI and data science in the Netherlands.
Subscribe on LinkedIn to get each new issue on the Saturday it comes out. Below is every issue so far, newest first, with the command its hour starts from. The packages are MIT licensed and archived with a DOI, and they're all on the software page of my site.
Issues
-
Issue 6 · 19 September 2026
533 of 4,873 real experiments had a significant winner. 134 survived being counted.
Of 4,873 real A/B tests, 533 produced a significant winner. Chance alone would have produced about 244 of them.
pip install abkit abkit demo --out results -
Issue 5 · 12 September 2026
106,134 arena votes rank 55 models. They separate one neighbouring pair out of 54.
Closing the median gap on this board would take roughly 8.9 million votes at the current design, against the 106 thousand collected.
pip install arenakit arenakit demo --out results -
Issue 4 · 5 September 2026
Three models rated their own confidence. A constant beat all three.
One of my models said 95 on 878 of its 1,000 answers and was right on 79% of them.
pip install calikit curl -O https://mohammadi.cv/assets/data/blog/calibration/llama3.1-8b.jsonl calikit audit llama3.1-8b.jsonl --rescale 0,100 -
Issue 3 · 29 August 2026
HumanEval catches a real 3-point gap about one time in four
Even at a 90% baseline, its floor is six points: the smallest true gap it catches four times in five.
pip install abeval abeval power --baseline 0.75 --delta 0.03 -
Issue 2 · 22 August 2026
27 of the 28 emotions in GoEmotions fall below the reliability floor
Given the per-rater reliability actually observed, you would need roughly 13 raters per item to produce aggregated labels that are 0.8 reliable. The dataset averages 3.6, which projects to 0.52.
pip install raterkit raterkit report ratings.jsonl --out audit -
Issue 1 · 16 August 2026
My LLM judge picked whichever answer came first, 91% of the time
The part I didn't expect: two of the three never once called a tie between two identical answers.
pip install judgekit judgekit report judge.jsonl --out audit
The Persian sister is for beginners
«یادگیری آخر هفته» is the Persian sister of this newsletter, written for people starting out. Every two weeks, on Fridays, it explains one basic idea from machine learning and language models in plain Persian, with an exercise that needs nothing installed. Many of its readers are in Iran on slow and expensive connections, so every lesson also works offline once it's downloaded.
The two take turns, so each weekend brings one of them. The Persian lessons have their own page, and the newsletter is on LinkedIn. The rest of my writing is on my site.