Test-Time Compute - Order One More Test

A model can be made better after it is trained, just by letting it answer for longer. Almost all of that gain comes from the thing that checks the answer.

Tokenizers - Out of Sorts

A model's vocabulary is fitted by counting pairs and frozen before training starts, and it decides what the model can spell, add up and charge you.

LLM Benchmarks - The Horse and the Rider

The famous benchmarks saturated and leaked. The agentic ones replacing them score a model and the code driving it together, and cannot tell the two apart.

MCTS and AlphaZero - Four Steps in the Dark, Part 2

AlphaGo gave tree search a policy and a value network. AlphaGo Zero made the search the network's teacher, in an update that sits close to PPO's.

MCTS and AlphaZero - Four Steps in the Dark, Part 1

Why game programs search a tree, why they fill it with random games, and what plain Monte Carlo tree search does on real Go and chess positions.

Jev - The Classifier Strikes Back

Jev returns typed, calibrated decisions instead of text. Underneath, it looks like a very good zero-shot classifier, and calibration is the claim to test.

Mamba - Is Attention All You Really Need? Part 2

Selection broke the convolution that made state space models fast to train. Here is how Mamba got the speed back, and why attention stayed in the room.

Mamba - Is Attention All You Really Need? Part 1

State space models come from control theory, run like an RNN and train like a CNN. Mamba's one change lets each word decide what the model remembers.

Knowledge Distillation - The Perfumer's Apprentice

The label says rose. The teacher's nose says rose, some jasmine, and definitely not vanilla, and that last part is what the small model actually learns.

Quantization - Four Bits and a Winter Coat

Cutting every weight from sixteen bits to four is nearly free. The handful of channels that shout a hundred times louder than the rest is what costs you.

GRPO and RLVR - The Answer Key and the Rank List

RLHF trained a model to mark the homework. RLVR checks the answer instead, and GRPO sacks the critic by grading every batch on a curve.

World Models - The Film Set That Never Runs Out of Street

Learned simulators generate the street instead of building it. Genie 3, Cosmos and GAIA-3 all look right, and looking right is the cheap part.