Test-Time Compute - Order One More Test
A model can be made better after it is trained, just by letting it answer for longer. Almost all of that gain comes from the thing that checks the answer.
Tokenizers - Out of Sorts
A model's vocabulary is fitted by counting pairs and frozen before training starts, and it decides what the model can spell, add up and charge you.
LLM Benchmarks - The Horse and the Rider
The famous benchmarks saturated and leaked. The agentic ones replacing them score a model and the code driving it together, and cannot tell the two apart.
MCTS and AlphaZero - Four Steps in the Dark, Part 2
AlphaGo gave tree search a policy and a value network. AlphaGo Zero made the search the network's teacher, in an update that sits close to PPO's.
MCTS and AlphaZero - Four Steps in the Dark, Part 1
Why game programs search a tree, why they fill it with random games, and what plain Monte Carlo tree search does on real Go and chess positions.
Jev - The Classifier Strikes Back
Jev returns typed, calibrated decisions instead of text. Underneath, it looks like a very good zero-shot classifier, and calibration is the claim to test.
Mamba - Is Attention All You Really Need? Part 2
Selection broke the convolution that made state space models fast to train. Here is how Mamba got the speed back, and why attention stayed in the room.
Mamba - Is Attention All You Really Need? Part 1
State space models come from control theory, run like an RNN and train like a CNN. Mamba's one change lets each word decide what the model remembers.
Knowledge Distillation - The Perfumer's Apprentice
The label says rose. The teacher's nose says rose, some jasmine, and definitely not vanilla, and that last part is what the small model actually learns.
Quantization - Four Bits and a Winter Coat
Cutting every weight from sixteen bits to four is nearly free. The handful of channels that shout a hundred times louder than the rest is what costs you.
GRPO and RLVR - The Answer Key and the Rank List
RLHF trained a model to mark the homework. RLVR checks the answer instead, and GRPO sacks the critic by grading every batch on a curve.
World Models - The Film Set That Never Runs Out of Street
Learned simulators generate the street instead of building it. Genie 3, Cosmos and GAIA-3 all look right, and looking right is the cheap part.