Latest writing
All posts →Test-Time Compute - Order One More Test
A model can be made better after it is trained, just by letting it answer for longer. Almost all of that gain comes from the thing that checks the answer.
Tokenizers - Out of Sorts
A model's vocabulary is fitted by counting pairs and frozen before training starts, and it decides what the model can spell, add up and charge you.
LLM Benchmarks - The Horse and the Rider
The famous benchmarks saturated and leaked. The agentic ones replacing them score a model and the code driving it together, and cannot tell the two apart.
MCTS and AlphaZero - Four Steps in the Dark, Part 2
AlphaGo gave tree search a policy and a value network. AlphaGo Zero made the search the network's teacher, in an update that sits close to PPO's.
Selected projects
All projects →Legal Research Assistant
On the cover: Legal Research Assistant Platform This project was delivered for MiAI.law, an Australian legal technology firm, to build a modern AI-driven system f...
AI-Driven Quantitative Risk Analysis
On the cover: Digitized Process Flow Diagram (PFD) with highlighted Elementary Process Sections (EPS) and graph This work was conducted at Saudi Aramco. The proje...