Latest writing
All posts →LLM Benchmarks - The Horse and the Rider
The famous benchmarks saturated and leaked. The agentic ones replacing them score a model and the code driving it together, and cannot tell the two apart.
MCTS and AlphaZero - Four Steps in the Dark, Part 2
AlphaGo gave tree search a policy and a value network. AlphaGo Zero made the search the network's teacher, in an update that sits close to PPO's.
MCTS and AlphaZero - Four Steps in the Dark, Part 1
Why game programs search a tree, why they fill it with random games, and what plain Monte Carlo tree search does on real Go and chess positions.
Jev - The Classifier Strikes Back
Jev returns typed, calibrated decisions instead of text. Underneath, it looks like a very good zero-shot classifier, and calibration is the claim to test.
Selected projects
All projects →Legal Research Assistant
On the cover: Legal Research Assistant Platform This project was delivered for MiAI.law, an Australian legal technology firm, to build a modern AI-driven system f...
AI-Driven Quantitative Risk Analysis
On the cover: Digitized Process Flow Diagram (PFD) with highlighted Elementary Process Sections (EPS) and graph This work was conducted at Saudi Aramco. The proje...