Latest writing
All posts →GRPO and RLVR - The Answer Key and the Rank List
RLHF trained a model to mark the homework. RLVR checks the answer instead, and GRPO sacks the critic by grading every batch on a curve.
World Models - The Film Set That Never Runs Out of Street
Learned simulators generate the street instead of building it. Genie 3, Cosmos and GAIA-3 all look right, and looking right is the cheap part.
Speculative Decoding - Is Attention All You Really Need?
One token costs a full read of every weight in the model. Speculative decoding guesses eight, checks them in a single read, and changes nothing.
Solving Chess - Paths, Nodes and the Proof Nobody Has
Musk and Chess.com were counting different things. Solving chess was never about the tree's size, but about proving which branches you can skip.
Selected projects
All projects →Legal Research Assistant
On the cover: Legal Research Assistant Platform This project was delivered for MiAI.law, an Australian legal technology firm, to build a modern AI-driven system f...
AI-Driven Quantitative Risk Analysis
On the cover: Digitized Process Flow Diagram (PFD) with highlighted Elementary Process Sections (EPS) and graph This work was conducted at Saudi Aramco. The proje...