Latest writing
All posts →Quantization - Four Bits and a Winter Coat
Cutting every weight from sixteen bits to four is nearly free. The handful of channels that shout a hundred times louder than the rest is what costs you.
GRPO and RLVR - The Answer Key and the Rank List
RLHF trained a model to mark the homework. RLVR checks the answer instead, and GRPO sacks the critic by grading every batch on a curve.
World Models - The Film Set That Never Runs Out of Street
Learned simulators generate the street instead of building it. Genie 3, Cosmos and GAIA-3 all look right, and looking right is the cheap part.
Speculative Decoding - Is Attention All You Really Need?
One token costs a full read of every weight in the model. Speculative decoding guesses eight, checks them in a single read, and changes nothing.
Selected projects
All projects →Legal Research Assistant
On the cover: Legal Research Assistant Platform This project was delivered for MiAI.law, an Australian legal technology firm, to build a modern AI-driven system f...
AI-Driven Quantitative Risk Analysis
On the cover: Digitized Process Flow Diagram (PFD) with highlighted Elementary Process Sections (EPS) and graph This work was conducted at Saudi Aramco. The proje...