Blog
Notes from inside the build.
What we learned shipping AI systems for real companies, including the parts that did not work.
Featured · AI Engineering
The eval is the product
Teams treat evaluation as a testing chore bolted on at the end. In an AI system it is the specification, and the only thing standing between a demo and something you can run unattended.
All posts
What retrieval gets wrong when your users write in Malay
Every retrieval benchmark you have read is English. Here is what actually breaks when the corpus is code-switched Bahasa Malaysia, and how to measure it before it embarrasses you.
Why we stopped billing hours
A timesheet rewards the wrong thing. Moving to outcome-based pricing changed what our engineers optimised for, and made the incentive honest.
The pilot that never ships
A pattern we see in almost every enterprise: an AI pilot that succeeds on its own terms and never reaches production. The cause is usually decided before the pilot starts.
The case for squads of four
Adding people to a delivery team stops helping much earlier than most organisations expect. What we optimise for instead.
Prefer this in your inbox?
One tactical breakdown every Tuesday: what we shipped, what broke, and what we would do differently.