Insights
Notes from the work
Writing for engineers, on evaluation, speech, and language model systems. Each piece exists because something in the work turned out differently than expected.
Engineering / Evaluation / Large Language Models / Machine Learning / Product Thinking / Research / Speech AI
Articles
- 4 min read
Fluency is not accuracy
Language models are optimised to sound right, which is a different objective from being right. In clinical documentation the gap between the two is the entire safety problem.
Speech AI / Research / Evaluation
- 3 min read
Your benchmark is not your production distribution
Benchmark scores predict production behaviour poorly, and the reasons are structural rather than accidental. What to measure instead.
Machine Learning / Evaluation / Engineering
- 3 min read
The case for running models on the edge
On device inference is usually framed as a compromise on capability. In regulated settings it is closer to the opposite, and it changes what you are able to sell.
Engineering / Product Thinking / Speech AI
- 4 min read
Keep the state outside the prompt
Long running assistants degrade when conversation history becomes the memory. Treating the prompt as a rendering of explicit state fixes quality and cost at once.
Large Language Models / Engineering / Product Thinking
Contact
If you are working on something where being wrong matters, I would like to hear about it.
I am open to consulting engagements, research collaborations, and conversations that do not have a clear outcome yet.