Lab
Experiments in progress
Benchmarks, prototypes, and investigations, published before they are finished. Each entry states the premise being tested and what is known so far, including the results that were not what I expected.
Active experiments
- RunningJune 2026
MLI MCP Server
Premise
Can a Model Context Protocol server give an assistant genuinely useful access to an inbox while keeping the permission surface small enough to reason about?
- Python
- Model Context Protocol
- Gmail API
- RunningMay 2026
Autonomous Content Agent
Premise
An agent that researches a topic, writes about it, and publishes across platforms without supervision. The real question is whether it can judge when something is not good enough to post.
- Python
- LLM orchestration
- Automation
- NLP
- BenchmarkApril 2026
Speech Model Quantisation Benchmark
Premise
A repeatable comparison of quantised speech recognition models on consumer hardware, measuring word error rate against latency and memory footprint.
- ONNX Runtime
- CTranslate2
- Whisper
- Python
- PrototypeMarch 2026
Multi Modal Agent Failure Handling
Premise
Built at the Google and Cerebral Valley hackathon. An agent orchestrating vision, text, and video models, with the actual objective being what happens when one of them fails.
- LangChain
- Node.js
- Vision models
Paused
On hold, and why
Work that stopped for a reason worth recording. Paused is a more honest status than complete.
Contact
If you are working on something where being wrong matters, I would like to hear about it.
I am open to consulting engagements, research collaborations, and conversations that do not have a clear outcome yet.