case studies
The two pieces of AI work I can show completely - built end to end, with the reasoning, the mistakes, and the numbers included.
Eval case study · built with n8n
The eval said 0.96. The product was dropping every number in the story.
I built a golden-set, LLM-as-judge evaluation layer for a personal news-digest pipeline - then validated the judge blind against my own labels. It caught a truncation bug that had been quietly cutting real accuracy from 0.96 to roughly 0.17, and measured a ±0.04 noise floor that most "prompt improvements" never clear.
Read the case study →iOS app · built with Claude Code
56 commits, zero Swift experience - the gym app I actually use
I'm not a programmer. I wrote the spec, made the product and data-model decisions, and reviewed every change - Claude Code wrote and committed the Swift. This is what that collaboration looked like in practice, including the design pass we redid three times and the feature I deliberately didn't build.
Read the case study →more AI work
Work I've delivered through my day-to-day role - summarized here where client confidentiality applies, ordered by how hands-on I was.