Senior Software Engineer · Zapier
Senior software engineer building applied AI systems, agent infrastructure, and developer tools.
For more than a decade I’ve built production systems — and, increasingly, the infrastructure and tooling other engineers rely on. My focus now is making LLM-powered products dependable: agent infrastructure, evaluation, retrieval, and long-running execution.
- 10+ years shipping production software
- Senior Software Engineer at Zapier
- TypeScript + Python
- AI agents, evaluation, retrieval & developer tooling
Selected work
All work →AI Workflow Kit
Portable, markdown-based AI skills that automate the non-coding parts of an engineer’s day — morning kickoffs, standups, meeting prep, reviews — across Claude Code, Cursor, and Codex.
Long-Running Coding Agents: A 6,000-Line Migration
Completed a 6,000-line Python-to-TypeScript migration with an AI coding agent by treating the model as a stateless executor and myself as the planner — using an external progress file to hold state across sessions.
Adaptive Prompt Selection & Evaluation
A design for moving beyond static prompts: retrieve candidate prompts by context, select with a multi-armed bandit, handle cold start, and close the loop with multidimensional evaluation and reward design.
MCP & Personal Workflow Infrastructure
A set of Model Context Protocol servers and automations that give AI assistants safe, scoped access to real tools — calendar, GitLab, meeting notes, Google Workspace, Notion — with clear boundaries on what should and shouldn’t be automated.
Selected writing
All writing →
Ground-Truthing Your LLM Judge: Does Your Eval Actually Track Reality?
LLM-as-judge is the default eval now — but a judge is just another unreliable model. Here is how to check whether its verdicts correlate with real outcomes, and catch the leniency bias most judges have.
I Built 50 AI Skills for the Parts of My Day I Never Got Around To
The biggest productivity gap isn't code — it's context gathering, note-taking, and status updates. I built 50 AI workflow skills to make those 15-minute tasks actually happen every day.
Reward Engineering and Evaluation: Making Your System Learn What Matters
This is the third post in our series on probabilistic prompt pipelines. In our first post, we explored why static prompts become bottlenecks. The second…
Building the Selector: Retrieval, Bandits, and Cold Start Solutions
This is the second post in our series on probabilistic prompt pipelines. In the first post, we explored why static prompts become bottlenecks and saw a…
Why Static Prompts Fail and How Probabilistic Selection Solves Real Problems
This is the first post in a four-part series on building production-ready probabilistic prompt pipelines. By the end of this post, you'll understand why…
Claude Code: The Hidden Testing Ground for AI Agents
I've been using Claude Code daily since it launched for general use. After months of working with it, I've developed a theory about what's actually…
What I work on
Applied AI
- Agentic workflows
- Tool use & orchestration
- Retrieval & RAG
- Evaluation & feedback loops
- Model integration & routing
- Human-in-the-loop systems
Product engineering
- Frontend architecture
- Backend services
- API design
- Data-rich interfaces
- Experimentation
- Production observability
Let’s compare notes
Working on reliable AI products, agent infrastructure, or developer tools? I’m always interested in comparing notes with serious builders.