Selected
work.
Real products and open lab builds. The decisions, the implementation, and the measured results.

Lab build: an AI support copilot retrofitted onto a live admin app in one PR
A portfolio lab build on Tessera (a fictional SaaS): an AI support copilot retrofitted onto an existing admin app in a single PR, human-approved before any mutation, with accuracy from a real, repeatable eval.
30/30 (100%)Tool-selection accuracy (30-command eval)
CFX: a forex-signals platform with 40+ screens across two mobile apps
A complete forex-signals product: separate admin and user mobile apps, 40+ polished screens, and one full reusable design system.
40+Screens designed in Figma
Lab build: a "chat with your docs" RAG that scores 27/27 retrieval hit@1 and runs live at $0/query
A portfolio lab build: a grounded 'chat with your data' RAG you drop into an existing Next.js app: 27/27 hit@1 on a 27-question eval, a 1.77s median answer at $0/query on the free Gemini path (or Claude Haiku 4.5 at $0.0035/query), with clickable citations and an honest 'I don't know'.
27/27 hit@1Retrieval accuracy (27-question eval)
BridgeSync: an AI progress-tracking SaaS built end-to-end, billing included
My university final-year project, built solo in four months: a SaaS that turns raw GitHub activity into client-ready progress reports (hybrid rule + LLM matching of pull requests to requirements), with auth, plans, and Stripe subscription billing done properly.
1,809Automated tests passing
Lab build: an invoice extractor that flags every low-confidence field for review
A portfolio lab build on Tessera (a fictional SaaS): upload an invoice, get review-ready structured JSON with per-field confidence. Accuracy graded on the third-party SROIE benchmark.
77.8%SROIE field exact-match (30 receipts, 90 fields)Lab build: an n8n pipeline that scores and routes inbound leads automatically
A portfolio lab build on Tessera (a fictional SaaS): an importable n8n workflow that enriches, scores, logs, and routes inbound B2B leads, with cost and latency measured on real runs.
Lab build · Tessera (fictional SaaS)Lab build: a public Telegram support bot that routes 30/30 and cites its source docs
A portfolio lab build on Tessera (a fictional SaaS): a public Telegram support agent that answers from the product's help docs with the source article cited, looks up real ticket status, and escalates billing disputes to a human. Hardened for public traffic (webhook secret, idempotent updates, per-chat rate limits, a daily LLM budget) and scoring 30/30 routing and 27/27 retrieval hit@1 at $0 on the free Gemini tier.
30/30 (100%)Routing accuracy (30-intent eval)
Lab build: a browser-native voice booking agent that went 73.3% → 100% task completion
A portfolio lab build on Tessera (a fictional SaaS): a browser-native, push-to-talk voice agent that books appointments. It can only offer slots the calendar actually returned, books atomically so exactly one caller wins a contested slot, and issues real confirmation codes; a committed eval caught the model silently reformatting slot IDs and drove task completion from 73.3% to 15/15.
15/15 (100%)Task completion (15-dialogue eval)Lab build: an MCP server that gives Claude safe hands on ops data, every write a dry-run first
A portfolio lab build on Tessera (a fictional SaaS): an MCP server that gives Claude seven tools over a support-ops database (tickets, customers, invoices). Reads are instant, but every mutation defaults to a dry-run preview and executes only on an explicit confirm re-call, with human-in-the-loop enforced by the tool contract rather than model goodwill. Race-safe, zero-key quickstart, 23/23 on a deterministic tool-contract suite.
23/23Tool-contract suite (deterministic)