Skills
What I’ve worked with, mostly on the ML side. Filter by area or by where I used it, and open a row to see what I actually did with it.
Machine learning
Instrumented optimizer for the Muon experiments, and the inference pipeline behind the CryptoFace web demo.
Testing whether Muon works because its Newton–Schulz step approximates the polar factor, or because of the band it pushes singular values into.
Constrained fits of compute-matched quintic polynomial families with SciPy, with bounded-iteration and positivity checks.
Coefficient solver and spectrum diagnostics for the Newton–Schulz study.
A web demo for CryptoFace’s patch-based face recognition: shows patch extraction, the 256-D embedding, and identity similarity between two photos.
CryptoFace (CVPR ’25) runs face recognition under fully homomorphic encryption, so raw images and features are never exposed.
Baselines, preprocessing and evaluation splits before reaching for anything deep.
Transformers, tokenizers and the Hub for pulling and fine-tuning open models.
Weights & Biases and MLflow for runs, sweeps and comparing configs.
matplotlib for research figures; the interactive charts in the agent-loop report are hand-built SVG in React.
LLMs & agents
A two-agent loop that pulls tasks from a dependency-ordered plan, verifies its own work and writes handoff notes when it stops.
Measured what fills a 128K window turn by turn, and kept most runs under the compaction threshold.
Schema validation for tool arguments, a cap on sequential tool calls, and filler speech while a slow tool runs.
Chat Completions, Realtime, and a new Responses API agent with stateful and stateless variants.
Brought Gemini to parity with the other providers: sequential tool calls and no dropped tool output.
Throttled document uploads into the retrieval pipeline, with the concurrency limit behind a feature flag.
An identity guardrail with basic sentiment classification, configurable per agency.
Task descriptions a small model can act on in isolation, and system prompts that steer the model away from hallucinated tool calls.
Inference & serving
Moving from Ollama to llama-server took prompt-cache reuse from about 7% to 97%.
The first serving setup for the agent loop, and the baseline the later setups are measured against.
Append-only prompts so each turn reuses the cache, and a q8 KV cache to fit two 128K contexts.
Two Qwen 27B models across two 3090s; the second agent added about 20% throughput.
Race a second completion request when the first passes a timeout, without losing usage accounting or cancellations.
Device probing for a fleet manager that reports schedulable GPU and Apple Silicon memory on each machine.
Speech & voice
Added Soniox as a WebSocket provider with Deepgram as fallback, and switched to ElevenLabs Scribe mid-call for spelled-out emails.
Pregenerated, cached filler phrases in sixteen languages, and routing long numbers to a higher-quality voice.
Delayed audio commits after voice activity ends, and recovery when a turn is interrupted by silence.
Call transfers, hangups and DTMF on live phone calls.
Engineering
Most of what I write: agents, research code, scripts.
Model Hub’s web app, and this site.
Model Hub’s agent: enrolls a machine and streams its hardware inventory over signed connections.
Competitive programming (CSES problem set).
Postgres for services; BigQuery for warehouse views.
Service backends, and the Flask server behind the CryptoFace demo.
Dashboards, demos, and this site.
Local stacks and deployment.
Secrets and provider keys wired through Terraform.
Per-call baggage on every span, so a trace shows which providers and models handled each turn.
LaunchDarkly variations for rolling out providers and behaviors per agency.
Scrapers that walk listing pages and pull structured data.