Systems, AI tools and products I've designed and built.
Selected work
AI & Agents · Systems
swift-qwen3.8-rtx3090
A reproducible pipeline that rebuilds a 27B finetune into a fast single-RTX-3090 variant. It derives the draft vocabulary and int4 GPTQ calibration from the model's own outputs: draftable-token coverage went from 96.7% to 99.8% and decode from 94.0 to 98.4 tok/s on Swift 1.0. On Swift 1.5 the 1.0 draft list covered the new model's tokens as well as a rebuilt one. A controlled, interleaved rerun found the same raw decode speed as 1.0 and speculative throughput level to a few percent behind. Most of the acceptance gap belongs to the model, and a union of the two draft lists removed a quote-workload cliff.
A Rust secret broker that lets coding agents rotate, verify and reconcile SOPS-managed credentials without ever receiving a value. Agents send a logical id over a Unix socket; a root daemon does the work from a root-owned registry and answers with booleans.
My Linux development workstation, which also serves local LLMs to my coding agents, runs image generation and acts as a CI runner. One GPU and one pool of RAM are shared across all of it, with cgroup limits and Grafana telemetry showing what limits each workload.
A Rust CLI that turns architecture rules into checks: component size, dependencies that point towards stability, and change hotspots from Git history. It runs as a pre-commit hook or in CI, so code written by agents has to pass the same rules as mine.
Rebuilds a 27B finetune as a fast single-RTX-3090 build, with draft vocabulary and int4 heads calibrated on its own outputs, rebuilt and measured for Swift 1.5.