Excloud
platform engineer
jun 2023 to aug 2026
remote
i owned systems end to end: VM lifecycle and host setup, SPDK/NVMe-oF storage, Kubernetes bootstrap and CSI, IAM, DNS, managed databases, and the APIs and tools used to run them.
the part i care about is failure behavior. most of the work was making partial transitions recover cleanly: reconciliation, idempotent attach/detach/resize paths, cleanup, retries, and the protocol bugs that only show up in production.
- compute / storage
- built the Go compute control plane over QEMU/KVM and the block storage engine on SPDK/NVMe-oF, including recovery for interrupted provisioning, attach/detach, resize, and cleanup.
- clusters / network
- shipped managed Kubernetes end to end: bootstrap, Cilium, CSI, Karpenter, and OIDC/JWKS, plus IAM, authoritative DNS, ARP/NDP, and the network plumbing underneath.
- api / tooling
- built the shared Go HTTP/OpenAPI layer and API-to-client pipeline feeding the SDK, 20-group CLI, and Terraform provider across a platform of 20+ internal services.
4,000+ accounts380 active VMs140 managed Postgres clusters20+ internal services