Golden-path tooling that lets product teams provision infrastructure and ship services without touching raw infra — across 50+ applications and 5 EKS clusters.
When every team needs the platform team to hand-provision infrastructure and wire up deploys, the platform team becomes the bottleneck and every team reinvents the same wheel slightly differently. The fix isn't more tickets — it's paved roads: standardized, reusable building blocks teams can self-serve.
I built the reusable layer that sits between raw AWS/Kubernetes and the product teams:
Availability doesn't come from hope — it comes from seeing problems before users do. I run the Prometheus + Grafana layer as dashboards-as-code with SLOs and tuned alerts, holding 99.95% service availability.
I work AI-first. Claude Code runs as a daily agent in the infrastructure repos — but on a strict plan → generate → review-every-diff discipline. The judgment stays human; the toil gets automated.
Incident investigation and troubleshooting time down ~30% — without letting an agent make an unreviewed change to IAM, state, or anything destructive.