Back to projects
project44 · Platform engineering

Internal Developer Platform @ project44

Golden-path tooling that lets product teams provision infrastructure and ship services without touching raw infra — across 50+ applications and 12 EKS clusters.

AWS EKSArgo CDTerraform HelmGitHub ActionsPrometheus GrafanaClaude Code

The problem: infrastructure as a bottleneck

When every team needs the platform team to hand-provision infrastructure and wire up deploys, the platform team becomes the bottleneck and every team reinvents the same wheel slightly differently. The fix isn't more tickets — it's paved roads: standardized, reusable building blocks teams can self-serve.

The golden path

I built the reusable layer that sits between raw AWS/Kubernetes and the product teams:

Product teams — a PR, not a ticket Golden-path building blocks Terraform modules Helm charts Argo CD ApplicationSets Actions workflows 12 EKS clusters · multiple AWS accounts · 50+ apps Prometheus + Grafana · SLO dashboards · alerts as code

Observability standards

Reliability doesn't come from hope — it comes from seeing problems before users do. I built the Prometheus and Grafana layer as dashboards and alerts as code, so every new service inherits a real observability baseline instead of depending on whoever remembers to build one.

Platform networking: ingress-nginx to Gateway API

Ingress ran out of room. Every real feature — rewrites, timeouts, canaries — lived in vendor annotations, and routing config shared one object across platform and application concerns. Migrating to Kubernetes Gateway API backed by AWS Load Balancer Controller fixed the ownership split.

AI-assisted operations

I work AI-first, on a strict plan then generate then review-every-diff discipline. Reusable agent workflows built with Claude Code, Codex, and MCP integrations take the toil; the judgment stays human.

RESULT

Incident investigation time down ~30% — without letting an agent make an unreviewed change to IAM, state, or anything on a destructive path.

Outcomes

Where it landed

50+
apps on the platform
12
EKS clusters
−40%
platform support requests
<30min
environment setup
−30%
incident investigation