About The Role
The role owns the infrastructure that keeps production systems running reliably at scale - CI/CD pipelines, Kubernetes clusters, cloud networking, and the observability stack that surfaces problems before customers notice them.
This is a hands-on position on a small platform team working directly with product engineers. Expect real ownership: the tooling built here is used by every engineering team, and reliability decisions made in this role directly affect uptime, deploy velocity, and infrastructure cost.
Key Responsibilities
- Build, maintain, and scale Kubernetes-based infrastructure across staging and production environments, including cluster upgrades, autoscaling policies, and node optimization
- Design and improve CI/CD pipelines using GitHub Actions or GitLab CI, with a focus on build speed, deployment safety, and automated rollback
- Own Infrastructure-as-Code using Terraform and Helm; manage state, module reuse, and drift detection across cloud accounts
- Build and maintain observability tooling with Prometheus, Grafana, and OpenTelemetry; define actionable SLOs and on-call alerting that minimizes noise
- Lead incident response and blameless postmortems; drive root-cause fixes and resilience patterns (circuit breakers, graceful degradation, multi-AZ failover)
- Optimize cloud spend on AWS or GCP - right-sizing instances, storage tiering, and reserved capacity planning - without sacrificing performance or reliability
- Partner with application teams on service design reviews, load testing, and production readiness for high-traffic launches
What We Are Looking For
- 3–7 years of experience in DevOps, SRE, or infrastructure engineering, including running production systems serving meaningful traffic
- Deep hands-on expertise with Kubernetes and containers in production - not just cluster setup, but day-2 operations: upgrades, debugging, and incident response
- Strong Infrastructure-as-Code skills with Terraform (or Pulumi); experience managing multi-environment cloud infrastructure on AWS or GCP
- Proficiency with Linux systems, networking fundamentals (DNS, TLS, load balancing, VPC design), and bash/Python scripting
- Experience building or maintaining CI/CD pipelines and GitOps workflows (ArgoCD, Flux, or similar)
- Track record of improving measurable reliability outcomes: MTTR, error budgets, on-call burden reduction, or cost optimization
- Bachelor's degree in Computer Science, Engineering, or equivalent practical experience. Bonus: experience with service meshes (Istio, Linkerd), secrets management (Vault), multi-region architectures, or SOC 2 / compliance-related infrastructure work