Deep expertise in SRE principles: automation-first operations, observability-driven engineering, error budget management, and resilience engineering.
Experience shaping SRE operating models, defining service-level objectives, governing automation standards, and leading platform and architecture reviews.
A track record of strengthening resilience through incident response, performance optimisation, capacity planning, and continuous improvement.
Experience building SRE capability, including coaching emerging leaders and influencing senior stakeholders across functions.
Experience designing or scaling observability, automation, self-healing or resilience platforms in a complex environment.
Credibility with engineering teams you do not manage. A large part of this role is influence rather than authority.
Client confirmation required: Confirm the observability, incident-management, service-mesh and chaos-engineering stack. Do not list competing platforms as though all are in use.
You should be comfortable with observability and cloud-native tooling such as OpenTelemetry, Prometheus, Grafana, Istio, Linkerd, and PagerDuty, and with chaos engineering tools such as Chaos Mesh or Litmus.
What may set you apart
You have built an SRE practice from a thin base rather than inheriting a mature one.
You have led incident response somewhere downtime was visible outside the company.
You combine technical depth with curiosity about the business, its products, and what drives its revenue.
A degree in computer science, engineering, information systems, or a related field is useful. Equivalent practical experience is valued just as highly.