We are seeking a
Senior Site Reliability Engineer to strengthen critical infrastructure reliability and accelerate delivery through robust DevOps practices. You will drive improvements across CI/CD, cloud platforms, and operational readiness while solving high-impact production issues.
Responsibilities
- Design reliability strategies for critical infrastructure and services
- Build and maintain CI/CD pipelines and release workflows using GitLab
- Automate infrastructure operations and tooling with Python
- Operate cloud environments across AWS and Azure to meet availability goals
- Harden and standardize infrastructure practices across networking, security, IAM, and compute
- Coordinate incident response during on-call shifts and restore service quickly
- Implement Kubernetes-based deployment and operational patterns to improve stability
- Analyze reliability signals and root causes to prevent repeat incidents
- Improve DevOps processes and engineering capabilities to enable faster change
- Partner with stakeholders to prioritize resilience work over short-term fixes
Requirements
- 3+ years of site reliability engineering experience in cloud environments
- Proven leadership ability to influence reliability practices across teams
- Enterprise-scale release management experience supporting frequent deployments
- Strong cloud platform knowledge across Amazon Web Services and Microsoft Azure
- Advanced Python programming skills for automation and tooling
- Solid Kubernetes skills using clusters as a developer
- Deep CI/CD and source control knowledge with GitLab or similar DevSecOps platforms
- Strong infrastructure fundamentals across networking, compute, security, IAM, and configuration automation
- Strong analytical skills for complex problem solving under pressure
- Upper-Intermediate English proficiency (B2)
- Reliable on-call readiness to assess and resolve business-critical issues
Nice to have
- Amazon Web Services expertise, including design patterns for resilient systems
- Microsoft Azure expertise, including governance and operational best practices
- AI Architecture experience applied to platform reliability and automation
- AI Solution Engineering experience for production-grade AI-enabled operations
- Gen AI Solutions Development experience focused on operational use cases
We offer
- International projects with top brands
- Work with global teams of highly skilled, diverse peers
- Healthcare benefits
- Employee financial programs
- Paid time off and sick leave
- Upskilling, reskilling and certification courses
- Unlimited access to the LinkedIn Learning library and 22,000+ courses
- Global career opportunities
- Volunteer and community involvement opportunities
- EPAM Employee Groups
- Award-winning culture recognized by Glassdoor, Newsweek and LinkedIn
EPAM is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, sexual orientation, gender identity or expression, disability, protected veteran status, or any other characteristic protected by applicable law.