Site Reliability Manager - Environment Strategy

EPAM Systems

London, UKhybridPosted Jul 10, 2026
Posting intelligenceActively listedReposted 18×, possible evergreen/ghost posting

Skills

cloudformationprometheusterraformgrafanacicdaws

About the role

We're looking for a Site Reliability Manager to join our team in London, United Kingdom in a hybrid working mode. In this role, you will lead a team focused on environment strategy, automation, patch governance and operational reliability for AWS-based platforms. Your responsibilities include setting roadmaps, driving best practices and delivering consistency across complex environments. You will manage people, budgets and schedules while ensuring strong technical standards, compliance and resilience across all production and non-production systems.

Responsibilities

Define and own the vision and roadmap for site reliability and environment strategy

Lead, mentor and develop a team of DevOps and environment engineers

Set and enforce standards for environment provisioning, lifecycle management and patch governance

Drive adoption of Infrastructure as Code and automation-first practices across environments

Oversee monitoring, alerting and operational readiness to meet availability and performance objectives

Partner with security, engineering and operations teams to reduce risk and ensure compliance

Build continuous improvement processes and establish key performance metrics for reliability

Requirements

8+ years of experience in site reliability, DevOps or platform engineering roles

Strong knowledge of AWS platforms and cloud infrastructure principles

Proven ability to manage teams and operational priorities in complex environments

Experience implementing automation using Terraform or CloudFormation

Experience with CI/CD pipelines and operational monitoring with tools such as CloudWatch, Prometheus and Grafana

Understanding of OS patching for Windows Server and RHEL as part of governance practices

Background in risk management, compliance frameworks and incident response processes

Excellent leadership, communication and stakeholder management skills

We offer

EPAM Employee Stock Purchase Plan (ESPP)

Protection benefits including life assurance, income protection and critical illness cover

Private medical insurance and dental care

Employee Assistance Program

Competitive group pension plan

Cyclescheme, Techscheme and season ticket loans

Various perks such as free Wednesday lunch in-office, on-site massages and regular social events

Learning and development opportunities including in-house training and coaching, professional certifications, and courses

If otherwise eligible, participation in the discretionary annual bonus program

If otherwise eligible and hired into a qualifying level, participation in the discretionary Long-Term Incentive (LTI) Program

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for DevOps / SRE roles in United Kingdom varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our DevOps / SRE hub for United Kingdom medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.