Principal/Senior Site Reliability Engineer

Roche

London, UKonsitePosted Jul 20, 2026
Posting intelligenceActively listedReposted 3×, possible evergreen/ghost posting

Skills

cloudformationkubernetesprometheustensorflowterraformairflowgrafanadatadogdockerpulumipythonazuresparkgooglecloudawsgoml

About the role

At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally. This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come. Join Roche, where every voice matters.

The Position

Join the Computational Sciences Center of Excellence as a Senior Site Reliability Engineer, where the platforms you build accelerate the discovery of transformative medicines. You will work alongside talented engineers in the Data& Digital Catalyst organisation to design resilient, cloud-based systems for MLOps and HPC workloads at global scale. This is a role for someone who wants their engineering craft to have real impact on science and patients.

The Opportunity:

You architect Infrastructure as Code using Terraform, Pulumi, or CloudFormation to provision and manage cloud infrastructure for MLOps and HPC workloads across global regions.

You design for resiliencebuilding disaster recovery and failover plans with auto-scaling and load balancing to keep critical systems available worldwide.

You strengthen reliability through chaos engineeringrunning experiments that validate systems and surface weaknesses before they become incidents.

You build deep observability with monitoring, logging, and alerting frameworks such as Prometheus, Grafana, Datadog, and ELK.

You provide technical leadershipto a team of engineers, fostering collaboration, innovation, and continuous improvement.

You partner across teams to align infrastructure with ML and HPC needs and to advance operational maturity through SLAs, SLOs, SLIs, and error budgets.

Who you are:

You bring deep expertise in Infrastructure as Code with proven success deploying Terraform, Pulumi, or CloudFormation in AWS, Azure, or GCP for MLOps and HPC workloads.

You understand cloud-native and on-prem architectures including autoscaling, serverless, and multi-region deployments, and you are hands-on with Docker, Kubernetes, and Kubeflow.

You are an expert in automationscripting confidently in Python, Bash, or Go, with a strong grasp of GPU-accelerated computing and HPC workload scaling.

You lead through influencecommunicating and mentoring with clarity, and solving complex problems with a methodical approach.

You hold a degree in Computer Science or a related technical fieldor bring equivalent experience in software and site reliability engineering.

Preferred:

Experience with distributed ML frameworks such as Horovod or TensorFlow Distributed.

Familiarity with data engineering pipelines such as Apache Airflow or Apache Spark.

Knowledge of chaos engineering tools and compliance frameworks such as GDPR, SOC 2, or ISO 27001.

Relocation benefits are available for this position.

Who we are

A healthier future drives us to innovate. Together, more than 100’000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact.

Let’s build a healthier future, together.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for DevOps / SRE roles in United Kingdom varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our DevOps / SRE hub for United Kingdom medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.