
Principal/Senior Site Reliability Engineer
Skills
About the role
At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally. This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come. Join Roche, where every voice matters.
The Position
Join the Computational Sciences Center of Excellence as a Senior Site Reliability Engineer, where the platforms you build accelerate the discovery of transformative medicines. You will work alongside talented engineers in the Data& Digital Catalyst organisation to design resilient, cloud-based systems for MLOps and HPC workloads at global scale. This is a role for someone who wants their engineering craft to have real impact on science and patients.
The Opportunity:
You architect Infrastructure as Code using Terraform, Pulumi, or CloudFormation to provision and manage cloud infrastructure for MLOps and HPC workloads across global regions.
You design for resiliencebuilding disaster recovery and failover plans with auto-scaling and load balancing to keep critical systems available worldwide.
You strengthen reliability through chaos engineeringrunning experiments that validate systems and surface weaknesses before they become incidents.
You build deep observability with monitoring, logging, and alerting frameworks such as Prometheus, Grafana, Datadog, and ELK.
You provide technical leadershipto a team of engineers, fostering collaboration, innovation, and continuous improvement.
You partner across teams to align infrastructure with ML and HPC needs and to advance operational maturity through SLAs, SLOs, SLIs, and error budgets.
Who you are:
You bring deep expertise in Infrastructure as Code with proven success deploying Terraform, Pulumi, or CloudFormation in AWS, Azure, or GCP for MLOps and HPC workloads.
You understand cloud-native and on-prem architectures including autoscaling, serverless, and multi-region deployments, and you are hands-on with Docker, Kubernetes, and Kubeflow.
You are an expert in automationscripting confidently in Python, Bash, or Go, with a strong grasp of GPU-accelerated computing and HPC workload scaling.
You lead through influencecommunicating and mentoring with clarity, and solving complex problems with a methodical approach.
You hold a degree in Computer Science or a related technical fieldor bring equivalent experience in software and site reliability engineering.
Preferred:
Experience with distributed ML frameworks such as Horovod or TensorFlow Distributed.
Familiarity with data engineering pipelines such as Apache Airflow or Apache Spark.
Knowledge of chaos engineering tools and compliance frameworks such as GDPR, SOC 2, or ISO 27001.
Relocation benefits are available for this position.
Who we are
A healthier future drives us to innovate. Together, more than 100’000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact.
Let’s build a healthier future, together.
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.