Senior Site Reliability Engineer

Roche

ESonsitePosted Jun 19, 2026
Posting intelligenceActively listed

Skills

kubernetesterraformpythonazurecicdaws

About the role

Bei Roche kannst du ganz du selbst sein und wirst für deine einzigartigen Qualitäten geschätzt. Unsere Kultur fördert persönlichen Ausdruck, offenen Dialog und echte Verbindungen. Hier wirst du für das, was du bist, wertgeschätzt, akzeptiert und respektiert. Dies schafft ein Umfeld, in dem du sowohl persönlich als auch beruflich wachsen kannst. Gemeinsam wollen wir Krankheiten vorbeugen, stoppen und heilen und sicherstellen, dass jeder Zugang zur Gesundheitsversorgung hat – heute und in Zukunft. Werde Teil von Roche, wo jede Stimme zählt.

Die Position

The Position

We are building a global Site Reliability Engineering (SRE) team to support critical commercial and internal platforms and applications. As an SRE, you will help design, build, and scale reliable distributed systems that power healthcare innovation worldwide.

This role is focused on reliability, scalability, automation and operational excellence. You will influence system design, define reliability standards and reduce operational toil through engineering solutions.

This role includes participation in a structured on-call rotation.

Who We Are

At Roche, we are passionate about transforming patients’ lives, and we are bold in both decision and action - we believe that good business means a better world. That is why we come to work every single day. We commit ourselves to scientific rigor, unassailable ethics and access to medical innovations for all. We do this today to build a better tomorrow.

Roche is strongly committed to a diverse and inclusive workplace. We strive to build teams that represent a range of backgrounds, perspectives and skills. Embracing diversity enables us to create a great place to work and to innovate for patients.

Step into the Future of IT with Roche!

As a seasoned Site Reliability Engineer (SRE) at Roche, you will leverage your deep software engineering expertise to propel our products to new heights of robustness, scalability and reliability. This isn't just a role—it's an invitation to shape the backbone of technological innovations forward.

Your Mission

Design and maintain cutting-edge tools, scripts and frameworks that automate repetitive tasks, streamline software deployment and manage expansive systems with unparalleled efficiency. Partner closely with forward-thinking development teams to architect and implement high-performance solutions that elevate system efficiency, optimize resource utilization and enhance deployment processes for superior uptime and user satisfaction.

Your Impact

Lead the charge in incident management and response. Detect system anomalies, troubleshoot swiftly and conduct thorough root cause analyses to prevent recurring issues.

Champion continuous improvement by refining monitoring and alerting mechanisms, conducting insightful post-incident reviews and embedding best practices in software lifecycle management. Your strategic foresight and meticulous planning will ensure our systems are not only reliable but also superlatively performant.

By joining our elite team, you will play a pivotal role in delivering seamless experiences to our end-users, exceeding business and customer demands, and solidifying Roche's reputation as a leader in IT innovation.

Your Core Responsibilities

Reliability Engineering& Architecture

Define and implement SLIs, SLOs, and error budgets with product and engineering teams

Conduct reliability reviews for new and existing services

Design scalable, fault-tolerant architectures in AWS and Azure environments

Lead capacity planning, performance and cost optimization initiatives

Improve system resilience through automation and self-healing patterns

Drive organizational observability maturity (metrics, logs, traces, alert quality)

Incident Management& Continuous Improvement

Perform complex root cause analysis and drive rapid mitigation

Participate in blameless postmortems and follow-through

Improve MTTR, reduce incident frequency, and elevate production standards

Collaborate seamlessly with engineering teams to enable timely and effective resolutions

Handle requests and incidents, create and maintain runbooks

Participation in a structured 24*7 on-call rotation

Automation& Platform Engineering

Reduce operational toil through tooling and automation (Python or similar)

Improve CI/CD reliability and deployment safety mechanisms

Build and maintain infrastructure-as-code (Terraform or equivalent)

Enhance Kubernetes platform reliability (EKS, AKS, or similar)

Cross-Functional Leadership

Partner with business, engineering, security, and cloud teams to embed reliability early in the software development life cycle

Mentor mid-level engineers and help shape SRE best practices

Championing a culture of ownership, accountability, and continuous improvement

Who You Are:

Minimum bachelor’s degree in computer science, Engineering, or a related field, or equivalent professional experience.

Experience in either site reliability engineering, software engineering or related fields with production on-call experience.

Solid experience with AWS and/or Azure, including setting up, monitoring, and maintaining cloud resources (incl. Kubernetes, EKS, AKS, GKE, etc knowledge).

Proficiency with observability tools

Hands-on experience with incident management tools

Proficiency in scripting languages for automation purposes

Demonstrated proficiency in troubleshooting, especially in cloud and distributed system environments

Excellent communication, teamwork and documentation skills, with a proactive and self-motivated approach to improving system reliability and operational efficiencies.

We value and encourage candidates from diverse backgrounds and experiences, believing that diverse perspectives drive innovation and success.

Excelling in both spoken and written English communication.

Where pay transparency applies, details are provided based on the primary posting location. For this role, the primary location is Sant Cugat del Vallès. If you are interested in additional locations where the role may be available, we will provide the relevant compensation details later in the hiring process.

Wer wir sind

Eine gesündere Zukunft treibt uns zur Innovation an. Mehr als 100.000 Mitarbeiter weltweit arbeiten gemeinsam daran, wissenschaftliche Fortschritte zu erzielen und sicherzustellen, dass jeder Zugang zur Gesundheitsversorgung hat – heute und für zukünftige Generationen. Durch unser Engagement werden über 26 Millionen Menschen mit unseren Medikamenten behandelt und mehr als 30 Milliarden Tests mit unseren Diagnostik-Produkten durchgeführt. Wir ermutigen uns gegenseitig, neue Möglichkeiten zu erkunden, Kreativität zu fördern und hohe Ziele zu setzen, um lebensverändernde Gesundheitslösungen zu liefern.

Gemeinsam können wir eine gesündere Zukunft gestalten.

Roche ist ein Arbeitgeber, der die Chancengleichheit fördert.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for DevOps / SRE roles in Spain varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our DevOps / SRE hub for Spain medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.