Senior Site Reliability Engineer

DigiCert

USonsite$125k-$145k/yrPosted Jun 22, 2026
Posting intelligenceActively listed

Skills

kubernetesprometheusterraformgrafanadatadogpythonazurecicdgooglecloudawsgo

About the role

Who we are

DigiCert is a global leader in intelligent trust. We protect the digital world by ensuring the security, privacy, and authenticity of every interaction. Our AI-powered DigiCert ONE platform unifies PKI, DNS, and certificate lifecycle management, to secure infrastructure, software, devices, messages, AI content and agents. Learn why more than 100,000 organizations, including 90% of the Fortune 500, choose DigiCert to stop today's threats and prepare for a quantum-safe future at www.digicert.com

Job summary

The Site Reliability Engineer (SRE) collaborates with development teams to embed reliability, scalability, and performance best practices throughout the software development lifecycle. This role bridges software engineering and cloud operations, ensuring mission-critical systems remain highly available and resilient. By integrating reliability early, the SRE fosters a culture of shared responsibility while enabling rapid and safe feature delivery.

What you will do

Design and build fault-tolerant, high-performing systems that meet Service Level Objectives (SLOs) and Service Level Agreements (SLAs).

Implement monitoring, alerting, distributed tracing, and logging to ensure real-time system health visibility and proactive issue resolution.

Act as a first responder for production incidents, conduct blameless postmortems, and drive root cause analysis (RCA) and corrective actions.

Develop self-healing, automated deployments, and scaling solutions to minimize toil and improve system efficiency.

Improve continuous integration and deployment pipelines to enable safe, rapid, and reliable feature rollouts.

Review code, debug issues, and perform quality assurance (QA) on software components to enhance system reliability and performance.

Work closely with development teams to ensure best practices in software architecture, coding standards, and operational readiness.

Forecast scalability needs and optimize cloud infrastructure costs while balancing performance and efficiency.

Ensure production environments meet security and compliance requirements, collaborating with teams to mitigate vulnerabilities and enforce best practices.

Work closely with development teams to embed reliability at every stage rather than treating it as an afterthought.

Use error budgets to balance feature velocity with system stability.

Implement observability and automation-first principles to measure system health and drive continuous improvement.

Leverage game days, chaos engineering, and resilience testing to validate system robustness and refine operational processes.

What you will have

Extensive experience in distributed systems, cloud-native architectures (AWS, GCP, Azure), and DevOps practices.

Proficiency in Kubernetes, Terraform, CI/CD pipelines, and Infrastructure as Code (IaC).

Strong scripting and automation skills in Python, Go, Bash, or similar languages.

Expertise in observability tools such as Prometheus, Grafana, Datadog, Splunk, New Relic, and OpenTelemetry.

Ability to troubleshoot complex production issues and drive scalable, resilient solutions.

Experience reviewing code, debugging applications, and conducting software testing to ensure high reliability and quality.

Benefits

Competitive compensation and comprehensive health, dental, and vision coverage

Retirement savings programs with company matching (401(k) or RRSP)

Generous paid time off, including holidays, and vacation

Paid parental leave and family support benefits

Life and disability coverage

Flexible spending and health savings options (where applicable)

Health and wellness support, including gym reimbursement and wellness programs

Employee Assistance Program with 24/7confidential support for employees and families

Education assistance and professional development opportunities

Access to LinkedIn Learning and continuous learning resources

Employee referral bonus program and additional company perks and discounts

Internal rewards and recognition platform (Motivosity) to celebrate and acknowledge project wins, milestone achievements, and the outstanding contributions of our colleagues

Business travel insurance and global employee support programs

#LI-RR1

Compensation

This DevOps / SRE role pays $125k-$145k/yr. Within typical range for devops / sre roles in United States.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for DevOps / SRE roles in United States varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our DevOps / SRE hub for United States medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.