Site Reliability Engineer
Skills
About the role
Key Responsibilities
Design, implement, and maintain highly available and scalable infrastructure.
Monitor system performance, availability, and reliability using observability tools.
Automate infrastructure provisioning, deployment, and operational tasks.
Build and maintain CI/CD pipelines.
Troubleshoot production issues and perform root cause analysis (RCA).
Participate in on-call rotations and respond to critical incidents.
Collaborate with development teams to improve application reliability and performance.
Define and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).
Implement disaster recovery, backup, and business continuity strategies.
Continuously improve system security, reliability, and operational efficiency.
Required Skills
Strong experience with Linux/Unix system administration and shell scripting.
Proficiency in one or more programming/scripting languages such as Python, Go, Bash, or Java.
Hands-on experience with cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform (GCP).
Experience with containerization and orchestration technologies, including Docker and Kubernetes.
Expertise in Infrastructure as Code (IaC) tools such as Terraform, CloudFormation, or Pulumi.
Experience building and maintaining CI/CD pipelines using tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.
Strong knowledge of monitoring, logging, and observability tools such as Prometheus, Grafana, ELK Stack, Datadog, Splunk, or New Relic.
Solid understanding of networking concepts, including TCP/IP, DNS, HTTP/HTTPS, load balancing, SSL/TLS, and firewalls.
Experience with version control systems such as Git.
Knowledge of configuration management and automation tools such as Ansible, Puppet, or Chef.
Experience in incident management, root cause analysis (RCA), and production support.
Understanding of high availability, scalability, disaster recovery, and business continuity planning.
Familiarity with SRE best practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), Service Level Agreements (SLAs), and error budgets.
Strong troubleshooting, analytical, and problem-solving skills.
Excellent communication, collaboration, and stakeholder management skills.
Ability to work in an on-call production support environment and manage critical incidents effectively.
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.