Site Reliability Engineer

Spait Infotech

Toronto, CAonsitePosted Jul 16, 2026
Posting intelligenceActively listed

Skills

cloudformationazure devopskubernetesprometheusterraformjenkinsansiblegrafanadatadogdockergithubgitlabpulumipythonazurecicdjavagooglecloudawsgo

About the role

Key Responsibilities

Design, implement, and maintain highly available and scalable infrastructure.

Monitor system performance, availability, and reliability using observability tools.

Automate infrastructure provisioning, deployment, and operational tasks.

Build and maintain CI/CD pipelines.

Troubleshoot production issues and perform root cause analysis (RCA).

Participate in on-call rotations and respond to critical incidents.

Collaborate with development teams to improve application reliability and performance.

Define and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Service Level Agreements (SLAs).

Implement disaster recovery, backup, and business continuity strategies.

Continuously improve system security, reliability, and operational efficiency.

Required Skills

Strong experience with Linux/Unix system administration and shell scripting.

Proficiency in one or more programming/scripting languages such as Python, Go, Bash, or Java.

Hands-on experience with cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform (GCP).

Experience with containerization and orchestration technologies, including Docker and Kubernetes.

Expertise in Infrastructure as Code (IaC) tools such as Terraform, CloudFormation, or Pulumi.

Experience building and maintaining CI/CD pipelines using tools such as Jenkins, GitHub Actions, GitLab CI, or Azure DevOps.

Strong knowledge of monitoring, logging, and observability tools such as Prometheus, Grafana, ELK Stack, Datadog, Splunk, or New Relic.

Solid understanding of networking concepts, including TCP/IP, DNS, HTTP/HTTPS, load balancing, SSL/TLS, and firewalls.

Experience with version control systems such as Git.

Knowledge of configuration management and automation tools such as Ansible, Puppet, or Chef.

Experience in incident management, root cause analysis (RCA), and production support.

Understanding of high availability, scalability, disaster recovery, and business continuity planning.

Familiarity with SRE best practices, including Service Level Indicators (SLIs), Service Level Objectives (SLOs), Service Level Agreements (SLAs), and error budgets.

Strong troubleshooting, analytical, and problem-solving skills.

Excellent communication, collaboration, and stakeholder management skills.

Ability to work in an on-call production support environment and manage critical incidents effectively.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for DevOps / SRE roles in Canada varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our DevOps / SRE hub for Canada medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.