AI

Site Reliability Engineer (SRE) – Azure AKS, Observability & Terraform

Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision!

Toronto, CAonsitePosted Jun 23, 2026
Posting intelligenceActively listedReposted 4×, possible evergreen/ghost posting

Skills

kubernetesterraformpagerdutynodeazurerediscicdjavascript

About the role

Site Reliability Engineer (SRE) – Azure AKS, Observability & Terraform

Key Responsibilities

Observability, SRE, DevOps roles with expertise in infrastructure and application reliability

Dynatrace, ELK, Splunk, PagerDuty

SLI/SLO frameworks

Azure Kubernetes Service (AKS), Terraform, Azure managed services

What will you do

Design and implement observability-as-code solutions using Terraform for monitoring pipelines, dashboards, and alerting across distributed systems

Drive observability improvements using Dynatrace, ELK, Splunk, PagerDuty for real-time performance insights and system visibility

Instrument applications for end-to-end observability including distributed tracing, metrics collection, and log aggregation across Node.js and .NET microservices and event-driven architectures

Troubleshoot complex production incidents across service layers, databases, caches, and APIs using SLI/SLO frameworks

Investigate and resolve Azure Kubernetes Service (AKS) infrastructure issues ensuring reliability and scalability of containerized workloads using Terraform and Azure services (SQL MI, Redis, Functions, Event Grid)

Translate business requirements into observable, resilient systems aligned to SLIs/SLOs

Automate operational tasks using Infrastructure-as-Code and CI/CD to reduce toil and improve resilience

Lead incident response and remediation for critical systems, including blameless postmortems and chaos engineering practices

Collaborate with development, platform, and business teams to improve availability, scalability, and operational excellence

What do you need to succeed

Must-have

8+ years experience in SRE, DevOps, or Observability roles focused on infrastructure and application reliability

Strong expertise in Dynatrace, ELK, Splunk, PagerDuty and observability principles (instrumentation, correlation IDs, SLIs/SLOs)

Advanced proficiency in Azure Kubernetes Service (AKS), Terraform, and Azure managed services (SQL MI, Redis, Functions, Event Grid)

Hands-on experience with observability instrumentation (distributed tracing, metrics, logs) across Node.js and .NET microservices and event-driven systems

Strong troubleshooting skills across distributed systems (services, databases, caches, APIs) in production environments

Incident management expertise using PagerDuty and ServiceNow, including high-severity incident resolution and RCA

Knowledge of incident, problem, and change management, SRE principles, blameless postmortems, and chaos engineering

Strong communication and leadership skills for cross-functional coordination and incident handling

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for DevOps / SRE roles in Canada varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our DevOps / SRE hub for Canada medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.