Site Reliability Engineer (SRE) – Azure AKS, Observability & Terraform
Astra-North Infoteck Inc. ~ Conquering today’s challenges, achieving tomorrow’s vision!
Skills
About the role
Site Reliability Engineer (SRE) – Azure AKS, Observability & Terraform
Key Responsibilities
Observability, SRE, DevOps roles with expertise in infrastructure and application reliability
Dynatrace, ELK, Splunk, PagerDuty
SLI/SLO frameworks
Azure Kubernetes Service (AKS), Terraform, Azure managed services
What will you do
Design and implement observability-as-code solutions using Terraform for monitoring pipelines, dashboards, and alerting across distributed systems
Drive observability improvements using Dynatrace, ELK, Splunk, PagerDuty for real-time performance insights and system visibility
Instrument applications for end-to-end observability including distributed tracing, metrics collection, and log aggregation across Node.js and .NET microservices and event-driven architectures
Troubleshoot complex production incidents across service layers, databases, caches, and APIs using SLI/SLO frameworks
Investigate and resolve Azure Kubernetes Service (AKS) infrastructure issues ensuring reliability and scalability of containerized workloads using Terraform and Azure services (SQL MI, Redis, Functions, Event Grid)
Translate business requirements into observable, resilient systems aligned to SLIs/SLOs
Automate operational tasks using Infrastructure-as-Code and CI/CD to reduce toil and improve resilience
Lead incident response and remediation for critical systems, including blameless postmortems and chaos engineering practices
Collaborate with development, platform, and business teams to improve availability, scalability, and operational excellence
What do you need to succeed
Must-have
8+ years experience in SRE, DevOps, or Observability roles focused on infrastructure and application reliability
Strong expertise in Dynatrace, ELK, Splunk, PagerDuty and observability principles (instrumentation, correlation IDs, SLIs/SLOs)
Advanced proficiency in Azure Kubernetes Service (AKS), Terraform, and Azure managed services (SQL MI, Redis, Functions, Event Grid)
Hands-on experience with observability instrumentation (distributed tracing, metrics, logs) across Node.js and .NET microservices and event-driven systems
Strong troubleshooting skills across distributed systems (services, databases, caches, APIs) in production environments
Incident management expertise using PagerDuty and ServiceNow, including high-severity incident resolution and RCA
Knowledge of incident, problem, and change management, SRE principles, blameless postmortems, and chaos engineering
Strong communication and leadership skills for cross-functional coordination and incident handling
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.