Software Development Engineer III - DevOps
Skills
About the role
Description
Responsibilities:
Own the reliability, scalability, and performance of production systems and services.
Design and implement highly available, fault-tolerant, and distributed infrastructure.
Define and drive observability strategy, including monitoring, logging, and alerting.
Build and maintain scalable CI/CD pipelines to enable fast and reliable deployments.
Automate infrastructure provisioning and operational workflows using IaC tools.
Lead incident management, root cause analysis (RCA), and implement preventive measures.
Define and track SLIs, SLOs, and SLAs aligned with business and product requirements.
Collaborate closely with engineering teams to improve system design, deployment processes,
and operational excellence.
Optimize cloud infrastructure for cost, performance, and efficiency.
Own and improve on-call processes; mentor engineers in handling production incidents.
Drive best practices for security, networking, and infrastructure reliability.
Participate in architecture and design reviews to ensure system resilience and scalability.
Document system architecture, runbooks, and operational processes.
Requirements:
5–8+ years of experience in DevOps, SRE, or Cloud Engineering roles.
Proven experience building software using Python, Go, Rust, or JavaScript, with strong scripting
capabilities.
Hands-on experience with cloud platforms such as AWS, GCP, or Azure.
Expertise in Infrastructure as Code tools like Terraform, Ansible, or CloudFormation.
Strong experience with containerization (Docker) and orchestration (Kubernetes).
Solid understanding of Linux systems, networking concepts, and security best practices (IAM,
VPNs, firewalls).
Experience with monitoring and observability tools like Prometheus, Grafana, ELK Stack, New
Relic, or Datadog.
Proven experience in building and maintaining CI/CD pipelines (Circle CI, ArgoCD, Jenkins,
GitHub Actions, GitLab CI, etc.).
Deep hands-on experience owning cloud security end-to-end.
Strong debugging and troubleshooting skills for complex production systems.
Comfortable operating in fast-moving engineering environments, balancing long-term
infrastructure investment with immediate operational needs.
Experience owning and continuously improving on-call processes - including rotation design,
escalation policies, runbook culture, and post-incident review cadence.
Nice to Have:
Experience with large-scale distributed systems and microservices architecture.
Experience building and running operators on Kubernetes / knowledge of internals.
Hands-on experience with cost optimization and capacity planning in cloud environments.
Experience mentoring engineers and driving engineering best practices.
Prior involvement in architectural reviews and cross-team technical decision-making.
Strong understanding of compliance, security standards, and DevSecOps practices.
Experience building internal developer platforms or self-service infrastructure tools.
Development exposure – experience contributing to backend services or application code (e.g.,
APIs, microservices), with strong software engineering fundamentals.
MLOps experience – familiarity with deploying, monitoring, and managing ML models in
production, including tools like ML pipelines, model versioning, and data workflows.
About Company:
TELUS Digital is the customer experience transformation partner to the world's most admired brands. Our diverse team weaves data, technology, and human ingenuity to deliver differentiated customer journeys, drive operational effectiveness, and scale AI solutions with meaningful value and positive impact.
We craft real-world solutions in the moments that matter, from customer acquisition to lifelong loyalty. Enabled by our global reach - spanning 78,000 experts in 33 countries - and deep industry expertise, we help over 600 organizations make the customer experience feel effortless.
Our solutions span Data & AI, Digital Experience & IT, CX Management and Trust & Safety. At the core of our innovation is Fuel iX™, an enterprise-grade generative AI platform that helps clients safely access and optimize leading LLMs to scale their own AI from pilot to production.
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.