Senior DevOps Engineer – Observability Platform
Skills
About the role
Company:
Qualcomm India Private Limited
Job Area:
Engineering Group, Engineering Group > Software Engineering
General Summary:
As a Senior DevOps Engineer – Observability Platform , you will be responsible for building and maintaining scalable, reliable infrastructure and deployment pipelines with a strong emphasis on observability - metrics, logs, and traces - across systems running on Kubernetes and AWS. You will work closely with development teams to improve development velocity while ensuring system reliability, security, and performance. This role is critical in providing a standardized, observability platform that gives both internal engineering teams and external, customer-facing services deep, reliable visibility into system health, performance, and reliability
Minimum Qualifications:
Bachelor's degree in Engineering, Information Systems, Computer Science, or related field and 2+ years of Software Engineering or related work experience.
OR
Master's degree in Engineering, Information Systems, Computer Science, or related field and 1+ year of Software Engineering or related work experience.
OR
PhD in Engineering, Information Systems, Computer Science, or related field.
2+ years of academic or work experience with Programming Language such as C, C++, Java, Python, etc.
Infrastructure Management : Design, implement, and maintain cloud-based infrastructure using Infrastructure as Code principles
Automation : Develop automation scripts and tools to streamline operations and eliminate manual processes
Containerization : Manage containerization strategies and orchestration using Docker and Kubernetes
Observability Platform: Design, build, and operate a standardized, self-service metrics, logs, and tracing platform (Prometheus, Grafana, Loki, OpenTelemetry) serving both internal teams and external, customer-facing services running on Kubernetes and AWS
Instrumentation & Telemetry: Partner with engineering teams to instrument applications and infrastructure, standardizing telemetry collection with OpenTelemetry
SLOs & Alerting: Define and maintain SLIs/SLOs and error budgets, build actionable dashboards, and tune alerting to maximize signal and reduce noise
Performance Optimization : Use observability data to analyze and optimize system performance, scalability, and cost-efficiency
Documentation : Create and maintain thorough documentation for infrastructure, deployment processes, and operational procedures
Incident Response & Escalation: Provide second-tier engineering escalation during business hours and own the telemetry, SLO, and alerting tooling that powers incident detection and reduces MTTD/MTTR; front-line 24/7 on-call is owned by the dedicated SRE team, not observability engineers. Lead post-mortem analysis for observability-platform incidents
Requirements
Qualifications
5+ years of experience in DevOps, Observability, or similar roles, including hands-on production experience operating Kubernetes based stack.
Advantage - Strong background in software development with security focus
Technical Skills
Cloud Platforms : Extensive hands-on experience with AWS, including its observability services (CloudWatch, X-Ray, Amazon Managed Service for Prometheus, Amazon Managed Grafana)
Infrastructure as Code : Proficiency with Terraform, AWS CloudFormation, or similar IaC tools
Containerization : Advanced knowledge of Docker and Kubernetes ecosystem
Observability Stack: Hands-on experience with Prometheus, Grafana, Loki, Tempo or Jaeger, OpenTelemetry, and Alertmanager; experience scaling metrics storage with Thanos, Mimir, or Cortex
Programming/Scripting : Strong coding skills in Python, Bash, or Go
Soft Skills
Problem-Solving: Excellent analytical and troubleshooting skills.
Communication: Strong verbal and written communication skills.
Collaboration: Ability to work effectively in a team environment and collaborate with cross-functional teams.
Leadership: Proven leadership skills and the ability to mentor junior engineers.
Remote Work: Comfortable working in a fully distributed, offshore setup and collaborating effectively with development teams across multiple locations.
Qualcomm expects its employees to abide by all applicable policies and procedures, including but not limited to security and other requirements regarding protection of Company confidential information and other confidential and/or proprietary information, to the extent those requirements are permissible under applicable law.
To all Staffing and Recruiting Agencies : Our Careers Site is only for individuals seeking a job at Qualcomm. Staffing and recruiting agencies and individuals being represented by an agency are not authorized to use this site or to submit profiles, applications or resumes, and any such submissions will be considered unsolicited. Qualcomm does not accept unsolicited resumes or applications from agencies. Please do not forward resumes to our jobs alias, Qualcomm employees or any other company location. Qualcomm is not responsible for any fees related to unsolicited resumes/applications.
If you would like more information about this role, please contact Qualcomm Careers .
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.