Sr. Monitoring and Observability Engineer
Skills
About the role
About Analog Devices
Analog Devices, Inc. (NASDAQ: ADI ) is a global semiconductor leader that bridges the physical and digital worlds to enable breakthroughs at the Intelligent Edge. ADI combines analog, digital, AI, and software technologies into solutions that combat climate change, reliably connect humans and the world, and help drive advancements in automation and robotics, mobility, healthcare, energy and data centers. With revenue of more than $11 billion in FY25, ADI ensures today's innovators stay Ahead of What's Possible. Learn more at www.analog.com and on LinkedIn and X .
Position Overview
We are seeking an experienced Senior Monitoring and Observability Engineer to lead our monitoring and observability initiatives. This role combines strategic responsibility for defining monitoring standards and KPIs with hands-on technical expertise in building and maintaining robust observability infrastructure. You will drive data-driven decision-making, mentor team members, and continuously enhance the user experience across our monitoring ecosystem.
Key Responsibilities
Design, implement, and maintain comprehensive monitoring strategies across Linux, Storage, Applications, and Networking platforms
Define and establish KPI/Metrics standards and best practices for the organization, ensuring consistency and measurability
Architect and deploy monitoring solutions using industry-leading tools including Prometheus, Grafana, Nagios, and SolarWinds
Develop and implement Kafka-based event streaming and monitoring pipelines for real-time data collection and analysis
Integrate and manage AI monitoring capabilities to enable intelligent alerting, anomaly detection, and predictive insights
Create and maintain comprehensive documentation and runbooks for monitoring infrastructure, troubleshooting procedures, and operational playbooks
Develop custom scripts and automation in bash, Python, or Go to enhance monitoring capabilities and reduce operational overhead
Establish and refine SLAs, SLOs, and alerting thresholds based on business requirements and operational data
Conduct performance analysis and optimization of monitoring infrastructure to reduce latency and improve system reliability
Collaborate with DevOps, Platform, and Application teams to ensure comprehensive observability across the entire stack
Stay current with emerging monitoring and observability trends, tools, and technologies
Drive continuous improvement initiatives by analyzing metrics, gathering feedback, and implementing enhancements to the user experience
Required Qualifications
Minimum 8+ years of professional experience in monitoring, observability, systems engineering, or DevOps roles
Proven experience leading and mentoring technical teams in monitoring or infrastructure domains
Deep hands-on expertise with opensource monitoring tools such as Prometheus, Grafana, and Nagios
Strong experience with enterprise monitoring platforms like SolarWinds
Proficiency in bash scripting and Linux system administration (RHEL/CentOS/Ubuntu preferred)
Experience with event-driven architectures and Apache Kafka for real-time data processing
Knowledge of monitoring and observability across Linux systems, storage infrastructure, applications, and network platforms
Understanding of metrics collection, time-series data management, and data-driven decision-making
Experience implementing AI/ML-based monitoring solutions or anomaly detection systems
Strong problem-solving skills and ability to work effectively in fast-paced, complex environments
Excellent communication skills with ability to present technical concepts to both technical and non-technical stakeholders
Experience with Infrastructure-as-Code (IaC) tools (Terraform, Ansible, etc.)
Preferred Qualifications
Experience with containerized environments and Kubernetes monitoring
Proficiency in Python or Go for custom tooling development
Experience with cloud platforms (AWS, Azure, GCP) and their native monitoring services
Experience with observability backends such as OpenTelemetry.
Knowledge of application performance monitoring (APM) tools
Experience with incident response and post-mortem processes
Certification in relevant areas (e.g., Kubernetes, cloud platforms, monitoring tools)
Contribution to open-source monitoring projects
Required Technical Skills
Monitoring Platforms: Prometheus, Grafana, Nagios, Azure Monitor, AWS CloudWatch, SolarWinds
Data Streaming: Apache Kafka, event processing pipelines
AI Monitoring: Anomaly detection, intelligent alerting, ML-based insights
Scripting Languages: Bash, Python, Go
Operating Systems: Linux (RHEL, CentOS, Ubuntu), system administration
Infrastructure: Storage systems, networking, application monitoring
Data Visualization: Grafana dashboards, custom visualizations
Cloud & Containerization: Docker, Kubernetes (optional but valuable)
Strategic Responsibilities
As a leader in this role, you will be responsible for:
Developing and communicating the monitoring and observability strategy aligned with organizational goals
Setting industry-standard KPIs and metrics that drive business value and operational excellence
Using data-driven approaches to identify improvement opportunities and measure success
Building a scalable, maintainable monitoring infrastructure that grows with the organization
Establishing clear performance baselines and continuous improvement metrics
Ensuring that monitoring initiatives directly improve user experience and system reliability
What We're Looking For
We seek a hands-on technical leader who is passionate about observability and excellence. Ideal candidates will:
Balance strategic thinking with hands-on technical expertise
Demonstrate a track record of building high-performing teams
Show deep curiosity about system behavior and data-driven insights
Possess strong written and verbal communication skills
Take pride in mentoring others and sharing knowledge
Approach challenges with creativity and persistence
Stay current with evolving monitoring and observability technologies
For positions requiring access to technical data, Analog Devices, Inc. may have to obtain export licensing approval from the U.S. Department of Commerce - Bureau of Industry and Security and/or the U.S. Department of State - Directorate of Defense Trade Controls. As such, applicants for this position – except US Citizens, US Permanent Residents, and protected individuals as defined by 8 U.S.C. 1324b(a)(3) – may have to go through an export licensing review process.
Job Req Type: Experienced
Required Travel: Yes, 10% of the time
Shift Type: 1st Shift/Days
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.