Lead I - Cloud Infrastructure Services-Appdynamics

UST

unknownPosted Jul 11, 2026
Posting intelligenceActively listedReposted 3×, possible evergreen/ghost posting

Skills

kubernetesprometheusgrafanapythonazurecicdaws

About the role

Role description

Senior Observability Engineer Experience

5+ years

Location

Any UST location

Mandatory Skills

Python

AppDynamics

AWS CloudWatch

Grafana

Observability Engineering

Site Reliability Engineering (SRE)

Monitoring, Logging & Tracing

REST APIs

CI/CD & DevOps

Automation (Python scripting)

Kubernetes (Preferred)

OpenTelemetry (Preferred)

ELK/OpenSearch or Prometheus (Preferred)

We are seeking a highly motivated and technically proficient Senior Observability Engineer to join our DSX Observability Engineering organization. The ideal candidate will have strong expertise in Observability platforms, Python development, and Site Reliability Engineering (SRE) practices to design, implement, and continuously improve monitoring, ing, logging, tracing, and reliability solutions across critical business applications and cloud platforms.

This role will be responsible for building scalable observability frameworks, driving operational excellence, improving service reliability, and enabling engineering teams with actionable insights into application and infrastructure performance.

Key Responsibilities

Observability Engineering

Design, implement, and maintain enterprise-wide observability solutions covering metrics, logs, traces, and user experience monitoring.

Develop and manage dashboards, s, and service health indicators using Grafana, CloudWatch, AppDynamics, ELK and related observability tools.

Define and implement SLIs, SLOs, and Error Budgets for critical services.

Drive adoption of observability best practices across engineering teams.

Establish standards for instrumentation, monitoring, ing, and incident response.

Build end-to-end observability solutions leveraging OpenTelemetry and cloud-native monitoring capabilities.

Analyze application and infrastructure performance data to identify trends, bottlenecks, and optimization opportunities.

Site Reliability Engineering (SRE)

Improve platform reliability, scalability, availability, and operational efficiency.

Participate in incident management, root cause analysis (RCA), and post-incident reviews.

Drive proactive reliability initiatives through automation and engineering improvements.

Support chaos engineering, resiliency testing, and disaster recovery exercises.

Define and monitor operational KPIs related to system health and reliability.

Collaborate with development teams to improve production readiness and operational excellence.

Python Development & Automation

Develop automation tools and frameworks using Python to streamline observability and operational workflows.

Build integrations between monitoring platforms, cloud services, ticketing systems, and collaboration tools.

Automate enrichment, reporting, dashboard provisioning, and operational tasks.

Develop scripts and APIs for data collection, analysis, and observability enhancements.

Create self-service solutions that improve developer productivity and operational visibility.

Cloud & Platform Monitoring

Implement monitoring solutions for cloud-native architectures hosted on AWS/AKS/Azure.

Configure and optimize CloudWatch metrics, logs, alarms, and dashboards.

Monitor distributed systems, microservices, APIs, containers, and serverless workloads.

Support observability for Kubernetes, containerized applications, and modern cloud platforms.

Ensure monitoring coverage for business-critical services and customer journeys.

Skills

site reliability engineering,aws cloud watch,python,appdynamics,grafana,kubernetes,

About UST

UST is a global digital transformation solutions provider. For more than 20 years, UST has worked side by side with the world’s best companies to make a real impact through transformation. Powered by technology, inspired by people and led by purpose, UST partners with their clients from design to operation. With deep domain expertise and a future-proof philosophy, UST embeds innovation and agility into their clients’ organizations. With over 30,000 employees in 30 countries, UST builds for boundless impact - touching billions of lives in the process.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation varies by seniority, employer size, and location. When this listing publishes a salary band you'll see it in the badge row above the description.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.