Lead Associate Principal, Software Engineering – DevOps / SRE

New York Technology Partners

Chicago, USonsitePosted Jun 9, 2026
Posting intelligenceActively listedReposted 3×, possible evergreen/ghost posting

Skills

kubernetesprometheusterraformpagerdutyopsgeniejenkinsansiblegrafanadatadoggopythonkafkascalacicdjavaaws

About the role

We are seeking a highly skilled Lead DevOps / Site Reliability Engineer (SRE) to drive site reliability, system observability, and operational excellence across our core platform. In this leadership role, you will collaborate with product, infrastructure, security, and operations teams to build and scale multi-region Kubernetes environments, automate infrastructure, and embed reliability into the software development lifecycle.

If you are passionate about eliminating toil, designing robust observability frameworks, and building highly available cloud architectures, we want you on our team!

What You Will Do:

Platform & Infrastructure Automation: Guide the implementation of CI/CD pipelines in Kubernetes environments. Direct the configuration and execution of Terraform and Ansible automation pipelines. Setup and manage common infrastructure platforms, including multi-region Kubernetes and Kafka clusters.

Site Reliability & Observability: Define and enforce SLOs, SLIs, and Error Budgets. Design and implement comprehensive observability frameworks (metrics, logs, distributed tracing). Lead blameless post-mortems and drive toil reduction initiatives.

Architecture & Scalability: Conduct capacity planning, load forecasting, and implement/validate infrastructure failover scenarios (Multi-AZ, multi-region). Confers with Enterprise Architecture teams to devise robust implementation architectures.

Security & Incident Response: Promote the highest security posture by collaborating on least-privilege access management (Hashicorp Vault). Own and continuously improve incident response runbooks, on-call rotations, and escalation procedures.

Cross-Functional Leadership: Partner with product teams to embed reliability and Non-Functional Requirements (NFRs) early in the SDLC. Follow and enable program-level Agile practices for efficient delivery.

The Technical Toolbox (Required Skills):

Cloud & Infrastructure: AWS (EC2), Kubernetes, Terraform, Ansible, Hashicorp Vault.

Observability & SRE Tooling: Prometheus, Grafana, OpenTelemetry, Datadog, PagerDuty, OpsGenie.

Messaging & Data Streaming: Kafka, Microservices architecture, streaming data-intensive applications.

Development & CI/CD: Jenkins, Git, Artifactory, API-based development. Fluency in at least one of: Java, Python, Scala, or Golang.

Data & Databases: SQL and NoSQL databases, in-memory data stores. Fluent in JSON, Protobuf, Avro.

Advanced Engineering: Chaos engineering principles and tooling (Chaos Monkey, Gremlin, LitmusChaos), Multi-AZ/multi-region failover architecture.

Frameworks: Familiarity with enterprise architecture frameworks (e.g., TOGAF).

What You Bring (Qualifications):

Experience: 7+ years of experience building large-scale, data-centric solutions, with 7+ years of recent experience on a DevOps or SRE team (or as a Product Owner for one).

Education: BS degree in Computer Science, a similar technical field, or equivalent practical experience.

SRE & Agile Mastery: Deep understanding of Google SRE practices (error budgets, toil elimination, reliability hierarchy) and Agile/Kanban methodologies.

Domain Knowledge: Broad experience in two or more of the following: web/mobile app development, Unix/Linux environments, event-driven systems, transaction processing, distributed/parallel systems, or public-cloud platforms.

Soft Skills: A self-starter who anticipates needs, comfortable communicating complex technical concepts to both technical and non-technical audiences, and a humble, collaborative team player.

Bonus Points (Preferred):

AWS Certified Solutions Architect / DevOps Engineer

Certified Kubernetes Administrator (CKA) / Kafka Certifications

Google Cloud Professional – Site Reliability Engineer (or equivalent SRE certification)

Project/Program management certifications

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for DevOps / SRE roles in United States varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our DevOps / SRE hub for United States medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.