Lead DevOps Engineer

New York Technology Partners

Chicago, UShybridPosted Jun 8, 2026
Posting intelligenceActively listedReposted 3×, possible evergreen/ghost posting

Skills

kubernetesprometheusterraformpagerdutyopsgeniejenkinsansiblegrafanadatadoggopythonkafkascalacicdjavaaws

About the role

Job Title: Lead Associate Principal, Software Engineering: DevOps

Location: Chicago, IL (Onsite – Hybrid - 3 days in a week)

Position Type: Fulltime permanent position

Responsibilities:

Qualifications:

The requirements listed are representative of the knowledge, skill, and/or ability required. Reasonable accommodations may be made to enable individuals with disabilities to perform the primary functions.

[Required] Understanding of Kanban and/or Agile methodologies

[Required] Familiarity with SRE principles as defined by Google SRE practices (error budgets, toil elimination, reliability hierarchy)

[Required] Able to succeed in a fast-paced environment with frequent changes

[Required] Comfortable communicating with both technical and non-technical audiences

[Required] Self-starter — takes initiative to research, learn, and deliver; anticipates the play

[Required] Team player — humble, collaborative, and focused on making the entire team succeed

Technical Skills & Background:

[Required] AWS EC2, Kubernetes, Kafka, Jenkins, Terraform, Ansible, Hashicorp Vault

[Required] Observability tooling such as Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent

[Required] Incident management platforms and on-call tooling (e.g., PagerDuty, OpsGenie)

[Required] Microservices and streaming data-intensive application architecture

[Required] Application architecture, networking, and security in the cloud

[Required] Setting up platforms in AWS for high-performance requirements

[Required] Broad experience in API-based development

[Required] Git and Artifactory for sourcing artifacts

[Required] Multi-AZ, multi-region failover architecture

[Required] Chaos engineering principles and tooling (e.g., Chaos Monkey, Gremlin, LitmusChaos)

[Required] Fluent with different data formats and structures: JSON, Protobuf, Avro

[Required] SQL and NoSQL databases, in-memory data stores

[Required] Java/Python/Scala/Golang software development

[Required] Two or more of the following: web/mobile application development, Unix/Linux environments, event-driven systems, transaction processing systems, distributed and parallel systems, large software system development, security software development, public-cloud platforms

[Required] Fluent in industry best practices, software patterns, and architecture principles

[Required] Enterprise architecture frameworks such as TOGAF

[Required] Ability to define and document architecture strategies, designs, and requirements across all enterprise architecture domains

[Required] Ability to define service-based, component architectures and demonstrate visualization of enterprise architecture concepts

Certifications:

[Preferred] AWS Certified Solutions Architect / DevOps Engineer

[Preferred] Kubernetes, Kafka certification

[Preferred] Google Cloud Professional — Site Reliability Engineer or equivalent SRE-focused certification

[Preferred] Project/program management certifications

Education & Training:

[Required] BS degree in Computer Science, similar technical field, or equivalent experience

[Required] 7+ years of experience building large-scale, data-centric solutions

[Required] 7+ years of recent experience participating on a DevOps or SRE team, or as product owner for such a team

Responsibilities:

To perform this job successfully, an individual must be able to perform each primary duty satisfactorily.

Guides the implementation using CI/CD pipelines in Kubernetes environment

Directs review, configuration, and execution of Terraform and Ansible automation pipelines delivered by product teams

Guides the setup of common infrastructure platforms like multi-region Kubernetes and Kafka clusters

Elicits requirements for application deployment and sizing to manage expected workloads

Defines and enforces Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets in collaboration with product teams

Leads blameless post-mortems and drives resolution of action items to reduce repeat incidents

Designs and implements observability frameworks covering metrics, logs, and distributed tracing across all platform services

Drives toil reduction initiatives by identifying and automating repetitive operational work

Partners with product teams to embed reliability requirements and non-functional requirements (NFRs) early in the software development lifecycle

Monitors application performance and tunes systems working with product teams

Confers with product team leads and practitioners to create deployment and reliability plans

Confers with Enterprise Architecture and Renaissance architecture teams to devise implementation architecture

Promotes standards across application configuration towards the highest security posture

Collaborates with access management and security teams on setting up roles and permissions using least privilege strategies

Collaborates with integration/performance testing teams to leverage integrated release testing in the Release Acceptance environment

Collaborates with production controls teams on monitoring, failover, logging, and alerting strategies

Owns and continuously improves incident response runbooks, on-call rotations, and escalation procedures

Conducts capacity planning and load forecasting to proactively address scalability needs

Implements and validates infrastructure failover scenarios

Confers with Network team on all connectivity plans and issue resolution (including between on-premises and AWS)

Follows and enables program-level agile practices for efficient collaboration and delivery

Develops documentation for ORT technical infrastructure, architecture, and reliability support

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for DevOps / SRE roles in United States varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our DevOps / SRE hub for United States medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.