Lead Site Reliability Engineer - Infrastructure & DevOps

J2B GLOBAL LLC

USonsitePosted Jul 20, 2026
Posting intelligenceActively listed

Skills

azure devopspostgreskubernetesprometheusterraformmongojenkinsgithubgitlabpythonazurekafkaredishelmcicdaws

About the role

Lead Site Reliability Engineer - Infrastructure & DevOps

Orlando, FL

Visa: US CitizenGreen CardGC-EADH1BH4-EADTNEAD

Job Title: Lead Site Reliability Engineer (SRE) Overview / Summary We are seeking a Lead Site Reliability Engineer to help drive the reliability, scalability, and operational excellence of a rapidly growing Generative AI platform. This role provides technical leadership while designing and supporting highly available cloud infrastructure powering modern AI and data-driven applications. The ideal candidate combines deep expertise in Site Reliability Engineering, cloud infrastructure, Kubernetes, and Infrastructure as Code with strong leadership skills. You will work alongside platform engineers, architects, and development teams to build resilient systems, improve automation, and ensure high availability across a multi-cloud environment. Key Responsibilities Lead the design, implementation, and support of highly available cloud infrastructure across Google Cloud Platform (primary), AWS, and Azure. Design, build, and maintain Kubernetes infrastructure using Helm and Terraform for Infrastructure as Code. Develop scalable platform solutions capable of maintaining 99.99% service availability. Lead and mentor Site Reliability Engineers and DevOps engineers by providing technical guidance and establishing engineering best practices. Plan, prioritize, and coordinate infrastructure initiatives within Agile delivery teams. Design and implement automated deployment pipelines using modern CI/CD tools, including Harness. Implement progressive deployment strategies such as blue/green deployments, canary releases, and feature flag rollouts. Build and enhance observability solutions using monitoring, logging, alerting, and distributed tracing technologies. Partner with engineering teams to review infrastructure sizing, capacity planning, and scalability requirements. Support production systems through backups, upgrades, patching, disaster recovery, and operational maintenance. Troubleshoot complex production issues across distributed systems and cloud-native applications. Evaluate emerging SRE and DevOps technologies and recommend improvements to platform reliability and operational efficiency. Ensure infrastructure aligns with security, governance, and compliance standards. Required Qualifications 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related infrastructure roles. Expert-level experience administering and operating Kubernetes in production environments. Strong experience with Helm for Kubernetes application management. Advanced experience using Terraform for Infrastructure as Code. Hands-on experience building automated deployment pipelines using Harness or comparable enterprise CI/CD platforms. Experience supporting production workloads across Google Cloud Platform, AWS, and Azure. Strong scripting and automation skills using Python, Bash, and YAML. Experience supporting production databases and messaging technologies, including PostgreSQL, Redis, Kafka, MongoDB, and Vault. Experience with enterprise CI/CD platforms such as GitHub Actions, GitLab CI, Jenkins, Azure DevOps, or Harness. Experience implementing observability solutions using technologies such as OpenTelemetry, Prometheus, Splunk, AppDynamics, or similar platforms. Strong troubleshooting skills within distributed systems and cloud-native environments. Experience working within Agile development environments. Excellent communication skills with the ability to explain complex technical concepts to both technical and non-technical audiences.

Refer

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for DevOps / SRE roles in United States varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our DevOps / SRE hub for United States medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.