Associate Director Engineering - SRE

HSBC

unknownPosted Jul 9, 2026
Posting intelligenceActively listedReposted 29×, possible evergreen/ghost posting

Skills

kubernetesprometheusterraformansiblegrafanadatadogdockerazurehelmcicdgooglecloudaws

About the role

If you’re looking for a career where you can make a real impression, join our Global Service Center (GSC)- HSBC and discover how valued you’ll be.

We are currently seeking an experienced professional to join our team in the role of

Associate Director Engineering - SRE

Role purpose

We are seeking a seasoned Senior Site Reliability Engineer – Manager (SRE) with IT experience to drive reliability, scalability, and performance initiatives across our critical production environments. As a senior member of our team, you will design and implement solutions to ensure the health, availability, and continuous improvement of our infrastructure and services, acting as a technical leader and mentor in SRE methodologies and best practices.

Main activities:

Lead complex troubleshooting and root cause analysis efforts for incidents impacting production, driving rapid resolution and long-term prevention.

Design, architect, and enhance scalable, highly available, and secure infrastructure leveraging cloud, container, and orchestration technologies (e.g., AWS/GCP/Azure, Kubernetes, Docker).

Champion the adoption and refinement of SRE practices - defining and measuring SLIs/SLOs, establishing error budgets, and automating operational processes to minimize toil.

Develop and maintain comprehensive monitoring, logging, and alerting systems using modern observability tools (e.g., Prometheus, Grafana, ELK, Datadog, Splunk).

Drive advancements in deployment automation, CI/CD pipelines, infrastructure-as-code (Terraform, Ansible, Helm, etc.), and configuration management.

Guide, mentor, and coach junior SREs and engineers, fostering a culture of knowledge sharing, reliability, and continuous learning.

Collaborate with software development, QA, product, and operations teams to embed reliability, scalability, and security considerations throughout the software development lifecycle.

Participate in and lead on-call rotations, review and improve incident response processes, and perform blameless postmortems.

Identify, prioritize, and lead large-scale system improvements and engineering projects that enhance reliability and operational efficiency.

Author and maintain thorough documentation, runbooks, and knowledge bases.

Requirements

Lead complex troubleshooting and root cause analysis efforts for incidents impacting production, driving rapid resolution and long-term prevention.

Design, architect, and enhance scalable, highly available, and secure infrastructure leveraging cloud, container, and orchestration technologies (e.g., AWS/GCP/Azure, Kubernetes, Docker).

Champion the adoption and refinement of SRE practices - defining and measuring SLIs/SLOs, establishing error budgets, and automating operational processes to minimize toil.

Develop and maintain comprehensive monitoring, logging, and alerting systems using modern observability tools (e.g., Prometheus, Grafana, ELK, Datadog, Splunk).

Drive advancements in deployment automation, CI/CD pipelines, infrastructure-as-code (Terraform, Ansible, Helm, etc.), and configuration management.

Guide, mentor, and coach junior SREs and engineers, fostering a culture of knowledge sharing, reliability, and continuous learning.

Collaborate with software development, QA, product, and operations teams to embed reliability, scalability, and security considerations throughout the software development lifecycle.

Participate in and lead on-call rotations, review and improve incident response processes, and perform blameless postmortems.

Identify, prioritize, and lead large-scale system improvements and engineering projects that enhance reliability and operational efficiency.

Author and maintain thorough documentation, runbooks, and knowledge bases.

Preferred Qualifications:

Experience with infrastructure-as-code (Terraform, Ansible, or similar tools)

Experience working in 24x7 or high-availability production environments

SRE and cloud certifications (e.g., GCP Professional SRE, AWS DevOps Engineer, CKA/CKAD).

Experience with microservices, distributed systems, and high-throughput architectures.

Experience with AIOps to optimize production operations

You’ll achieve more when you join HSBC!

At HSBC we offer our colleagues a greater number of days so that they can fully enjoy their wedding, take care of the new member of the family, or grieve the loss of a family member. Our paid leave package is at the forefront in Mexico, now you have one more reason to be HSBC and proudly live a culture of well-being, balance and care

Personal data held by the Bank relating to employment applications will be used in accordance with our Privacy Statement, which is available on our website.

Issued by HSBC Electronic Data Processing (México) Private LTD

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation varies by seniority, employer size, and location. When this listing publishes a salary band you'll see it in the badge row above the description.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.