
Associate Director Engineering - SRE
Skills
About the role
If you’re looking for a career where you can make a real impression, join our Global Service Center (GSC)- HSBC and discover how valued you’ll be.
We are currently seeking an experienced professional to join our team in the role of
Associate Director Engineering - SRE
Role purpose
We are seeking a seasoned Senior Site Reliability Engineer – Manager (SRE) with IT experience to drive reliability, scalability, and performance initiatives across our critical production environments. As a senior member of our team, you will design and implement solutions to ensure the health, availability, and continuous improvement of our infrastructure and services, acting as a technical leader and mentor in SRE methodologies and best practices.
Main activities:
Lead complex troubleshooting and root cause analysis efforts for incidents impacting production, driving rapid resolution and long-term prevention.
Design, architect, and enhance scalable, highly available, and secure infrastructure leveraging cloud, container, and orchestration technologies (e.g., AWS/GCP/Azure, Kubernetes, Docker).
Champion the adoption and refinement of SRE practices - defining and measuring SLIs/SLOs, establishing error budgets, and automating operational processes to minimize toil.
Develop and maintain comprehensive monitoring, logging, and alerting systems using modern observability tools (e.g., Prometheus, Grafana, ELK, Datadog, Splunk).
Drive advancements in deployment automation, CI/CD pipelines, infrastructure-as-code (Terraform, Ansible, Helm, etc.), and configuration management.
Guide, mentor, and coach junior SREs and engineers, fostering a culture of knowledge sharing, reliability, and continuous learning.
Collaborate with software development, QA, product, and operations teams to embed reliability, scalability, and security considerations throughout the software development lifecycle.
Participate in and lead on-call rotations, review and improve incident response processes, and perform blameless postmortems.
Identify, prioritize, and lead large-scale system improvements and engineering projects that enhance reliability and operational efficiency.
Author and maintain thorough documentation, runbooks, and knowledge bases.
Requirements
Lead complex troubleshooting and root cause analysis efforts for incidents impacting production, driving rapid resolution and long-term prevention.
Design, architect, and enhance scalable, highly available, and secure infrastructure leveraging cloud, container, and orchestration technologies (e.g., AWS/GCP/Azure, Kubernetes, Docker).
Champion the adoption and refinement of SRE practices - defining and measuring SLIs/SLOs, establishing error budgets, and automating operational processes to minimize toil.
Develop and maintain comprehensive monitoring, logging, and alerting systems using modern observability tools (e.g., Prometheus, Grafana, ELK, Datadog, Splunk).
Drive advancements in deployment automation, CI/CD pipelines, infrastructure-as-code (Terraform, Ansible, Helm, etc.), and configuration management.
Guide, mentor, and coach junior SREs and engineers, fostering a culture of knowledge sharing, reliability, and continuous learning.
Collaborate with software development, QA, product, and operations teams to embed reliability, scalability, and security considerations throughout the software development lifecycle.
Participate in and lead on-call rotations, review and improve incident response processes, and perform blameless postmortems.
Identify, prioritize, and lead large-scale system improvements and engineering projects that enhance reliability and operational efficiency.
Author and maintain thorough documentation, runbooks, and knowledge bases.
Preferred Qualifications:
Experience with infrastructure-as-code (Terraform, Ansible, or similar tools)
Experience working in 24x7 or high-availability production environments
SRE and cloud certifications (e.g., GCP Professional SRE, AWS DevOps Engineer, CKA/CKAD).
Experience with microservices, distributed systems, and high-throughput architectures.
Experience with AIOps to optimize production operations
You’ll achieve more when you join HSBC!
At HSBC we offer our colleagues a greater number of days so that they can fully enjoy their wedding, take care of the new member of the family, or grieve the loss of a family member. Our paid leave package is at the forefront in Mexico, now you have one more reason to be HSBC and proudly live a culture of well-being, balance and care
Personal data held by the Bank relating to employment applications will be used in accordance with our Privacy Statement, which is available on our website.
Issued by HSBC Electronic Data Processing (México) Private LTD
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.