SRE Engineer

Citi

Chennai, INonsitePosted Jul 7, 2026
Posting intelligenceActively listedReposted 4×, possible evergreen/ghost posting

Skills

ansiblegrafanaoraclemysqlcicdjava

About the role

Discover your future at Citi

Working at Citi is far more than just a job. A career with us means joining a team of more than 230,000 dedicated people from around the globe. At Citi, you’ll have the opportunity to grow your career, give back to your community and make a real impact.

Job Overview

The Apps Support Intmd Analyst is a developing professional role. Deals with most problems independently and has some latitude to solve complex problems. Integrates in-depth specialty area knowledge with a solid understanding of industry standards and practices. Good understanding of how the team and area integrate with others in accomplishing the objectives of the subfunction/ job family. Applies analytical thinking and knowledge of data analysis tools and methodologies. Requires attention to detail when making judgments and recommendations based on the analysis of factual information. Typically deals with variable issues with potentially broader business impact. Applies professional judgment when interpreting data and results. Breaks down information in a systematic and communicable manner. Developed communication and diplomacy skills are required in order to exchange potentially complex/sensitive information. Moderate but direct impact through close contact with the businesses' core activities. Quality and timeliness of service provided will affect the effectiveness of own team and other closely related teams.

Responsibilities:

System Stability and Availability: Ensure the reliability, availability, and scalability of our Java-based applications and databases, running on both on-premise VMs and cloud environments.

Infrastructure Management: Manage and apply Infrastructure as Code (IaC) principles using tools like Ansible to maintain consistent and repeatable environments.

CI/CD and Automation: Oversee and enhance existing CI/CD pipelines to support reliable software delivery. Identify opportunities for automation to reduce toil and improve operational efficiency.

Monitoring and Observability: Manage and enhance comprehensive monitoring, logging, and alerting solutions (e.g., Open Telemetry, Grafana, ELK stack, GCO) for proactive identification, troubleshooting, and resolution of issues.

Incident Management: Lead incident response, conduct blameless post-mortems, and drive the implementation of corrective actions to prevent recurrence and improve system resilience.

Database Reliability: Uphold the reliability and performance of our database systems (e.g., BigData, MySQL, Oracle). Focus on performance tuning, backup and recovery strategies, and disaster recovery preparedness.

Performance Analysis: Proactively identify, analyze, and address performance bottlenecks across the application and infrastructure stack to ensure optimal performance.

Security and Compliance: Apply and enforce security best practices throughout the infrastructure and application lifecycle. Ensure compliance with industry standards and internal policies.

Leveraging AI in Day-to-Day Tasks:

AI-Powered Monitoring: Utilize AI and machine learning models to enhance monitoring capabilities, enabling predictive alerting and anomaly detection to identify potential issues before they impact users.

Automated Root Cause Analysis: Leverage AI-driven tools to accelerate root cause analysis during incidents by analyzing logs, metrics, and traces to pinpoint the source of the problem.

Intelligent Automation: Utilize AI-powered automation to handle complex operational tasks, such as intelligent resource scaling, automated remediation of common issues, and predictive capacity planning.

ChatOps and AI Assistants: Work with AI-powered chatbots and assistants within operational workflows to streamline communication, handle routine queries, and gain real-time insights during incident response.

Collaboration and Leadership:

Cross-Functional Collaboration: Work closely with development, QA, and security teams to foster a culture of reliability and ensure services meet high standards for availability and performance.

Mentorship: Mentor junior SREs and share your expertise to help grow the team's capabilities in system stability and monitoring.

Technical Guidance: Provide expert technical guidance on reliability best practices and contribute to the long-term stability of our infrastructure.

Qualifications:

5-8 years experience

Basic knowledge or interest about apps support procedures, concepts and of other technical areas.

Participation in some process improvements.

Previous experience or interest in standardization of procedures and practices.

Basic Business knowledge/ understanding of financial markets and products.

Knowledge/ experience of problem Management Tools.

Understands of how own sub-function integrates within the function and commercial awareness

Evaluates (sometimes complex) situations using multiple sources of information Developed communication and diplomacy skills to persuade and influence

Good customer service, communication and interpersonal skills

Good knowledge of the business and its technology strategy

Consistently demonstrates clear and concise written and verbal communication skills

Knowledge of issue tracking and reporting using tools

Good all-round team member

Effectively share information with other support team members and with other technology teams

Ability to plan and organize workload

Ability to communicate appropriately to relevant stakeholder

Education:

Bachelor’s/University degree or equivalent experience

-

Job Family Group:

Technology

-

Job Family:

Applications Support

-

Time Type:

Full time

-

Most Relevant Skills

Please see the requirements listed above.

-

Other Relevant Skills

For complementary skills, please see above and/or contact the recruiter.

-

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for DevOps / SRE roles in India varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our DevOps / SRE hub for India medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.