Senior Site Reliability Engineer

Experian

unknownPosted Jun 19, 2026
Posting intelligenceActively listedReposted 6×, possible evergreen/ghost posting

Skills

prometheusgrafanaawsml

About the role

Company Description

Experian is a global data and technology company, powering opportunities for people and businesses around the world. We help to redefine lending practices, uncover and prevent fraud, simplify healthcare, create marketing solutions, and gain deeper insights into the automotive market, all using our unique combination of data, analytics and software. We also assist millions of people to realize their financial goals and help them save time and money.

We operate across a range of markets, from financial services to healthcare, automotive, agribusiness, insurance, and many more industry segments.

We invest in people and new advanced technologies to unlock the power of data. As a FTSE 100 Index company listed on the London Stock Exchange (EXPN), we have a team of 22,500 people across 32 countries. Our corporate headquarters are in Dublin, Ireland. Learn more at experianplc.com.

Job Description

We are looking for a Site Reliability Engineer to improve the reliability, and performance of business-critical systems. Reporting into our Head of SRE you will focus on AWS cloud infrastructure, DevOps tooling, and core SRE practices within a distributed, production environment.

Main Responsibilities:

Leadership & Strategy

Define and implement SRE best practices across the organization.

Proven expertise in production support, engineering, disaster recovery (DCR), automation, and cloud operations

Mentor and guide a team of SREs, fostering growth.

Collaborate with senior stakeholders to align reliability goals with business objectives.

Reliability & Performance

Establish SLIs, SLOs, and SLAs for critical services and ensure adherence.

Drive initiatives to improve system resilience and reduce operational toil.

Excellent in designing systems that detect and remediate issues without manual intervention – Self Healing systems, Runbook automation

Exposure to tools like Gremlin, Chaos Monkey, AWS FIS to simulate outages and improve fault tolerance

Incident Management

Act as the primary point of escalation for critical production issues and lead major incident response, root cause analysis, and postmortems.

Perform detailed post-incident investigations to identify underlying causes. Document findings and share learnings to prevent recurrence.

Implement preventive measures and continuous improvement processes.

Observability

Champion monitoring, logging, and alerting strategies using tools like Prometheus, Grafana, ELK, and AWS CloudWatch.

Build real-time dashboards to visualize system health and reliability metrics.

Configure intelligent alerting based on anomaly detection and thresholds.

Combine metrics, logs, and traces to enable root cause analysis and reduce Mean Time to Resolution (MTTR).

Knowledge of AIOps or ML-based anomaly detection for proactive reliability management.

Collaboration

Work closely with development teams to integrate reliability into application design and deployment

Promote a culture of shared responsibility for uptime and performance across engineering teams.

Qualifications

Deep expertise with various AWS services. Advanced knowledge of monitoring and observability tools.

Strong leadership capabilities with a focus on setting clear direction, aligning team efforts with organizational goals, and maintaining high levels of motivation and engagement across the team.

Excellent communication skills, with the ability to articulate complex ideas, solutions, and feedback clearly to both technical and non-technical stakeholders. Adept at managing conflict constructively and facilitating consensus.

Proven track record of building secure, mission-critical, high-volume transaction web-based software systems, preferably in regulated environments (finance and insurance industries).

Hands on technologist working in software development including leading an SRE team.

Additional Information

Hybrid working, 2 days a week our Nottingham Office

Great compensation package and discretionary bonus

Core benefits include pension, bupa healthcare, sharesave scheme and more

25 days annual leave with 8 bank holidays and 3 volunteering days. You can purchase additional annual leave.

Our uniqueness is that we celebrate yours. Experian's culture and people are important differentiators. We take our people agenda very seriously and focus on what matters; DEI, work/life balance, development, authenticity, collaboration, wellness, reward & recognition, volunteering... the list goes on. Experian's people first approach is award-winning; World's Best Workplaces™ 2024 (Fortune Top 25), Great Place To Work™ in 24 countries, and Glassdoor Best Places to Work 2024 to name a few. Check out Experian Life on social or our Careers Site to understand why.

Experian Careers - Creating a better tomorrow together

Find out what its like to work for Experian by clicking here

#LI-Hybrid

Experian Careers - Creating a better tomorrow together

Find out what its like to work for Experian by clicking here

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation varies by seniority, employer size, and location. When this listing publishes a salary band you'll see it in the badge row above the description.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.