Principal Site Reliability Engineer – Performance (A&D, Ultra HA/Exadata)

IFS

Ottawa, CAonsitePosted Aug 24, 2026
Posting intelligenceActively listed

Skills

cloudformationelasticsearchkubernetesprometheusregressionterraformansiblegrafanadatadogdockeroraclepythonazurekafkajiracicdgooglecloudawsgo

About the role

Company Description

IFS is a billion-dollar revenue company with 7000+ employees on all continents. Our leading AI technology is the backbone of our award-winning enterprise software solutions, enabling our customers to be their best when it really matters–at the Moment of Service™. Our commitment to internal AI adoption has allowed us to stay at the forefront of technological advancements, ensuring our colleagues can unlock their creativity and productivity, and our solutions are always cutting-edge.

At IFS, we’re flexible, we’re innovative, and we’re focused not only on how we can engage with our customers but on how we can make a real change and have a worldwide impact. We help solve some of society’s greatest challenges, fostering a better future through our agility, collaboration, and trust.

We celebrate diversity and understand our responsibility to reflect on the diverse world we work in. We are committed to promoting an inclusive workforce that fully represents the many different cultures, backgrounds, and viewpoints of our customers, our partners, and our communities. As a truly international company serving people from around the globe, we realize that our success is tantamount to the respect we have for those different points of view.

By joining our team, you will have the opportunity to be part of a global, diverse environment; you will be joining a winning team with a commitment to sustainability; and a company where we get things done so that you can make a positive impact on the world.

We’re looking for innovative and original thinkers to work in an environment where you can #MakeYourMoment so that we can help others make theirs. With the power of our AI-driven solutions, we empower our team to change the status quo and make a real difference.

If you want to change the status quo, we’ll help you make your moment. Join Team Purple. Join IFS.

Job Description

Job Description

The role of IFS Principal Site Reliability Engineer – Performance Engineering (Principal-SRE) exists within the Unified Support organization/division and serves as a technical and strategic leader for Site Reliability Engineering practices. By being part of the shift operation and broader SRE leadership, a Principal-SRE will drive 24x7x365 support excellence to IFS customers across the globe within the Aerospace & Defense (A&D) industry vertical while establishing and elevating SRE practices across the organization. This role sits within a dedicated performance engineering team supporting our Ultra High Availability (Ultra HA) on Exadata customers, where sustained database and application performance is a contractual commitment rather than a best-effort goal.

As Principal SRE – Performance Engineering, you will be responsible for architecting and evolving the reliability frameworks, automation platforms, and operational strategies that underpin IFS's cloud-native products and infrastructure. This role involves setting technical direction, leading cross-functional teams, and partnering with R&D, Product, and Operations leadership to ensure that reliability is embedded at every level of service delivery. You will be an influential technical leader who drives organizational transformation through SRE principles, automation, and continuous improvement culture.

Strategic Impact Areas

Reliability Architecture & Strategy: Define and evolve the SRE operating model for the A&D vertical, including incident management, capacity planning, and disaster recovery strategies. Architect resilience patterns and frameworks for cloud-based applications supporting critical Aerospace & Defense operations. Establish Service Level Objectives (SLOs) and Error Budgets aligned with customer contractual requirements and business objectives. Drive adoption of chaos engineering and resilience testing across production systems.

Technical Leadership & Platform Engineering: Lead the design and implementation of enterprise-scale monitoring, observability, and automated remediation platforms. Establish and maintain infrastructure-as-code standards and CI/CD pipelines for reliable deployments. Champion modern SRE tooling and methodologies across Unified Support and R&D organizations. Define technical architecture for Kubernetes, multi-cloud orchestration, and cloud-native application patterns. Architect solutions for performance optimization, cost efficiency, and security at scale (3,200+ concurrent users, 6,048+ database connections).

Ultra HA & Exadata Performance Engineering: Own the performance and availability posture of Ultra HA environments running on Oracle Exadata, including proactive database health monitoring, workload and wait-event analysis, and end-to-end application performance management. Establish deep observability across the Oracle stack and the application tiers it serves, using Elastic, Grafana, and OpenTelemetry-based instrumentation to correlate database behaviour with user-facing latency. Lead performance troubleshooting for the most demanding customer workloads and drive tuning, capacity, and remediation decisions ahead of SLA impact.

Organizational Transformation & Culture: Drive automation-first culture, eliminating toil and enabling teams to focus on high-value, creative work. Lead post-incident reviews and establish blameless culture focused on systemic improvement. Mentor and develop SRE team members, establishing career progression paths and technical mastery standards. Establish communities of practice for knowledge sharing across Unified Support and R&D organizations.

Key Responsibilities

Strategic Leadership & Architecture

Set technical direction and establish long-term reliability strategies for IFS cloud products supporting the A&D industry

Partner with R&D, Product, and Customer Success leadership to define reliability requirements and drive architectural decisions

Design and oversee implementation of enterprise-scale observability, automation, and incident response platforms

Lead architectural assessments and provide recommendations on infrastructure resilience, capacity planning, and disaster recovery

Drive adoption of SRE best practices across product teams, establishing metrics and accountability for reliability

Performance Engineering & Ultra HA Operations

Own proactive performance monitoring for Ultra HA on Exadata customers - database health, workload profiles, wait events, and application response times - and act on trends before they breach SLA

Lead deep performance investigations spanning the Oracle database, application servers, and infrastructure, producing evidence-based tuning and capacity recommendations

Design and maintain the performance observability stack (Elastic, Grafana, OpenTelemetry/APM), including dashboards, baselines, and alert thresholds tied to customer SLOs

Run performance baselining, load testing, and pre-release validation for Ultra HA environments, and quantify the impact of changes before they reach production

Produce customer-facing performance reporting and participate in service reviews for Ultra HA accounts

Incident Management & Operational Excellence

Lead resolution of long-running, complex, and critical incidents affecting A&D customers; establish escalation frameworks and decision-making protocols

Establish incident command systems, on-call rotations, and escalation procedures that balance response speed with team sustainability

Drive post-incident review processes that focus on identifying systemic improvements rather than individual blame

Establish incident trends analysis and drive prevention of recurrence through root cause elimination

Automation & Continuous Improvement

Champion the elimination of toil through intelligent automation, leveraging AI to reduce redundant, admin-heavy tasks

Design and implement self-healing systems, automated remediation, and predictive alerting to minimize manual intervention

Establish continuous improvement programs that measure, track, and drive down Mean Time To Detect (MTTD) and Mean Time To Recovery (MTTR)

Drive infrastructure-as-code adoption, standardization, and versioning across cloud platforms (Azure, AWS, GCP)

Establish runbook automation and GitOps practices for reliable, auditable deployments

Documentation & Knowledge Management

Establish knowledge management frameworks and internal KBAs that guide support operations for cloud-based applications

Create and maintain architecture documentation, disaster recovery playbooks, and operational runbooks

Lead documentation standards that ensure knowledge is accessible, accurate, and actionable for all support tiers

Identify systemic knowledge gaps and drive closure through training, documentation, and process improvement

Cross-Functional Collaboration

Liaise with R&D and Unified Support Engineering to define automation requirements and platform capabilities for cloud applications

Partner with customer success and professional services teams to translate customer reliability requirements into technical strategies

Lead working groups and architecture reviews with multiple stakeholders to ensure reliability-first design decisions

Establish feedback loops between support operations and product development to drive product reliability improvements

Team Development & Mentorship

Mentor and develop SRE team members, establishing technical mastery standards and career progression frameworks

Lead technical training programs on cloud operations, Kubernetes, performance engineering, and SRE methodologies

Foster a culture of learning, experimentation, and continuous improvement within the SRE organization

Establish on-call culture that balances operational needs with team well-being and sustainable practices

Qualifications

Required Qualifications

Education & Certifications (Core Requirements)

University degree or equivalent professional qualification in Software Engineering, Computer Science, Information Technology, or similar discipline

Demonstrated expertise in modern ticket/service desk tooling (ServiceNow, Jira Service Desk, or equivalent)

ITIL, ISO 20000, or equivalent IT service delivery framework certification; or demonstrated mastery through practical application

Experience (Mandatory)

Minimum 10+ years of progressive experience in cloud computing services, enterprise IT delivery, or site reliability engineering

Minimum 7-8 years in hands-on SRE, DevOps, or cloud operations roles with demonstrated impact on reliability, automation, and operational efficiency

Proven experience leading SRE teams and establishing SRE practices across organizations

Deep understanding of low-level concepts in cloud computing (resource management, networking, storage, security)

Demonstrated expertise in performance engineering including load testing, capacity planning, and system optimization

Experience supporting mission-critical, high-availability systems (preferably in Aerospace, Defense, Finance, or similar industries)

Hands-on experience operating or supporting Oracle Exadata platforms in production, ideally under Ultra High Availability or comparable premium-SLA commitments

Demonstrated experience in a dedicated performance engineering or performance-analysis function, owning latency, throughput, and capacity outcomes for named enterprise customers

Track record of designing and implementing large-scale automation platforms that reduce toil and improve MTTR

Demonstrated ability to establish and drive organizational change around reliability and operational excellence

Required Technical Skills

The successful candidate will demonstrate mastery in most of the following areas:

Cloud Infrastructure & Operations

Expert-level cloud service administration and operations (Azure, AWS, GCP)

Kubernetes/Docker architecture, operations, and troubleshooting at production scale

Network administration and troubleshooting (DNS, Load Balancing, VPN, firewall rules)

Enterprise storage solutions, block storage, and object storage optimization

Oracle Database Monitoring, Administration & Performance

Expert-level Oracle database administration and performance tuning

Expert-level Oracle database monitoring and diagnostics - AWR, ASH, ADDM, Statspack, SQL trace/TKPROF, and Oracle Enterprise Manager (OEM) - with the ability to move from symptom to root cause under production pressure

Advanced Oracle performance troubleshooting: wait-event and contention analysis, execution-plan regression, SQL and PL/SQL tuning, optimizer statistics, partitioning, and index strategy

Oracle high-availability architectures for Ultra HA workloads - RAC, Data Guard/Active Data Guard, ASM, and RMAN - including failover testing and recovery validation

Understanding of connection pooling, query optimization, and capacity planning for high-concurrency systems (3,000+ connections)

Experience with database backup, recovery, and disaster recovery strategies

Hands-on Oracle Exadata administration and performance management (Smart Scan, storage indexes, IORM, flash cache, cell-level diagnostics) - required for Ultra HA customer environments

Application & Web Server Administration

Web server administration and troubleshooting (Wildfly, WebLogic, Nginx)

Application performance monitoring (APM) and optimization across the full request path - instrumentation, transaction tracing, and latency breakdown from web tier through middleware to the database

SSL/TLS certificate management and security best practices

Infrastructure-as-Code & Automation

Advanced proficiency in BASH, PowerShell, Python, or Go scripting

Infrastructure-as-Code expertise (Terraform, Ansible, CloudFormation)

CI/CD pipeline design and implementation

GitOps and declarative infrastructure practices

Observability, Monitoring & APM

Design and implementation of enterprise-scale monitoring solutions

Hands-on experience with Elasticsearch and the Elastic Stack (Elasticsearch, Logstash/Beats, Kibana) for log aggregation, search, and analysis at scale; equivalent experience with Splunk or DataDog also considered

Grafana dashboard design and metrics visualization, including building performance dashboards that surface Oracle, infrastructure, and application metrics side by side (Prometheus or equivalent metrics backends)

OpenTelemetry-based instrumentation and distributed tracing, or equivalent APM tooling (Dynatrace, AppDynamics, New Relic, Elastic APM)

Ability to correlate database, infrastructure, and application telemetry into a single performance narrative, and to build proactive alerting on leading indicators of degradation rather than outright failure

Debugging & Troubleshooting

Expert ability to debug complex, multi-tier applications

Deep understanding of application server internals and JVM diagnostics

Network packet analysis and protocol debugging

Linux system troubleshooting and performance analysis

Additional Information

Soft Skills

Strategic thinking and ability to translate business objectives into technical strategies

Executive communication skills with ability to present complex technical concepts to non-technical stakeholders

Team leadership and mentorship abilities with proven track record of developing technical talent

Change management and organizational influence skills

Problem-solving with ability to innovate new approaches to complex challenges

Decision-making under pressure with ability to weigh trade-offs between reliability, cost, and feature delivery

Emotional intelligence and ability to build trust across diverse, international teams

Proactive ownership of work items and willingness to challenge status quo for improvement

Self-learning and ability to stay current with rapidly evolving cloud technologies and SRE practices

Excellent communication in English (verbal and written) for cross-functional collaboration

Desirable Qualifications & Nice-to-Have Skills

Advanced cloud security best practices (Zero Trust, Identity & Access Management, encryption)

Extensive experience supporting large Aerospace & Defense customers and understanding of industry-specific compliance requirements

Performance engineering mastery including advanced load testing, capacity planning, and cost optimization strategies

Experience with Disaster Recovery (DR) planning and execution, including multi-region failover and RTO/RPO optimization

Expertise in cost optimization and FinOps practices for cloud infrastructure

Hands-on experience with event streaming platforms (Kafka, RabbitMQ) and asynchronous architecture patterns

API design and management experience

Machine Learning/AI applications for observability, anomaly detection, or predictive remediation

Certifications such as AWS Solutions Architect Professional, Azure Solutions Architect Expert, or CKA (Certified Kubernetes Administrator)

Oracle certifications (OCP Database Administrator, Exadata Database Machine Certified Implementation Specialist) or Elastic/Grafana certifications

Experience contributing to OpenTelemetry, Elastic, or Grafana ecosystems, or building custom exporters and instrumentation libraries

Published articles, conference talks, or open-source contributions in SRE or DevOps domains

Contribution to SRE community through mentorship, speaking, or knowledge sharing

Role Scope & Impact

Reporting Structure

Reports to: Senior Leadership (Engineering Manager, Technical Director, or VP Engineering)

Team Leadership: Manages SRE team of 3-5 engineers; influences broader organization

Scope: Responsible for reliability and performance strategy across the A&D vertical, touching 6,000+ customer deployments, with direct ownership of performance outcomes for Ultra HA on Exadata accounts

Key Performance Indicators

System reliability metrics (uptime %, SLA attainment, error rates)

Performance metrics for Ultra HA accounts (application response time, database wait time, throughput at peak, capacity headroom)

Observability coverage and alert quality (instrumented services, false-positive rate, share of issues detected before customer impact)

Operational efficiency (MTTR reduction, incident frequency, automation coverage)

Team development (mentee growth, retention, skill advancement)

Business impact (customer satisfaction, churn reduction, revenue protection)

Questions about this role

Click "Apply with AI Applyd" above and you are done. Your resume is rewritten for this advert, the screening questions are answered, and it is submitted on IFS's own hiring system. No retyping your history, no fourteen tabs, no evening lost.

Compensation for DevOps / SRE roles in Canada varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our DevOps / SRE hub for Canada medians across recent openings.

You never touch the form - the application is filled and submitted for you on IFS's own hiring system. It is not marked sent when we press submit. It is marked sent when a confirmation from their system arrives at the address we apply with, and your dashboard shows which stage each application is at until then.

Twelve applicant tracking systems have a real apply path: Workday, Greenhouse, Lever, Ashby, Workable, iCIMS, Personio, Recruitee, Teamtailor, Rippling, Breezy and SmartRecruiters. Your application goes in on the employer's own hiring system, never into an aggregator queue.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.