Senior Site Reliability Engineer (Networking)

Oracle

Zapopan, MXunknownPosted Jul 27, 2026
Posting intelligenceActively listedReposted 3×, possible evergreen/ghost posting

Skills

prometheusterraformansiblegrafanaoraclepythonazuregooglecloudaws

About the role

We are hiring a Senior Site Reliability Engineer (Networking) to join OCI Corporate Network and Security Operations. This role is responsible for improving the reliability, scalability, performance, and operational excellence of Oracle's global corporate network infrastructure. The engineer will collaborate with Network Engineering, Security, Software Development, and Cloud Operations teams to automate operations, improve service resilience, resolve complex production incidents, and ensure highly available network services supporting Oracle's enterprise environment.

The successful candidate combines Site Reliability Engineering principles with strong expertise in enterprise networking, automation, cloud infrastructure, and operational excellence. This role requires a proactive mindset focused on improving service reliability through automation, observability, capacity planning, and continuous improvement while supporting mission-critical production environments.

Key Responsibilities

Capacity Ingestion and Management

Takes proactive steps to design and architect infrastructure and/or services according to reliability, availability, scalability, and performance requirements.

Designs and improves highly available enterprise network infrastructure supporting Oracle Corporate Network and Security Operations.

Forecasts infrastructure and network capacity requirements and responds to capacity needs to ensure systems can support current and future workloads.

Collaborates with Software Development, Network Engineering, and Security teams to develop reliable, scalable, and resilient infrastructure and services.

Independently identifies opportunities for innovation and drives prototyping initiatives, including onboarding new infrastructure, services, and automation capabilities.

Incident and Service Lifecycle Management

Performs data collection, triage, technical analysis, and issue resolution to maintain and optimize infrastructure and network reliability.

Independently monitors production services and enterprise network infrastructure, maintaining up-to-date knowledge of service health and performance.

Supports Corporate Network and Security Incident, Change, Capacity, and Problem Management processes.

Leverages comprehensive knowledge to perform incident response, root cause analysis (RCA), maintenance activities, software upgrades, security updates, backup, and recovery.

Performs deep troubleshooting across enterprise and cloud network environments, including Layer 2/Layer 3 technologies, routing, switching, DNS, VPN, load balancing, and firewall platforms.

Provides health and performance reporting while proactively identifying trends and opportunities to improve operational reliability.

Performs provisioning activities supporting infrastructure, applications, and network services.

Performs standard and non-standard decommissioning activities when required.

Drives corrective and preventive actions following production incidents to eliminate recurring operational issues.

Automation

Identifies opportunities for automation and evaluates operational benefits.

Develops automation solutions using Python, Ansible, Infrastructure as Code, and related technologies to improve operational efficiency and network reliability.

Develops automation tools and scripts to gather metrics, monitor services, analyze system behavior, mitigate operational risks, and remediate production issues.

Supports Zero Touch Provisioning and automated infrastructure lifecycle management where applicable.

Independently validates automation solutions through testing to ensure expected functionality and operational readiness.

Technical Communication and Guidance

Communicates the scale, capacity, security, performance characteristics, and operational requirements of services and enterprise network infrastructure.

Identifies and explains the operational impact of infrastructure, tooling, and service changes before production implementation.

Partners with Network Engineering, Security, Cloud Operations, and Software Engineering teams during production changes, maintenance activities, and operational improvements.

Effectively communicates operational risks, capacity constraints, and reliability improvements to technical stakeholders.

Troubleshooting and Resolution

Provides operational support for Oracle technology services, escalating incidents and complex operational issues as appropriate.

Resolves technical issues spanning multiple infrastructure and network services while maintaining established Service Level Objectives (SLOs).

Troubleshoots complex production issues involving enterprise routing, switching, firewalls, VPN technologies, DNS, load balancers, Linux systems, cloud networking, and distributed infrastructure.

Documents incidents, performs comprehensive Root Cause Analysis (RCA), and independently executes post-incident reviews to prevent recurrence.

Participates in a 24x7 on-call rotation supporting mission-critical enterprise network services.

Innovation and Continuous Improvement

Evaluates emerging networking technologies, cloud capabilities, observability platforms, and automation solutions to improve operational efficiency.

Independently identifies and implements improvements addressing performance bottlenecks, scalability challenges, and infrastructure optimization opportunities.

Continuously improves monitoring, observability, automation, and operational tooling supporting enterprise network environments.

Maintains knowledge of Site Reliability Engineering and enterprise networking trends, sharing knowledge and best practices across the organization.

Performs standard and non-standard production analysis to support operational excellence and informed business decisions.

Required Qualifications

Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.

6–10 years of experience supporting enterprise or cloud network environments and mission-critical production infrastructure.

Strong knowledge of enterprise networking technologies, including TCP/IP, Layer 2/Layer 3 switching, routing, firewalls, VPNs, load balancers, DNS/DHCP, and enterprise network operations.

Hands-on experience with Incident, Change, Capacity, and Problem Management, Linux administration, Python, Ansible, infrastructure automation, cloud networking (OCI preferred; AWS, Azure, or GCP), and monitoring and observability platforms.

Strong analytical, troubleshooting, and problem-solving skills, with the ability to collaborate effectively across cross-functional engineering teams.

Excellent verbal and written communication skills and willingness to participate in a 24x7 on-call support rotation.

Preferred Qualifications

Experience with modern networking, cloud, and automation technologies such as OCI Networking, Terraform, Git, Infrastructure as Code (IaC), Zero Touch Provisioning, Grafana, Prometheus, Splunk, ELK, Cisco Nexus, Juniper, Palo Alto Networks, F5 Load Balancers, BGP, and OSPF is highly desirable. Experience designing and implementing automation solutions to improve operational efficiency, scalability, and network reliability in large-scale enterprise or cloud environments is a plus.

Core Responsibilities

Planning & Execution

Independently manages work, monitoring timelines and deliverables to ensure projects or initiatives stay on track and meet requirements. Proactively prioritizes work and adapts to resource or timeline shifts, suggesting adjustments to maintain project efficiency.

Collaboration & Partnership

Collaborates across teams to align on expectations and achieve shared objectives. Builds and maintains a comprehensive understanding of business, stakeholder, and/or customer needs to build and support effective partnerships. Actively listens to diverse perspectives and asks questions to ensure understanding of others.

Problem Solving

Independently identifies and addresses standard and non-standard issues in accordance with standard practices, escalating more complex issues as appropriate. Analyzes data and/or information from multiple sources to troubleshoot standard and non-standard errors. Contributes to knowledge sharing and best practices.

Continuous Learning

Embraces continuous learning by actively seeking to build knowledge and new skills and/or tools and staying current with industry trends and best practices. Seeks out and leverages feedback and training to improve skills. Contributes to a culture of continuous learning and knowledge sharing with team members.

Continuous Improvement

Develops ideas and recommends updates to increase the efficiency and effectiveness of processes, protocols, and workflows within a team. Seeks input from team members on alternative approaches and methods for improving work.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for DevOps / SRE roles in Mexico varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our DevOps / SRE hub for Mexico medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.