Site Reliability Engineer

RSAWEB

Cape Town, ZAonsitePosted Jul 22, 2026
Posting intelligenceActively listedReposted 2×, possible evergreen/ghost posting

Skills

cloudformationkubernetesprometheusterraformjenkinsansiblegrafanadatadogdockergithubgitlabpulumipythonazurecsscicdgooglecloudawsgo

About the role

Job Information

Date Opened

22/07/2026

Job Type

Full time

Industry

Systems Engineering

Work Experience

5 years

Education Level

Degree/B-Tech

City

Cape Town

Province

Western Cape

Country

South Africa

Postal Code

7405

Job Description

Established in 2001, RSAWEB is South Africa’s fastest growing internet service provider (ISP) with a focus on providing connectivity to home customers, and a wide array of technology solutions to businesses. We are obsessed about ensuring all our customers receive the best possible digital experience and exceptional customer service. Thousands of customers have given RSAWEB a 5-star rating, with an average rating of 4.7 out of 5 on Google – the best-rated ISP in South Africa. We are extremely proud of winning KFM’s Best of the Cape Awards: Best ISP in 2021 and 2022 being one of the fastest streaming ISPs on Netflix and a consistently top-rated ISP on MyBroadband. These accolades are not for nothing, as we constantly strive to improve our products, services, and solutions to enhance each customer’s experience. Having invested heavily in infrastructure, RSAWEB has built a strong presence in South Africa with Data Centres in Johannesburg and Cape Town.

Our Products and Services:

Fibre-to-the-Home (FTTH)

Fibre-to-the-Business (FTTB)

Enterprise connectivity

Mobile connectivity and data management

Cloud infrastructure and more!

At RSAWEB, we are passionate about using our creativity, to provide innovative solutions and services, that allow our customers to succeed in all areas of life. We believe that we are in the business of connecting customers and businesses with each other and a world of infinite possibility and opportunity, through technology. Our mission transcends our values through every customer, every interaction, every connection, every day.

Our values:

We Build Trust and Ownership

We Honour & Respect People

We Cultivate Passion & Creativity

We Innovate Feverishly

We Go the Extra Mile

We Believe in Humility

We Communicate Openly & Honestly

We Make it Fun

We Teach, Grow & Learn

We Do More, With Less

Role Purpose:

The Site Reliability Engineer (SRE) is responsible for ensuring the reliability, performance, scalability, and availability of RSAWEB's platforms, network services, and customer-facing systems. This role blends software engineering, infrastructure automation, and operations to deliver highly reliable services and improve the efficiency of technical teams.

Key Responsibilities

1. Reliability & System Performance

Maintain high availability and performance across platforms, services, and infrastructure.

Define, measure, and improve SLIs/SLOs/SLAs for critical systems.

Troubleshoot system and network reliability issues proactively.

2. Automation & DevOps Enablement

Build automation for deployments, monitoring, configuration, and operational tasks.

Improve CI/CD pipelines and assist engineers with release engineering.

Reduce manual work (toil) by implementing self-service tools and automation workflows.

3. Infrastructure Engineering

Deploy, manage, and optimise cloud and on-prem infrastructure (Linux servers, virtualisation, containers).

Work with network teams to ensure resilient integration between systems and ISP network elements.

Manage and scale containerised platforms (Docker, Kubernetes).

4. Observability & Monitoring

Implement and maintain monitoring, alerting, and logging solutions (e.g., Prometheus, Grafana, ELK, Datadog).

Ensure actionable, low-noise alerting and system dashboards.

Use metrics to identify performance bottlenecks and reliability risks.

5. Incident Management

Participate in incident response, including root cause analysis and corrective actions.

Improve monitoring and automation to prevent repeated issues.

Assist with on-call rotations to support critical services.

6. Security & Compliance

Implement security best practices across systems and deployments.

Support vulnerability scanning, patching, and secure configurations.

Ensure compliance with internal and industry standards (ISO, POPIA, etc).

7. Collaboration & Support

Work closely with Network Engineering, DevOps, Software Development, and NOC teams.

Provide technical guidance in system design, scalability, and reliability improvements.

Improve operational processes through documentation and automation.

Requirements

Minimum Qualifications

Diploma or degree in Computer Science, Engineering, Information Technology, or related field.

Relevant certifications (AWS/Azure/GCP, Linux, Kubernetes, Terraform) are beneficial.

Experience Requirements

3–5+ years in SRE, DevOps, Systems Engineering, or Infrastructure roles.

Experience supporting large-scale, mission-critical environments (preferably ISP or telecom).

Strong background in Linux (CentOS, Ubuntu, Debian) administration.

Experience with container orchestration and Infrastructure as Code.

Technical Skills

Strong scripting skills (Python, Bash, Go preferred).

CI/CD tools: GitHub Actions, GitLab CI, Jenkins, ArgoCD, etc.

IaC: Terraform, Ansible, Pulumi, CloudFormation.

Cloud platforms: AWS / Azure / GCP (or private cloud / OpenStack).

Monitoring: Prometheus, Grafana, Zabbix, ELK, Datadog.

Networking fundamentals: DNS, DHCP, firewalls, load balancing, routing.

Databases: SQL and NoSQL basics.

Knowledge of ISP infrastructure such as BNGs, RADIUS, DNS clusters (advantage).

Benefits

Medical Aid (Discovery)

Reduced Gap Cover Rates (Turnberry Premier)

Retirement Annuity Contribution (Allan Gray)

Medical Insurance (Momentum - Health4Me)

Discounted Internet Connectivity

Free Employee Wellness Programme (Lyra Wellbeing, formerly ICAS)

Exposure to latest industry technologies and standards

Lastly, a work environment that rivals the very best!

If you have not heard from us within 2 weeks of submitting your application, please consider your application as unsuccessful.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for DevOps / SRE roles in South Africa varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our DevOps / SRE hub for South Africa medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.