DevOps / Site Reliability Engineer

Joom

Lisbon, PTonsitePosted Feb 18, 2026
Posting intelligenceMay be filled, listed long agoReposted 96×, possible evergreen/ghost posting

At a glance

Highlights

  • Office-first with flexible hours
  • Remote up to 52 days per year
  • Comprehensive health and wellness benefits
  • Relocation assistance provided
  • Opportunities for promotion and professional training

Why this role might suit you

The position provides influence over a large Kubernetes fleet, opportunities to work with AI‑assisted automation, and a supportive environment that values sleep‑friendly duty rotation, while offering relocation assistance and professional growth across international offices.

Skills

kuberneteslinuxterraformterragruntansiblegopythongcpawsprometheusvictoria-metricsgrafanakeycloakfreeipahashicorp-vaultjfrog-artifactorybgpipsecargo-cdfluxfinops

About the role

Joom Group is an international tech-centric group of e-commerce companies founded in 2016 in Latvia. We are here to transform the largest industry in the world, global trade, making it more transparent, efficient, and technology-driven.

Today, Joom Group brings together the following businesses: Joom, a platform for shopping from all over the world; JoomPro, the first end-to-end cross-border B2B marketplace, with successful operations in Brazil and plans to to other markets; JoomPulse, data platform that provides analytics and recommendations for marketplace sellers; and Onfy, a pharmaceutical marketplace in Germany. Joom Group’s offices are located in China, Brazil, Portugal, Latvia, and Germany, with headquarters in Lisbon, Portugal. We work as one international team, sharing knowledge and collaborating across countries, businesses, and products.

We are looking for an experienced engineer to join our Infrastructure team. The team is responsible for the support and development of the company's key infrastructure components. We run and evolve a fleet of ~20 Kubernetes clusters, where the vast majority of our infrastructure resides.

This role is about impact: the architectural decisions you make and the tools you implement will directly influence the stability and workflow of multiple product teams and business units.

Our Approach

High efficiency: We are a team of 8 engineers managing a huge infrastructure platform. We achieve this through automation and self-service, not manual toil

Reliability: We take stability seriously. We can easily add an "8" to our SLO targets. We want you to help us add a "9"

Code-first: We strictly follow IaC/CasC approaches to maintain consistency and scalability as well as write missing automation in Go or Python

Engineering culture: We maintain a high bar for code quality, yet we dislike bureaucracy. We trust your engineering judgment to balance safety with speed when the situation requires it

Sleep-friendly duty: Our duty rotation acts primarily as an incoming interface during business hours. Night calls are exceptional events (around once in a few months). We value our sleep and aim to keep it that way

AI-curious: We actively explore AI tools to automate routine tasks, speed up troubleshooting and navigate across large codebases. AI assists with drafts and context gathering; engineers review and own every change

Technological Stack

Cloud: GCP (primary), AWS

Core: Kubernetes (foundation of our fleet), Linux (Ubuntu)

IaC / CasC: Terraform, Terragrunt, Ansible

Programming languages: Go, Python

Observability: Prometheus, Victoria Metrics, Grafana

Other services: Keycloak (OIDC/SSO), FreeIPA, HashiCorp Vault, JFrog Artifactory

Requirements

Production experience: solid background in DevOps/Infrastructure roles at scale

Linux proficiency: solid understanding of Linux system internals

Kubernetes expertise: deep understanding of operating Kubernetes at scale

Coding: ability to write automation and tooling in Go or Python

Cloud Knowledge: experience with GCP, AWS or other cloud platforms

Observability: experience with modern monitoring stacks (Prometheus, Victoria Metrics)

Preferred

Experience with networking fundamentals (BGP, IPsec) – rarely needed, but good to know

Experience with GitOps (ArgoCD, Flux)

Experience with FinOps

We offer

Compensation package: base salary and performance-based bonuses

Office-first: flexible hours with a possibility to work remotely 52 days per year

Care & Wellbeing: health insurance (including dental care) for employees and their children, daily meal allowance, and 100% paid sick leave

Relocation Support: full assistance for a smooth relocation for the employees and their family

Team & Growth: collaboration with colleagues across Portugal, Brazil, Latvia and China, with opportunities for promotions, professional trainings, and English & Portuguese courses

Community & Engagement: annual team building activities, knowledge-sharing workshops, and a strong sense of team work

Before applying for the above position please review our Candidate Privacy Notice here. By responding to the vacancy, you acknowledge that you have read our Privacy notice.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for DevOps / SRE roles in Portugal varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our DevOps / SRE hub for Portugal medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.