Senior Site Reliability Engineer (SRE) - Core Messaging Infrastructure - STACKIT (m/f/d)

Schwarz Digits KG

unknownPosted Jul 9, 2026
Posting intelligenceActively listedReposted 16×, possible evergreen/ghost posting

Skills

kubernetesansiblepythonkafkago

About the role

Location

Am Campus 1

74177 Bad Friedrichshall

Employment Area

IT - Cloud Services

Level

Experienced professional

Working Model

Full-time

Reference ID

3944

Schwarz Digits creates the technological foundation for digital sovereignty in Europe. As the IT and digital division of the Schwarz Group, we develop and manage the IT infrastructures for the retail divisions Lidl and Kaufland, as well as Schwarz Production and PreZero. At the same time, we operate as an independent provider in the external market to support companies across Europe in their digital transformation. We bundle our core services in the areas of Cloud, Cyber Security, Data & AI, Communication, and Workspace.

Join us and contribute to digital sovereignty in Europe. With us, you will work at the intersection of agility and security: You will benefit from fast decision-making processes, enjoy genuine creative freedom in your projects, and be able to build upon the stable foundation of the Schwarz Group.

We are looking for a Senior Engineer to build, scale, and own the central nervous system of our cloud infrastructure: a highly resilient, high-throughput message and event platform. As our engineering organization scales rapidly, we are transitioning to a real-time, event-driven architecture to ensure seamless communication between the control plane components of all our products. You will empower dozens of product teams by providing an outstanding developer experience, enabling them to seamlessly publish and consume millions of events per day.

Your Tasks

You design, deploy, and manage highly available, distributed message broker clusters (such as Apache Kafka, Solace, or NATS) across multiple data centers.

You ensure the reliability, performance, and fault tolerance of the messaging infrastructure by implementing robust disaster recovery and failover strategies and tune operating system configurations for low-latency delivery.

You automate the provisioning, scaling, and configuration of messaging clusters.

You build comprehensive monitoring, alerting, and logging dashboards to track cluster health, throughput, and latency.

You define best practices for application developers and build a self-service platform that makes it easy for internal teams to independently configure their integrations.

Your Profile

You bring solid experience in managing large-scale distributed systems in production, coming from a background like Site Reliability Engineering or Platform Engineering.

You have deep, hands-on administrative experience with enterprise brokers like Apache Kafka or Solace, and you bring experience in managing infrastructure on Kubernetes using the Operator Pattern as well as managing virtual machines using tools like Ansible.

You bring fluency in coding with Python, Go, or Bash, alongside a strong understanding of Linux performance tuning and networking protocols (such as Transmission Control Protocol/Internet Protocol or Domain Name System).

Ideally, you have a profound understanding of event-driven architecture patterns and event streaming concepts to help design scalable, real-time data pipelines.

Your English, ideally combined with German, is the basis for successful communication in our international, agile teams.

Our benefits

Your contact

Heiko Kiefer

Recruiter

E-Mail:recruiting@mail.schwarz

What happens once you applied?

1

Application

We love simplicity. Apply quickly and easily digitally, even without registration.

2

Checkup

Together with the department, we take a look at your documents.

3

Interview

When you first get to know each other, the focus is on both your personality and your professional suitability.

4

Contract offer

Have we convinced each other? Then your employment contract is on its way to you digitally by e-mail.

5

Welcome

Your first day of work with is about to begin and your individual onboarding in the team begins. We look forward to seeing you!

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation varies by seniority, employer size, and location. When this listing publishes a salary band you'll see it in the badge row above the description.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.