Site Reliability Engineer / SRE - Cloud Storage - STACKIT (m/f/d)

Schwarz Digits KG

unknownPosted Jul 9, 2026
Posting intelligenceActively listedReposted 16×, possible evergreen/ghost posting

Skills

elasticsearchprometheusansiblegrafanagopythonkubernetes

About the role

Location

Am Campus 1

74177 Bad Friedrichshall

Employment Area

IT - Cloud Services

Level

Experienced professional

Working Model

Full-time

Reference ID

705

Schwarz Digits creates the technological foundation for digital sovereignty in Europe. As the IT and digital division of the Schwarz Group, we develop and manage the IT infrastructures for the retail divisions Lidl and Kaufland, as well as Schwarz Production and PreZero. At the same time, we operate as an independent provider in the external market to support companies across Europe in their digital transformation. We bundle our core services in the areas of Cloud, Cyber Security, Data & AI, Communication, and Workspace.

Join us and contribute to digital sovereignty in Europe. With us, you will work at the intersection of agility and security: You will benefit from fast decision-making processes, enjoy genuine creative freedom in your projects, and be able to build upon the stable foundation of the Schwarz Group.

Your tasks

Stability & Reliability: You are responsible for maintaining and optimizing the stability and availability of our highly available, resilient storage infrastructure (block, object, backup and file storage). You ensure this through proactive monitoring, solving occurring faults on your own responsibility and avoiding their occurrence in the future

Automation: You automate the provisioning and operating processes in the storage environment with your own aspiration to become a little better every day and to continuously optimize our products

Architecture: With your team, you are responsible for a robust and efficient storage architecture - because it is important to you to build a long-term stable and reliable solution that our customers will be happy to use

End-to-end responsibility: Identifying with the products we provide to our customers is very important to us. Therefore, we actively live an end-to-end responsibility and receive support from many internal STACKIT service teams to refine our services

Performance and capacity planning: You will analyze and optimize the performance of our existing systems with regard to future scaling of the landscape. This also includes forward-looking capacity planning

Incident and post-mortem analysis: You are responsible for processing (major) incidents with storage participation as part of the incident & problem management process of STACKIT with the aim of deriving mitigating measures for the future and then successfully implementing them

Your Profile

You want to make a difference and play a significant role in shaping the solution with state-of-the-art cloud technologies

You have extensive experience in the market environment with various storage products (e.g. NetApp, Cohesity, Pure, Ceph) in the area of block, object, backup or file storage and have good knowledge of cloud environments and their architectures

You are an expert in the operation of storage infrastructure (e.g. solution scenarios, provision, scaling, migration, incident response) and their automation (e.g. using Golang / Python, Bash, Ansible)

You are familiar with containerized system landscapes of the storage environment (e.g. k8s)

You have experience in monitoring, alerting and logging to ensure complete system monitoring (e.g. Prometheus, Grafana, Elasticsearch)

You are already working with APIs and are developing them further (e.g. REST API with Golang and Python)

You enjoy the challenges of operating storage systems (e.g. protocols, troubleshooting, performance analyzes, high availability, lifecycle)

You bring with you passion and enthusiasm for new technologies and topics related to various storage systems

You like to be part of a motivated team that always strives for improvement and continuously develops itself (and the products)

Your excellent communication skills in German and English form the basis for successful cooperation in international, agile teams

Our benefits

Your contact

Heiko Kiefer

Recruiter

E-Mail:recruiting@mail.schwarz

What happens once you applied?

1

Application

We love simplicity. Apply quickly and easily digitally, even without registration.

2

Checkup

Together with the department, we take a look at your documents.

3

Interview

When you first get to know each other, the focus is on both your personality and your professional suitability.

4

Contract offer

Have we convinced each other? Then your employment contract is on its way to you digitally by e-mail.

5

Welcome

Your first day of work with is about to begin and your individual onboarding in the team begins. We look forward to seeing you!

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation varies by seniority, employer size, and location. When this listing publishes a salary band you'll see it in the badge row above the description.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.