KG

Systemadministrator (m/w/d)

KYAANI GmbH

Berlin, DEonsitePosted Aug 4, 2026
Posting intelligenceActively listed

Skills

kubernetesprometheusgrafanadockerpythonopenaillm

About the role

About the Role We are building a highly secure, on-premise Large Language Model (LLM) infrastructure that operates entirely offline. We are looking for a AI Infrastructure Engineer who operates at the intersection of bare-metal hardware management and MLOps. In this role, you will be responsible for building, optimizing, and maintaining the physical and software architecture required to serve massive AI models securely and efficiently without relying on external cloud APIs.

Key Responsibilities

Bare-Metal GPU Management: Install, configure, and troubleshoot Linux servers, Nvidia GPU drivers, CUDA toolkits, and cuDNN versions to ensure maximum hardware utilization.

Local Inference Deployment: Deploy and manage high-performance local inference servers such as vLLM, Hugging Face TGI (Text Generation Inference), or TensorRT-LLM.

Model Optimization: Manage VRAM efficiently through model quantization (GGUF, AWQ, EXL2) and calculate hardware requirements for various model sizes and batch loads.

Containerization & Orchestration: Package model weights and inference engines into Docker containers and orchestrate deployments using Kubernetes (K8s).

API & Integration: Set up and maintain local reverse proxies (e.g., LiteLLM) to ensure the offline models provide a standard, OpenAI-compatible API for internal developers.

Security & Monitoring: Architect and maintain air-gapped network environments with zero outbound internet access. Implement robust monitoring using Prometheus and Grafana to track GPU temperatures, power limits, and VRAM usage.

Must-Have Qualifications

Extensive experience as a Linux System Administrator managing bare-metal servers.

Deep, hands-on understanding of Nvidia GPU architecture, multi-GPU orchestration (NVLink, NCCL), and resolving driver/CUDA conflicts.

Proven experience deploying open-source LLMs (Llama 3, Mistral, Mixtral) locally using frameworks like vLLM or Ollama.

Strong proficiency in containerization (Docker) and orchestration (Kubernetes).

Solid scripting skills in Python and Bash for automation and MLOps pipelines.

Strong understanding of network security, specifically designing and maintaining air-gapped systems.

Nice-to-Have

Experience with custom model fine-tuning infrastructure.

Familiarity with building retrieval-augmented generation (RAG) pipelines on local hardware.

Location - Onsite Start Date - Immediately Type - Mini job (10 hours a week)

Job Type: Part-time

Pay: Up to 13,90€ per hour

Work Location: In person

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for Platform Engineer roles in Germany varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our Platform Engineer hub for Germany medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.