Software Engineer SRE (Site Reliability Engineer)
Skills
About the role
Overview
At NetApp, we have a history of helping customers turn challenges into business opportunities. That’s because we bring new thinking to age-old problems, like how to use data most effectively in the most efficient possible way. As an Engineer with NetApp, you’ll have the opportunity to work with modern cloud and container orchestration technologies in a production setting. You’ll play an important role in scaling systems sustainably through automation and evolving them by pushing for changes to improve reliability and velocity.
Own Every Moment at NetApp
At NetApp, your ideas power innovation. We lead in intelligent data infrastructure - delivering unified storage, integrated data services, and solutions that help organizations unlock the full potential of their data, from AI to multicloud.
Ready to innovate and contribute to our path to $10B? Here, you'll collaborate with passionate teams, tackle real-world challenges, and see your impact in how customers transform and grow. If you're ready to bring curiosity, creativity, and drive to every moment, NetApp is where your journey begins.
Job Summary
We're looking for a Site Reliability Engineer, focused on building and operating the data and AI/ML infrastructure platform that powers NetApp's cloud-native data services. You'll work at the intersection of software engineering and infrastructure operations - designing systems for reliability, driving automation, and ensuring our platforms meet the highest availability standards for customers worldwide.
This is an infrastructure-focused SRE role, you'll own the reliability of large-scale Kubernetes clusters (including GPU workloads), streaming data pipelines (Kafka), and analytical compute infrastructure (Spark, Dremio) across hybrid-cloud and multi-cloud environments.
Job Requirements
5+ years in SRE, DevOps, Platform Engineering, or Infrastructure Engineering roles
Extensive experience with Linux (RHEL/CentOS), including shells, filesystems, kernel tuning, networking, and performance optimization
Deep expertise with Kubernetes at scale, including cluster administration, troubleshooting, networking, storage, RBAC, and lifecycle management (on-premises and Rancher Kubernetes)
Hands-on experience operating GPU workloads on Kubernetes, including NVIDIA GPU Operator, device plugins, scheduling, and resource management
Strong experience managing Confluent Kafka in production, including operations, monitoring, performance tuning, and disaster recovery
Experience operating Apache Spark and/or Dremio, including cluster management, job scheduling, scaling, and performance optimization
Proficiency in Infrastructure as Code using Terraform, Helm, and GitOps workflows with ArgoCD/FluxCD
Proficiency in scripting and automation using Shell, Ansible, and Python, with a strong automation-first mindset
Experience with scheduling and orchestration tools such as cron jobs and Apache Airflow
Deep familiarity with monitoring and observability tools, including Dynatrace, Grafana, and Prometheus
Solid understanding of SQL and NoSQL databases, including operations, backup, and monitoring
Experience designing and maintaining CI/CD pipelines and release processes
Expertise in AWS cloud platforms and hybrid-cloud integration
Strong systems thinking, with an understanding of how infrastructure design choices impact failure modes, scalability, and recovery
Strong incident management skills and post-mortem facilitation experience
Excellent written communication skills for design documents, runbooks, post-mortems, and operational documentation
Nice to Have
Knowledge of Generative AI tools and frameworks, including the application of AI-based predictive analytics and automation in infrastructure operations
Familiarity with ML platforms such as Kubeflow, MLflow, and Ray, as well as AI/ML training infrastructure
Experience with Kafka Streams, ksqlDB, or Apache Flink
Education
5-8 years of relevant experience.
Bachelor of Science Degree in Computer Science, Electrical Engineering, or a related field; a Master’s Degree is preferred.
At NetApp, we embrace a hybrid working environment designed to strengthen connection, collaboration, and culture for all employees. This means that most roles will have some level of in-office and/or in-person expectations, which will be shared during the recruitment process.
Why You'll Thrive at NetApp
At NetApp, you won't wait for the perfect moment - you'll make it. The early planning, the extra thought, the bold idea that turns good into great: That's how our people operate and how we continue to push the boundaries of data infrastructure.
NetApp is the trusted partner for organizations transforming data into opportunity. As the only enterprise-grade storage service natively embedded in Google Cloud, AWS, and Microsoft Azure, we empower customers to run everything from traditional workloads to enterprise AI with unmatched performance, resilience, and security.
Our culture
We celebrate mold breakers, bold thinkers, and problem solvers. We reward initiative, impact, and ownership. We provide flexibility so you can balance professional ambition with your personal life. Here, differences are not just welcomed - they drive everything we do.
If you're ready to innovate, rise to the challenge, and own every moment - make your next move your best one. now.
Submitting an application
To ensure a streamlined and fair hiring process for all candidates, our team only reviews applications submitted through our company website. This practice allows us to track, assess, and respond to applicants efficiently. Emailing our employees, recruiters, or Human Resources personnel directly will not influence your application.
Our values
Put the customer at the center. Care for each other and our communities. Think and act like owners. Build belonging every day. Embrace a growth mindset.
Benefits
Volunteer time off
40 hours of paid volunteer time each year.
Well-being
Employee Assistance Program, fitness, and mental health resources to help employees be their best.
Time away
Paid time off for vacation and to recharge.
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.