NE

Staff DevOps Engineer

NEXXA

CAremote countryPosted Aug 27, 2026
Posting intelligenceActively listed

Skills

kubernetesdatabricksprometheussnowflaketerraformbigqueryredshiftcirclecijenkinsgrafanadatadoggithubgitlabpulumipythonazurecicdgooglecloudawsgoml

About the role

Nexxa is building the best AI systems for heavy industries - enabling machines, systems, and operations to think, decide, and act autonomously across manufacturing, large-scale infrastructure, logistics, and legacy environments.

Our mission is to translate deep technical breakthroughs into operational reality, solving some of the hardest systems-level problems in industry.

About the Role

We're looking for a Senior/Staff DevOps Engineer who has spent the last several years building and operating the infrastructure that lets AI and industrial systems run reliably at scale. You understand what it takes to keep production ML and data workloads fast, observable, and resilient - from GPU-backed training and inference clusters to the pipelines that connect them to real-world industrial environments.

This role is ideal for candidates who want deep infrastructure ownership at a company where uptime, latency, and reliability directly affect physical operations - not just software. You'll partner closely with AI, data, and product engineering teams to make sure the systems they build can actually run in production, safely and at scale.

What You'll Do

Own and evolve Nexxa's core infrastructure - compute, networking, storage, and deployment systems - end-to-end

Design and operate CI/CD pipelines that support fast, safe iteration across AI, data, and product engineering teams

Build and maintain infrastructure-as-code (e.g., Terraform, Pulumi) for reproducible, auditable environments across cloud and on-prem/edge deployments

Architect and manage Kubernetes-based platforms for training, inference, and application workloads, including GPU scheduling and autoscaling

Partner with data and AI teams to support the infrastructure behind:

Data warehouses and lakehouse architectures (e.g., Snowflake, BigQuery, Redshift, Databricks)

Feature stores, embedding indices, and retrieval pipelines

Model training, evaluation, and serving infrastructure

Define and drive observability practices - metrics, logging, tracing, and alerting - across distributed systems

Establish and enforce reliability practices: SLOs/SLIs, incident response, postmortems, and on-call rotations

Design for security and compliance across cloud infrastructure, secrets management, and access control, particularly relevant to industrial and legacy-environment integrations

Make pragmatic tradeoffs across cost, latency, reliability, and developer velocity

Collaborate with engineering leadership to define infrastructure roadmap and platform strategy

Mentor engineers on infrastructure best practices and raise the bar for operational excellence across the org

Required Qualifications

6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering roles

Deep hands-on experience with:

Cloud platforms (AWS, GCP, or Azure) at production scale

Kubernetes in production, including GPU workload scheduling

Infrastructure-as-code tooling (Terraform, Pulumi, or equivalent)

CI/CD systems (e.g., GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD)

Strong track record designing and operating observability stacks (e.g., Prometheus, Grafana, Datadog, OpenTelemetry)

Experience supporting ML/AI infrastructure - training clusters, model serving, data pipelines - a strong plus

Excellent scripting/programming skills (Python, Go, or Bash) for automation and tooling

Proven ability to independently scope and lead infrastructure projects from design through production rollout

Strong incident management instincts - you can lead through an outage calmly and drive toward root cause

Preferred Qualifications

Experience operating infrastructure that bridges cloud and edge/on-prem environments, especially in industrial or manufacturing contexts

Familiarity with data warehouse/lakehouse platforms (Snowflake, BigQuery, Redshift, Databricks)

Experience with service mesh, zero-trust networking, or compliance frameworks relevant to industrial/critical infrastructure (e.g., SOC 2, IEC 62443)

History of building internal developer platforms or self-service infrastructure tooling

Experience scaling infrastructure teams or setting technical direction at a Staff level

What Success Looks Like

You can own ambiguous, high-stakes infrastructure problems end-to-end

Systems you build stay reliable as usage and scale grow - you design for the next order of magnitude, not just today

You bring strong technical judgment on tradeoffs between reliability, cost, and speed

You raise the bar for operational rigor and engineering discipline across the team

You help define what's next for the platform, not just execute what's known

Why Join Nexxa.ai?

Innovative Environment: Play a critical role in transforming heavy industries through groundbreaking AI and automation technologies

Collaborative Culture: Be part of a team that values innovation, discipline, and continuous improvement

Professional Growth: Benefit from significant opportunities for career development and advancement

Competitive Compensation: Enjoy a comprehensive salary and equity package reflective of your expertise and contributions

If you're passionate about building the infrastructure that powers advanced AI solutions in the real world, we'd love to connect.

Questions about this role

Click "Apply with AI Applyd" above and you are done. Your resume is rewritten for this advert, the screening questions are answered, and it is submitted on NEXXA's own hiring system. No retyping your history, no fourteen tabs, no evening lost.

Compensation for DevOps / SRE roles in Canada varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our DevOps / SRE hub for Canada medians across recent openings.

You never touch the form - the application is filled and submitted for you on NEXXA's own hiring system. It is not marked sent when we press submit. It is marked sent when a confirmation from their system arrives at the address we apply with, and your dashboard shows which stage each application is at until then.

Twelve applicant tracking systems have a real apply path: Workday, Greenhouse, Lever, Ashby, Workable, iCIMS, Personio, Recruitee, Teamtailor, Rippling, Breezy and SmartRecruiters. Your application goes in on the employer's own hiring system, never into an aggregator queue.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.