ML Infrastructure Engineer, Training

dyna-robotics

Redwood City, USonsite$220k-$320k/yrPosted Mar 31, 2026
Posting intelligenceMay be filled, listed long ago

Skills

kubernetespytorchgooglecloudawsml

About the role

Dyna Robotics builds general-purpose robots powered by a proprietary embodied AI foundation model with top-in-industry generalization and real-world performance. Already deployed with customers across multiple industries, our robots do commercial-grade work in the physical world. Our team comes from Google DeepMind, Meta, and Cruise, and we're backed by CRV, First Round, and other leading investors.

THE ROLE

As a ML Training Infrastructure Engineer, you will architect and build the systems that turn our multi-cloud GPU fleet into a training engine our researchers love. Your charter is singular and broad: own training infrastructure end-to-end so that every GPU is busy, every run is reproducible, and every researcher's next experiment is one command away.

WHAT YOU’LL DO

- Scale Distributed Training: Architect and own the infrastructure for large-scale GPU clusters. You’ll implement sharding, activation checkpointing, and memory optimization (ZeRO, FSDP) to enable the training of massive multimodal models.

- Optimize Researcher Ergonomics: Build a research codebase and job scheduling system (Kubernetes/SLURM) that prioritizes fast iteration, automated retries, and seamless failure recovery.

- High-Performance Data Handling: Design high-throughput pipelines to ingest and transform terabytes of multimodal robot data (video, proprioception, 3D signals), ensuring dataloaders never starve the GPUs.

- Production Inference: Build low-latency inference pipelines for real-time robot control. You’ll apply quantization, distillation, and model compilation (TensorRT, Triton) to move models from the lab to the physical world.

- Deep Systems Profiling: Dive into the weeds of GPU utilization, I/O bottlenecks, and memory fragmentation to squeeze every bit of performance out of our expanding compute fleet.

WHAT YOU’LL BRING

- 7+ Years of Engineering: With a track record of leading technical projects in high-performance computing (HPC) or ML infrastructure.

- ML Systems Mastery: Deep experience with PyTorch and distributed training frameworks (DeepSpeed, Accelerate). You understand the nuances of mixed precision and gradient accumulation.

- Infrastructure Expertise: Hands-on experience managing cloud GPU environments (GCP/AWS) and container orchestration (Kubernetes).

- Low-Level Intuition: A fundamental understanding of distributed systems, including race conditions, memory management, and NCCL/inter-node communication.

- Ownership Mindset: You don't just "deploy" code; you design, build, and operate systems end-to-end to unblock fast-moving research.

BONUS POINTS FOR

- Experience with Robotics Data Formats (MCAP, Protobuf) or multimodal models (VLAs).

- Deep ML systems experience: custom kernels (Triton), compilers, or runtime optimization.

- Experience as a founding or early-stage infrastructure hire.

Don’t let a checklist stop you. Data shows that underrepresented groups often only apply if they meet 100% of the criteria. We value problem-solving and grit over keyword matching. If you’re passionate about the intersection of geometry and robotics, we want to hear from you—even if you don't check every box.

Compensation

This MLOps Engineer role pays $220k-$320k/yr. Within typical range for mlops engineer roles in United States.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for MLOps Engineer roles in United States varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our MLOps Engineer hub for United States medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.