Principal Technical ML Engineer
Skills
About the role
Job Description Notes: The ML team currently consists of two Principal-level ML engineers and 1 senior-level engineer This is a hands-on IC role for someone who is ML-first but thinks like a product engineer - someone who can build and ship LLM-driven features while staying grounded in how real customers use them. The team is moving fast on evals, optimization, and novel LLM workflows, and needs someone who can own that work end to end without adding to the load of the existing ML team. Essentially we want someone who can ship on the engineering side and has solid ML Ops experience. So priority in order is: 1) can write good production code 2) ML Ops experience 3) wants to be customer facing is a nice-to-have, no longer a must-have. Circuit summed it up like this: "We have specialized models deployed on our own infrastructure to solve specific problems in our pipeline. These touch indexing, retrieval, and inference, and our dependency on self-hosted models is expecting to grow. We want to take the majority of this burden from the infrastructure team and own it on the ML team. You'd be the primary owner of this, while also shipping ML related product features." Description: As a Principal Technical ML Engineer, you are a strong software engineer first, with deep hands-on experience in ML systems and model serving infrastructure. This is not a research role and it is not an architecture role. It is a hands-on Principal IC role with high ownership where all your technical judgment shows up in shipped systems. You will work closely with other AI/ML engineers to understand what they are building and drive it through to shipped product. The gap between promising research and working software is where you operate. You bring the engineering discipline to close that gap reliably and repeatedly. Key Responsibilities Own the engineering work that moves ML capabilities from research into production. You write the code, you ship the feature, you are accountable for it working. Work across the AI/ML team to understand their work and make sound engineering judgments about what is ready to ship and how to get it there. Build and maintain model serving and evaluation infrastructure. You understand the stack deeply, including how inference works, why serving choices matter, and what the tradeoffs are. Lead implementation of ML-driven features, coordinating with the research team and the rest of engineering to get things shipped. Debug production systems including edge cases and failure modes in ML pipelines and retrieval systems. Establish and improve observability, debugging, and testing practices for ML systems. Build and maintain evaluation systems, including datasets, scoring approaches, and repeatable testing to detect regressions. Improve the reliability, testability, and maintainability of the ML codebase without slowing down iteration. Use AI coding tools fluently as a core part of your daily workflow, with the judgment to evaluate what they produce and catch errors before they ship. Experience Strong software engineering background with a track record of owning and shipping production systems end to end. Genuine MLOps depth. You have run model serving and/or evaluation infrastructure in production, understand systems like KServe, and have real opinions about the serving stack. You are not someone whose experience stops at configuring managed cloud ML services. You understand the serving stack deeply enough to make tradeoffs, debug failures, and improve it. Enough ML knowledge to be a real partner to the AI/ML team. You can follow what they are building, ask the right questions, and make sound engineering decisions about productionizing it. You are not training models. Experience with LLMs, RAG systems, or agent-based workflows at a level deeper than stringing together API calls. Strong proficiency in Python. Comfortable in Go or willing to get there quickly. Experience integrating multiple systems, APIs, and data sources into cohesive product functionality. Experience designing or working with evaluation systems for ML quality. Experience debugging production ML systems including handling edge cases and failure modes. Daily, fluent use of AI coding tools as a core part of your engineering workflow. Human Skills Ownership mindset with a focus on delivering working systems in production. Bias toward shipping. You take ambiguous problems and drive them to working solutions without needing the path fully defined for you. Engineering judgment about when research is ready to productionize and what it takes to get it there. Critical evaluation of AI-generated output. You spot incorrect logic and subtle bugs before they reach production. Systems thinking, including attention to correctness and failure modes. Ability to operate effectively in fast-moving environments with high ownership. Low ego, with a focus on team outcomes. What We Offer Early-Stage Ownership: Join at the ground floor of a company with real traction and momentum. Empowered Culture: We value autonomy, candor, and craft. You'll be trusted to lead. Cutting-Edge Tech: Work with the latest in AI, backend systems, and intelligent infrastructure. Meaningful Impact: Shape a platform that transforms how organizations activate knowledge. Holistic Benefits: Competitive comp, equity, 100% paid healthcare, 401K, flexible PTO, and a team that truly cares.
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.