Machine Learning Infrastructure Engineer, Technology

Point72

New York, USonsite$185k-$300k/yrPosted Aug 5, 2026
Posting intelligenceActively listed

Skills

kubernetesterraformairflowpythonazurec++cicdrustawsgoml

About the role

EXPERIENCE

Experienced Professionals

LOCATION

New York

FOCUS

Software & System Engineering

BUSINESS

Point72

A Career with Point72’s Technology Team

As Point72 reimagines the future of investing, our Technology team is constantly evolving our firm’s IT infrastructure and engineering capabilities, positioning us at the forefront of a rapidly evolving technology landscape. We’re a team of experts who experiment and work to discover new ways to harness open-source solutions, modern cloud architectures, and sophisticated Artificial Intelligence (AI) solutions, while embracing enterprise agile methodologies. Our commitment to building and innovating in the AI space provides the framework intended to drive smarter decision making and enhance how we build and operate our platforms and applications.

As a member of Point72’s Technology team, we encourage and support your professional development from day one - helping you advance your technical skills, contribute innovative ideas, and satisfy your own intellectual curiosity - all while delivering real business impact for our multi-billion-dollar global business.

What you’ll do

Design and implement high-performance infrastructure to support large-scale generative AI and machine learning workloads, enabling faster model iteration and real business impact

Design and operate distributed systems for model training, hyperparameter tuning, inference, and data preprocessing pipelines to deliver reliable end-to-end machine learning (ML) workflows

Collaborate with ML researchers and engineers to produce models, optimizing compute utilization, training throughput, and inference latency

Develop and automate deployment, orchestration, and CI/CD pipelines for models and data workflows using container orchestration and infrastructure-as-code (IaC)

Implement observability, monitoring, and cost-management strategies for GPU and accelerator compute environments to maintain predictable performance and spend

Evaluate, integrate, and benchmark emerging hardware and software technologies across cloud and on-prem environments to improve scalability and throughput

Drive security, compliance, and operational runbooks for GenAI infrastructure including access controls, secrets management, and incident response procedures

Troubleshoot, profile, and optimize performance across GPU and CPU compute stacks to remove bottlenecks and increase reliability

Document architecture, operational practices, and mentor engineers to team capability and accelerate adoption of production-ready GenAI infrastructure

What’s required

Bachelor's or master's degree in computer science, electrical engineering, or a related technical field

3–7 years of experience building and maintaining scalable compute or machine learning infrastructure systems

Deep understanding of distributed systems, container orchestration (Kubernetes), and public cloud platforms such as AWS, Google Cloud Platform, or Azure

Hands-on experience with machine learning operations and infrastructure tools such as MLflow, Ray, Airflow, Kubeflow, and Terraform

Strong understanding of reinforcement learning concepts and their infrastructure implications

Proficiency in Python and systems-level programming in one or more languages such as Go, C++, or Rust

Strong debugging, performance profiling, and optimization skills across GPU and CPU compute stacks

Experience implementing monitoring, observability, and cost-optimization for GPU/accelerator-based compute environments

Excellent collaboration and communication skills with a systems-thinking mindset

Commitment to the highest ethical standards

We take care of our people

We invest in our people, their careers, their health, and their well-being. When you work here, we provide:

Fully-paid health care benefits

Generous parental and family leave policies

Mental and physical wellness programs

Volunteer opportunities

Non-profit matching gift program

Support for employee-led affinity groups representing women, minorities and the LGBT+ community

Tuition assistance

A 401(k) savings program with an employer match and more

About Point72

Point72 Asset Management is a global firm led by Steven Cohen that invests in multiple asset classes and strategies worldwide. Resting on more than a quarter-century of investing experience, we seek to be the industry’s premier asset manager through delivering superior risk-adjusted returns, adhering to the highest ethical standards, and offering the greatest opportunities to the industry’s brightest talent. We’re inventing the future of finance by revolutionizing how we develop our people and how we use data to shape our thinking. For more information, visit www.Point72.com/working-here

The annual base salary range for this role is $185,000-$300,000 (USD) , which does not include discretionary bonus compensation or our comprehensive benefits package. Actual compensation offered to the successful candidate may vary from posted hiring range based upon geographic location, work experience, education, and/or skill level, among other things.

Compensation

This Platform Engineer role pays $185k-$300k/yr. Within typical range for platform engineer roles in United States.

Questions about this role

Click "Apply with AI Applyd" above and you are done. Your resume is rewritten for this advert, the screening questions are answered, and it is submitted on Point72's own hiring system. No retyping your history, no fourteen tabs, no evening lost.

Compensation for Platform Engineer roles in United States varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our Platform Engineer hub for United States medians across recent openings.

You never touch the form - the application is filled and submitted for you on Point72's own hiring system. It is not marked sent when we press submit. It is marked sent when a confirmation from their system arrives at the address we apply with, and your dashboard shows which stage each application is at until then.

Twelve applicant tracking systems have a real apply path: Workday, Greenhouse, Lever, Ashby, Workable, iCIMS, Personio, Recruitee, Teamtailor, Rippling, Breezy and SmartRecruiters. Your application goes in on the employer's own hiring system, never into an aggregator queue.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.