AI Accelerator, Software Principal Engineer- Full-Stack

Ampere Computing

Santa Clara, UShybrid$195k-$292k/yrPosted Jul 23, 2026
Posting intelligenceActively listed

Skills

pytorch

About the role

Description

Invent the future with us.

Ampere is a semiconductor design company for a new era, leading the future of computing with an innovative approach to CPU design focused on high-performance, energy efficient AI compute.

As a pioneer in the new frontier of energy efficient high-performance computing, Ampere is part of the Softbank Group of companies driving sustainable computing for AI, Cloud, and edge applications.

Join us at Ampere and work alongside a passionate and growing team - we’d love to have you apply!

About the Role:

We are looking for an engineer with strong experience in PyTorch-based AI deployment, accelerated inference execution, and systems integration across software components. The role involves working on temporal and multi-modal workloads, optimizing execution on target platforms, and building infrastructure to run AI models reliably in production environments.

What You’ll Achieve:

Deploy and validate different AI models across supported inference environments (local, on-prem, or edge/accelerated platforms), optimizing runtime performance and reliability.

Develop and tune inference graph transformations, including torch.export-based graph workflows.

Collaborate with platform and infrastructure teams to improve model execution efficiency on supported AI accelerators and compute environments.

Integrate model inference pipelines into scalable services, including batching, streaming, and runtime orchestration.

Build and maintain middleware and communication layers to support modular, scalable system integration.

Support long-term platform development for end-to-end inference services and production readiness.

Relevant Technical Areas:

Temporal model architectures

Multi-frame or sequence embeddings (e.g., video/text sequences)

Attention-based models

Multi-modal workloads (e.g., text, imaging, and other feature modalities)

Automated labeling, evaluation, and validation workflows

CPU/runtime performance optimization for system components

Publish-subscribe middleware and distributed communication systems

About You:

Bachelors degree in Computer Science, Mathematics or a related technical field & 8 years of related experience; or Master's degree & 6 years

Strong hands-on experience with PyTorch

Experience deploying AI models to accelerated or constrained environments (e.g., edge, cloud GPU, or specialized accelerators)

Familiarity with graph optimization and model-performance tuning

Experience working with middleware, messaging, or distributed communication layers

Good understanding of hardware/software interaction in AI systems

Experience collaborating with hardware or platform partners

What We’ll Offer:

At Ampere we believe in taking care of our employees and providing a competitive total rewards package that includes base pay, cash long-term incentive, and comprehensive benefits. The full base pay range for this role is between $195,000 and $292,000. Our benefits include health, wellness, and financial programs that support employees through every stage of life.

Benefit highlights include:

Premium medical insurance, dental insurance, vision insurance, as well as income protection and a 401K retirement plan, so that you can feel secure in your health and financial future.

Unlimited Flextime and 10+ paid holidays so that you can embrace a healthy work-life balance.

A variety of healthy snacks, energizing espresso, and refreshing drinks to keep you fueled and focused throughout the day.

And there is much more than compensation and benefits. At Ampere, we foster an inclusive culture that empowers our employees to do more and grow more. We are excited to share more about our career opportunities with you through the interview process. Our benefits include health, wellness, and financial programs that support employees through every stage of life.

#LI-CB1

#LI-DR

#LI-Hybrid

Compensation

This Full-Stack Engineer role pays $195k-$292k/yr. Within typical range for full-stack engineer roles in United States.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for Full-Stack Engineer roles in United States varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our Full-Stack Engineer hub for United States medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.