AI/ML Engineer – Model Optimization & Acceleration

Programming.com

Bengaluru, INonsitePosted Jul 22, 2026
Posting intelligenceActively listedReposted 8×, possible evergreen/ghost posting

Skills

pytorchpythonc++ml

About the role

AI/ML Engineer – Model Optimization & Acceleration (8–10 Years)

Location: Bengaluru, India

Experience: 8–10 Years ( If you have experience more than 8 years only apply then)

Open Positions: 3

We are looking for an experienced AI/ML Engineer to optimize and deploy machine learning models across heterogeneous platforms (CPU, GPU, and NPU). If you're passionate about building high-performance, production-ready AI systems and working on cutting-edge technologies, we'd love to hear from you!

Key Responsibilities

Optimize AI models including LLMs, Diffusion Models, CNNs, Computer Vision, Multi-modal, and Speech Models.

Port models across frameworks (PyTorch → ONNX → Runtime).

Deploy and optimize models on GPU/NPU hardware accelerators.

Improve inference latency, throughput, and memory efficiency.

Implement quantization, model compression, and performance tuning.

Profile, benchmark, and debug AI system performance.

Required Skills

Strong expertise in PyTorch and ONNX

Proficiency in Python and C++

Experience with CUDA, ROCm, or GPU acceleration

Strong understanding of Transformers, CNNs, Deep Learning

Hands-on experience in Model Optimization, Quantization, Inference Optimization, and Performance Tuning

Good to Have

Edge AI / Embedded AI deployment

Generative AI or Multi-modal AI

Distributed inference or streaming pipelines

TensorRT, OpenVINO (preferred)

Pay: ₹2,500,000.00 - ₹3,000,000.00 per year

Experience:

AI/ML Engineer – Model Optimization & Acceleration: 8 years (Preferred)

Work Location: In person

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for Machine Learning Engineer roles in India varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our Machine Learning Engineer hub for India medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.