AI/ML Engineer – Model Optimization & Acceleration
Skills
About the role
AI/ML Engineer – Model Optimization & Acceleration (8–10 Years)
Location: Bengaluru, India
Experience: 8–10 Years ( If you have experience more than 8 years only apply then)
Open Positions: 3
We are looking for an experienced AI/ML Engineer to optimize and deploy machine learning models across heterogeneous platforms (CPU, GPU, and NPU). If you're passionate about building high-performance, production-ready AI systems and working on cutting-edge technologies, we'd love to hear from you!
Key Responsibilities
Optimize AI models including LLMs, Diffusion Models, CNNs, Computer Vision, Multi-modal, and Speech Models.
Port models across frameworks (PyTorch → ONNX → Runtime).
Deploy and optimize models on GPU/NPU hardware accelerators.
Improve inference latency, throughput, and memory efficiency.
Implement quantization, model compression, and performance tuning.
Profile, benchmark, and debug AI system performance.
Required Skills
Strong expertise in PyTorch and ONNX
Proficiency in Python and C++
Experience with CUDA, ROCm, or GPU acceleration
Strong understanding of Transformers, CNNs, Deep Learning
Hands-on experience in Model Optimization, Quantization, Inference Optimization, and Performance Tuning
Good to Have
Edge AI / Embedded AI deployment
Generative AI or Multi-modal AI
Distributed inference or streaming pipelines
TensorRT, OpenVINO (preferred)
Pay: ₹2,500,000.00 - ₹3,000,000.00 per year
Experience:
AI/ML Engineer – Model Optimization & Acceleration: 8 years (Preferred)
Work Location: In person
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.