MLOps Engineer
Skills
About the role
MLOps Engineer
Company: Reliance Enterprise Intelligence Ltd (REIL)
Function: Engineering – AI & Data
Location: Mumbai (Full-Time)
Experience: 5–7 Years
About REIL
Reliance Enterprise Intelligence Ltd (REIL) is a joint venture between Reliance Industries and Meta, combining Reliance’s unparalleled scale and deep enterprise domain knowledge with Meta’s world-class AI and technology capabilities. REIL is currently developing a cutting-edge enterprise AI platform focused on financial compliance and intelligence.
Role Overview
We are looking for an experienced MLOps Engineer to own the operational infrastructure of our AI platform. In this role, you will be responsible for deploying models reliably into production, building the infrastructure that keeps them running at scale, and ensuring that model performance is continuously tracked and maintained. You will set the MLOps standards for the squad and own the full lifecycle from model handoff to production monitoring and retraining.
Key Responsibilities1. ML Infrastructure & Tooling
Own the Infrastructure: Set up and manage the ML infrastructure on the data platform, including the model registry, experiment tracking, feature store, and serving infrastructure.
Define Standards: Enforce strict MLOps standards for model versioning, experiment tracking, and production readiness criteria.
CI/CD Pipelines: Build and maintain automated CI/CD pipelines for model development and deployment to ensure reliable, repeatable shipping.
2. Model Deployment & Serving
Production Deployment: Own the full pipeline from an ML Engineer’s trained model to a stable, monitored, and live serving endpoint.
Infrastructure Design: Build model serving infrastructure capable of handling real-time inference requirements with optimized latency and throughput.
Resilience & Rollbacks: Implement graceful fallback mechanisms (degrading to rule-based logic when confidence thresholds drop) and manage seamless model versioning and rollbacks.
3. Monitoring & Drift Detection
Dashboards & Alerts: Build and maintain performance dashboards tracking accuracy, latency, throughput, and data drift indicators, with automated alerts for anomalies.
Automated Retraining: Establish automated retraining pipelines triggered by drift detection or performance degradation to keep models compliant and up to date.
Feedback Loops: Maintain a tight feedback loop between application-layer validator decisions and model retraining workflows.
4. Production Reliability
Operational Readiness: Own production readiness procedures, including system monitoring, alerting setups, runbooks, and incident response protocols.
Cross-Functional Sync: Collaborate closely with Backend Engineers to ensure model inference endpoints meet the latency and reliability requirements of integrated systems.
Capacity Planning: Conduct regular load testing to guarantee the serving infrastructure seamlessly scales with growing transaction volumes.
Required Qualifications
Education: B.E. / B.Tech / M.Tech in Computer Science, Information Technology, or a related field.
Experience:
5+ years of overall engineering experience, with at least 3 years dedicated to MLOps or ML infrastructure in a high-scale production environment (not research or experimentation contexts).
Proven experience building model monitoring and automated retraining pipelines.
Technical Skills:
ML Platforms: MLflow (experiment tracking/model registry); Databricks MLOps stack is strongly preferred.
Model Serving: TorchServe, TensorFlow Serving, BentoML, or equivalent (Databricks Model Serving is a plus).
Infrastructure & IaC: Docker, Kubernetes, and Terraform (or equivalent).
CI/CD: GitHub Actions, GitLab CI, or equivalent.
Monitoring: Prometheus, Grafana, and data drift detection frameworks.
Languages: High proficiency in Python and familiarity with Bash scripting.
Cloud: Experience with at least one major cloud provider (AWS, Azure, or GCP).
Preferred Qualifications
Experience operating ML systems with real-time inference requirements in an enterprise setting.
Familiarity with feature stores to maintain training and serving consistency.
Prior experience in fintech, compliance, or highly regulated environments requiring strict model explainability and auditability.
Experience with A/B testing and shadow deployment patterns for safe model rollouts.
Core CompetenciesTechnical CompetenciesBehavioral Competencies Model Deployment & Serving Operational Ownership MLOps Infrastructure & Tooling Proactive Risk Management Monitoring & Drift Detection Attention to Detail CI/CD for ML Systems Cross-Functional Collaboration Production Reliability Engineering Structured Problem Solving
Pay: ₹322,930.16 - ₹2,000,000.00 per year
Benefits:
Health insurance
Paid sick time
Provident Fund
Work Location: In person
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.