AI Data Engineer

EXL Service

Noida, INonsitePosted Jul 8, 2026
Posting intelligenceActively listed

Skills

databrickssnowflakelangchainpythonopenaiflaskazurecicdgooglecloudnaturallanguageprocessingawsllmml

About the role

Job Description: Key Responsibilities

Design and develop LLM-based applications using single-agent or simple multi-agent patterns for business use cases

Build and maintain RAG pipelines : data ingestion chunking embeddings retrieval response generation

Implement prompt engineering techniques (prompt templates, chaining, basic tool/function calling)

Develop backend services/APIs for AI applications using Python frameworks (FastAPI / Flask / Streamlit)

Integrate AI solutions with enterprise systems, databases, and APIs

Apply basic guardrails and validation checks to improve response quality and reduce hallucination

Work with Data Engineering teams to ensure data quality, pipeline efficiency, and proper documentation

Collaborate with MLOps teams for deployment, monitoring, and iterative improvements

Document solutions, reusable components, and best practices

Must-Have Skills

Experience

4–6 years total experience , with 1+ year hands-on experience in GenAI / LLM-based applications

LLM / GenAI & Agentic Engineering

Strong hands-on experience with:

LLMs (Claude, OpenAI, etc.)

RAG pipelines and retrieval optimisation

GPT + Agentic AI implementation experience

Experience with:

LangChain, LangGraph, or similar frameworks

Agent orchestration and tool-calling architectures

Deep understanding of: LLM limitations, evaluation, and optimisation strategies

Core Engineering

Strong Python/Pyspark engineering expertise (production-grade development) with proven API integration experience

Deep data analysis experience and handling large volume of data

Fabric/Azure Databricks/Snowflake data engineering integration skills

Good exposure to:

Cloud platforms (Azure/AWS/GCP)

SQL

Containers, CI/CD, monitoring

Data / AI Foundations (Mandatory)

Prior experience in one or more:

Data Engineering (ETL/ELT, pipelines, orchestration)

Data Science / ML lifecycle (especially NLP)

Analytics engineering / data products

Good-to-Have / Preferred

Exposure to model fine-tuning (LoRA/PEFT) or prompt optimisation techniques

Experience with evaluation of LLM outputs (quality, relevance, latency)

Understanding of enterprise data privacy and security considerations in GenAI

Exposure to Azure AI / Azure OpenAI / AI Search ecosystems

Experience working on real client-facing AI solutions or POCs

Responsibilities: Key Responsibilities

Design and develop LLM-based applications using single-agent or simple multi-agent patterns for business use cases

Build and maintain RAG pipelines : data ingestion chunking embeddings retrieval response generation

Implement prompt engineering techniques (prompt templates, chaining, basic tool/function calling)

Develop backend services/APIs for AI applications using Python frameworks (FastAPI / Flask / Streamlit)

Integrate AI solutions with enterprise systems, databases, and APIs

Apply basic guardrails and validation checks to improve response quality and reduce hallucination

Work with Data Engineering teams to ensure data quality, pipeline efficiency, and proper documentation

Collaborate with MLOps teams for deployment, monitoring, and iterative improvements

Document solutions, reusable components, and best practices

Must-Have Skills

Experience

4–6 years total experience , with 1+ year hands-on experience in GenAI / LLM-based applications

LLM / GenAI & Agentic Engineering

Strong hands-on experience with:

LLMs (Claude, OpenAI, etc.)

RAG pipelines and retrieval optimisation

GPT + Agentic AI implementation experience

Experience with:

LangChain, LangGraph, or similar frameworks

Agent orchestration and tool-calling architectures

Deep understanding of: LLM limitations, evaluation, and optimisation strategies

Core Engineering

Strong Python/Pyspark engineering expertise (production-grade development) with proven API integration experience

Deep data analysis experience and handling large volume of data

Fabric/Azure Databricks/Snowflake data engineering integration skills

Good exposure to:

Cloud platforms (Azure/AWS/GCP)

SQL

Containers, CI/CD, monitoring

Data / AI Foundations (Mandatory)

Prior experience in one or more:

Data Engineering (ETL/ELT, pipelines, orchestration)

Data Science / ML lifecycle (especially NLP)

Analytics engineering / data products

Good-to-Have / Preferred

Exposure to model fine-tuning (LoRA/PEFT) or prompt optimisation techniques

Experience with evaluation of LLM outputs (quality, relevance, latency)

Understanding of enterprise data privacy and security considerations in GenAI

Exposure to Azure AI / Azure OpenAI / AI Search ecosystems

Experience working on real client-facing AI solutions or POCs

Qualifications: Key Responsibilities

Design and develop LLM-based applications using single-agent or simple multi-agent patterns for business use cases

Build and maintain RAG pipelines : data ingestion chunking embeddings retrieval response generation

Implement prompt engineering techniques (prompt templates, chaining, basic tool/function calling)

Develop backend services/APIs for AI applications using Python frameworks (FastAPI / Flask / Streamlit)

Integrate AI solutions with enterprise systems, databases, and APIs

Apply basic guardrails and validation checks to improve response quality and reduce hallucination

Work with Data Engineering teams to ensure data quality, pipeline efficiency, and proper documentation

Collaborate with MLOps teams for deployment, monitoring, and iterative improvements

Document solutions, reusable components, and best practices

Must-Have Skills

Experience

4–6 years total experience , with 1+ year hands-on experience in GenAI / LLM-based applications

LLM / GenAI & Agentic Engineering

Strong hands-on experience with:

LLMs (Claude, OpenAI, etc.)

RAG pipelines and retrieval optimisation

GPT + Agentic AI implementation experience

Experience with:

LangChain, LangGraph, or similar frameworks

Agent orchestration and tool-calling architectures

Deep understanding of: LLM limitations, evaluation, and optimisation strategies

Core Engineering

Strong Python/Pyspark engineering expertise (production-grade development) with proven API integration experience

Deep data analysis experience and handling large volume of data

Fabric/Azure Databricks/Snowflake data engineering integration skills

Good exposure to:

Cloud platforms (Azure/AWS/GCP)

SQL

Containers, CI/CD, monitoring

Data / AI Foundations (Mandatory)

Prior experience in one or more:

Data Engineering (ETL/ELT, pipelines, orchestration)

Data Science / ML lifecycle (especially NLP)

Analytics engineering / data products

Good-to-Have / Preferred

Exposure to model fine-tuning (LoRA/PEFT) or prompt optimisation techniques

Experience with evaluation of LLM outputs (quality, relevance, latency)

Understanding of enterprise data privacy and security considerations in GenAI

Exposure to Azure AI / Azure OpenAI / AI Search ecosystems

Experience working on real client-facing AI solutions or POCs

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for Data Engineer roles in India varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our Data Engineer hub for India medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.