AI Data Engineer

WebSenor

Noida, INremote countryPosted Jul 22, 2026
Posting intelligenceActively listed

Skills

databricksairflowpythonopenaiazuresparkkafkacicdllmml

About the role

Senior Data Engineer – GenAI & Unstructured Data Pipelines

Job Type: Full-time Experience: 6–8 Years Location: Offshore (Remote)

Job Summary

We are seeking an experienced Senior Data Engineer – GenAI & Unstructured Data Pipelines to build next-generation AI data platforms that power Large Language Model (LLM) applications. The ideal candidate will have strong expertise in designing scalable data pipelines for unstructured and multi-modal data while enabling Retrieval-Augmented Generation (RAG), embeddings, vector search, and AI copilots.

This role is ideal for professionals passionate about modern data engineering, cloud-native architectures, and Generative AI technologies.

Key Responsibilities

Design, build, and maintain scalable data pipelines for structured, unstructured, and multi-modal data.

Own and support the end-to-end machine learning lifecycle, including data ingestion, feature engineering, model training, evaluation, deployment, monitoring, retraining, and rollback.

Develop production-grade ML pipelines using Azure native services with CI/CD automation and best practices.

Utilize Azure Machine Learning and MLflow for experiment tracking, model registry, and governed deployments across environments.

Design and implement Generative AI solutions using Azure OpenAI, embeddings, vector search, and Retrieval-Augmented Generation (RAG).

Build Agentic AI workflows with multi-step reasoning, tool integration, observability, guardrails, reliability, and cost optimization.

Develop scalable batch and streaming data pipelines using Azure Databricks, Apache Spark, and Kafka.

Create efficient RAG pipelines, including document chunking, embedding generation, indexing, and retrieval workflows.

Design and manage vector databases and search solutions such as Azure AI Search and Pinecone.

Build high-performance batch and real-time data ingestion pipelines using Spark and Kafka.

Required Skills & Qualifications

6–8 years of experience in Data Engineering.

Strong hands-on expertise with:

Python

PySpark

SQL

Apache Spark

Apache Airflow

Apache Kafka

Experience processing and managing unstructured data, including JSON, logs, documents, and PDFs.

Strong experience with cloud platforms, preferably Microsoft Azure.

Hands-on exposure to Generative AI technologies, including:

Large Language Models (LLMs)

Retrieval-Augmented Generation (RAG)

Embeddings

Vector Databases

Azure OpenAI

Azure AI Search

Pinecone

Experience working with Azure Databricks and modern data engineering frameworks.

Knowledge of CI/CD pipelines, MLflow, and Azure Machine Learning is highly preferred.

Preferred Skills

Experience building AI copilots and Agentic AI applications.

Understanding of data governance, monitoring, and production-grade ML operations.

Excellent analytical, problem-solving, and communication skills.

Work Location: Hybrid remote in Noida, Uttar Pradesh (Noida)

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for Data Engineer roles in India varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our Data Engineer hub for India medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.