Lead Data Scientist-(VLM-Multimodal AI ,CV, Deep Learning)

HERE Technologies

Mumbai, INonsitePosted Jul 6, 2026
Posting intelligenceActively listed

Skills

pytorchpythonml

About the role

What's the role?:

We are seeking an experienced Lead Data Scientist to drive the development of advanced Vision Foundation Models, Vision-Language Models, and multimodal AI systems for large-scale image and video understanding.

You will lead the design, development, and deployment of scalable AI solutions, translating cutting-edge research into real-world applications.

Key Responsibilities

Design, train, and optimize large-scale vision foundation models across image and video modalities

Develop multimodal AI systems using architectures such as Vision Transformers (ViT), SAM, DINOv3, CLIP, and VLMs

Apply self-supervised learning, transfer learning, and fine-tuning approaches for downstream tasks

Build and enhance Vision-Language Models for visual reasoning and multimodal understanding

Develop Retrieval-Augmented Generation (RAG) pipelines and multimodal knowledge retrieval systems

Work with embeddings, vector databases, and semantic search frameworks

Build scalable pipelines for training, evaluation, and deployment

Manage large-scale image, video, and multimodal datasets

Optimize distributed training workflows and model performance

Translate research into production-ready solutions and explore emerging approaches in multimodal AI and generative AI

Evaluate model quality, robustness, and retrieval effectiveness

Who are you?:

You bring strong expertise in computer vision, foundation models, and multimodal AI systems, along with the ability to deliver scalable solutions from research to production.

Master’s or PhD in Computer Science, Artificial Intelligence, Machine Learning, or a related field

Extensive experience in deep learning, computer vision, or multimodal AI

Strong programming skills in Python and experience with PyTorch

Deep understanding of computer vision, Vision Transformers, self-supervised learning, Vision-Language Models, and multimodal systems

Hands-on experience with foundation models such as SAM, DINOv3, CLIP, BLIP/BLIP-2, LLaVA, or diffusion-based vision models

Experience building RAG pipelines, semantic retrieval systems, and working with embeddings and vector databases such as FAISS, Milvus, Pinecone, or Weaviate

Experience working with large-scale image and video datasets and distributed training environments

Familiarity with GPU acceleration and scalable ML infrastructure

Exposure to generative AI, multimodal reasoning systems, or large-scale perception systems

Contributions to research, publications, or open-source projects are valued

What Do We Offer?

Opportunity to work on cutting-edge AI and multimodal technologies

A collaborative, inclusive, and innovation-driven work environment

Opportunities to learn, grow, and advance your career

Exposure to large-scale, real-world AI challenges and global impact

Competitive compensation and performance-based bonus

Flexible and hybrid working options

Employee wellness programs and professional development support

Who are we?:

HERE Technologies is a location data and technology platform company. We empower our customers to achieve better outcomes – from helping a city manage its infrastructure or a business optimize its assets to guiding drivers to their destination safely.

At HERE we take it upon ourselves to be the change we wish to see. We create solutions that fuel innovation, provide opportunity and foster inclusion to improve people’s lives. If you are inspired by an open world and driven to create positive change, join us. Learn more about us on our YouTube Channel.

About the Team

You will be part of a highly collaborative AI/ML team focused on developing next-generation Vision Foundation Models (VFMs), Vision-Language Models (VLMs), and multimodal AI systems. The team works at the intersection of research and scalable production systems, driving innovation in large-scale image and video understanding.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for Data Scientist roles in India varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our Data Scientist hub for India medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.