Data Scientist - Enterprise Search

Roche

Madrid, ESonsitePosted Jul 23, 2026
Posting intelligenceActively listedReposted 12×, possible evergreen/ghost posting

Skills

elasticsearchscikitlearntensorflowpytorchpythonazurenaturallanguageprocessingawsllmml

About the role

Bei Roche kannst du ganz du selbst sein und wirst für deine einzigartigen Qualitäten geschätzt. Unsere Kultur fördert persönlichen Ausdruck, offenen Dialog und echte Verbindungen. Hier wirst du für das, was du bist, wertgeschätzt, akzeptiert und respektiert. Dies schafft ein Umfeld, in dem du sowohl persönlich als auch beruflich wachsen kannst. Gemeinsam wollen wir Krankheiten vorbeugen, stoppen und heilen und sicherstellen, dass jeder Zugang zur Gesundheitsversorgung hat – heute und in Zukunft. Werde Teil von Roche, wo jede Stimme zählt.

Die Position

TheData Scientist - Enterprise Search role is responsible for contributing to the design and development of the next-generation Enterprise Search and information retrieval architectures.

This role will build context-aware, generative AI-driven search systems optimized for agentic readiness—enabling autonomous tool orchestration, deep semantic understanding, and intelligent information retrieval at an enterprise scale.

The data scientist will develop advanced AI solutions, with a strong focus on Generative AI, LLM-based applications, and scalable data services. This requires to work with large datasets, develop and evaluate machine learning models, and collaborate with cross-functional teams to improve the accuracy, coverage, relevance, and performance of search algorithms at enterprise scale.

This role involves direct communication with project stakeholders and contributes to team best practices, while identifying optimization opportunities that enhance the impact of moderately complex data solutions within larger product architectures. You will leverage advanced technical skills to translate business needs into actionable data science initiatives.

Job Responsibilities

Generative AI, Agentic AI and LLM Optimization

Model Development& Experimentation: Lead exploratory data analysis, feature engineering, model selection, training, validation, and performance evaluation for machine learning and AI-enabled solutions. Design and evaluate multiple modeling approaches, establish appropriate evaluation metrics, and optimize models for scalability, reliability, and business impact.

Experimentation and Innovation: lead experimental projects and drive innovation in enterprise search, exploring novel approaches like GraphRAG or agentic search patterns.

Agentic AI Search: Develop and deploy intelligent agentic architectures that can interact with and enhance the enterprise search experience.

RAG Experimentations (RAG Evaluation Framework): Design and conduct Retrieval-Augmented Generation experiments to evaluate and improve search relevance and performance.

LLM Model Evaluation: Evaluate the performance of Large Language Models in various enterprise search contexts, ensuring they meet business requirements and performance standards.

Advanced Prompt Engineering: Design and optimize prompts to programmatically enhance the interaction and effectiveness of search queries and responses.

Business Problem Solving& Decision Support

Partner closely with business stakeholders and product teams to translate complex business challenges into analytical approaches, ML solutions, and scalable intelligence capabilities.

Support data-driven prioritization and strategic decision making through actionable insights, predictive models, and operational intelligence.

Conducts A/B testing and experiments to assess the performance of search models and algorithms.

Consultancy Provide expert consultancy on data science and machine learning best practices, guiding internal teams and stakeholders.

PoC and knowledge sharing: Design and lead proof-of-value (PoV) projects, conduct knowledge-sharing sessions, and deliver impactful demos to showcase capabilities and gather feedback.

Data Engineering& Vector Databases

Data Engineering& Processing: Work with structured and unstructured data, building efficient pipelines for data ingestion, preprocessing, and feature engineering.

Vector database experimentation: Conduct experiments with vector databases to improve the efficiency and accuracy of our search systems.

Managing embeddings: Implement and optimize techniques for embeddings generation, indexation and retrieval to support advanced search queries and retrieval capabilities.

Retrieval& relevance tuning: Develop, evaluate, and tune retrieval algorithms to optimize search precision, recall, and relevance for diverse datasets.

Data quality: Design data enhancement modules that extract, enrich, and validate document content and metadata, directly improving downstream model context, search recall, and agentic reasoning.

Model Lifecycle and integration

Deployment, Testing, and Training of ML Models and Endpoints: Develop, deploy, and continuously refine machine learning models and endpoints to enhance search functionalities. Conduct rigorous testing and validation to ensure model accuracy and reliability.

Development of ML Models for search: Create and deploy advanced runnable models, such as entity extraction models and metadata augmentation, to the capabilities of our search solutions.

MLOps& Monitoring: Implement best practices for model deployment, versioning, monitoring, and performance optimization.

API& MCP Interoperability: Interface with enterprise search engines, platforms, and other APIs to enhance our search functionalities and integrations. Familiarity with MCP and protocols for agent interoperability.

Qualifications

Education / Experience

Master’s degree or PhD in Data Science, Computer Science, Statistics, Mathematics, Engineering, Artificial Intelligence, or a related quantitative field.

Demonstrated experience as a rising expert developing predictive models and leading specific analytical modules or project components.

Proven track record of taking full accountability for the quality and timely delivery of analytical tasks and troubleshooting complex data issues independently.

Experience working effectively on moderately complex data science problems and understanding how contributions fit into medium-sized data architectures.

Technical Skills

Shows strong proficiency in programming languages, particularly Python.

Has proven experience as a Data Scientist, preferably with a focus on information retrieval and NLP.

Has a solid understanding of natural language processing (NLP) techniques and tools.

Possesses hands-on experience with machine learning and deep learning frameworks and libraries (e.g., TensorFlow, PyTorch, scikit-learn).

Familiarity with cloud platforms and services, particularly AWS or Microsoft Azure.

Familiarity with version control systems (e.g., Git) and agile development practices.

Past experience with search engines and technologies (e.g. Elasticsearch, Solr, or Lucene) and solid understanding of search algorithms, information retrieval, and relevancy tuning is a plus.

Demonstrates excellent analytical and problem-solving skills, with the ability to work with large, complex datasets.

Proven ability to translate well-defined business questions into clear analytical problems and technical solutions.

Additional Qualifications

Strong communication and collaboration skills, with the ability to manage direct communication with immediate project stakeholders.

Ability to actively integrate feedback from technical peers and junior team members.

Proactive mindset to identify potential optimizations or new analytical approaches within the project scope.

Ability to work autonomously to achieve goals and deliver results, while actively collaborating with team members to meet shared team objectives.

Experience in healthcare, pharmaceutical, or other regulated industries is a plus.

******This vacancy will remain open until August 20.

To ensure a fair selection process, applications will be reviewed and candidates will be shortlisted only after that date

Wer wir sind

Eine gesündere Zukunft treibt uns zur Innovation an. Mehr als 100.000 Mitarbeiter weltweit arbeiten gemeinsam daran, wissenschaftliche Fortschritte zu erzielen und sicherzustellen, dass jeder Zugang zur Gesundheitsversorgung hat – heute und für zukünftige Generationen. Durch unser Engagement werden über 26 Millionen Menschen mit unseren Medikamenten behandelt und mehr als 30 Milliarden Tests mit unseren Diagnostik-Produkten durchgeführt. Wir ermutigen uns gegenseitig, neue Möglichkeiten zu erkunden, Kreativität zu fördern und hohe Ziele zu setzen, um lebensverändernde Gesundheitslösungen zu liefern.

Gemeinsam können wir eine gesündere Zukunft gestalten.

Roche ist ein Arbeitgeber, der die Chancengleichheit fördert.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for Data Scientist roles in Spain varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our Data Scientist hub for Spain medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.