Mid-Senior Data Scientist (Life Sciences)

Capgemini

Fundão, BRunknownPosted Jul 22, 2026
Posting intelligenceActively listedReposted 31×, possible evergreen/ghost posting

Skills

pythonneo4jml

About the role

At Capgemini Engineering, the world leader in engineering services, we bring together a global team of engineers, scientists, and architects to help the world’s most innovative companies unleash their potential. From autonomous cars to life-saving robots, our digital and software technology experts think outside the box as they provide unique R&D and engineering services across all industries. Join us for a career full of opportunities. Where you can make a difference. Where no two days are the same.

YOUR ROLE

We are looking for a Mid-Senior Data Scientist with a strong background in biomedical sciences and data engineering to join our growing team in Portugal. In this role, you will work at the heart of cutting-edge life sciences projects, helping to design, build, and govern the data foundations that power scientific discovery and product development. You will act as a bridge between scientific, engineering, and product stakeholders - translating complex biological knowledge into robust, scalable, and FAIR data solutions. In this role you will play a key role in:

Designing and governing biomedical data models: leading source-to-canonical mapping, ontology alignment, and schema governance including versioning, changelogs, and downstream impact assessments to ensure data integrity and scientific accuracy across complex biomedical domains

Building and maintaining data pipelines: developing robust, schema-driven pipelines in Python and SQL, performing exploratory data analysis, and implementing validation frameworks that support high-quality, reproducible scientific workflows

Driving knowledge graph development: applying hands-on experience with RDF, OWL, SPARQL, and property graph modelling tools such as Neo4j and GraphDB to build and enrich knowledge graphs that connect biomedical entities across diverse data sources

Applying machine learning and Gen AI: leveraging applied ML experience and familiarity with Gen AI tools (including text generation APIs, chatbots, and enterprise search solutions) to extract insight and value from scientific data at scale

Championing FAIR data principles: designing and delivering FAIR data products, leading harmonisation efforts across multiple source systems, and ensuring persistent identifiers and provenance are embedded into every data product

Aligning stakeholders across disciplines: driving alignment between scientific, engineering, and product teams through clear communication, structured documentation, and a solutions-focused mindset that keeps complex projects moving forward

Working with biomedical ontologies and controlled vocabularies: applying deep knowledge of resources such as Ensembl, UniProt, and Gene Ontology, including judgment on when and how to extend or map them to real-world data challenges

YOUR PROFILE

MSc or PhD in Bioinformatics, Biomedical Engineering, Molecular/Cell Biology, Neuroscience, Genetics, or related field

Excellent stakeholder management - able to drive alignment across scientific, engineering, and product stakeholders

Data modelling and harmonization experience across complex biomedical domains including source-to-canonical mapping, ontology alignment, persistent identifiers, and provenance.

Strong Python and SQL skills; comfortable building data pipelines and performing exploratory data analysis

Applied data science / ML experience relevant to knowledge graph or scientific data work

Strong communication skills with both technical and business stakeholders

Critical thinking, intellectual curiosity, and impact-driven mindset

Ability to adapt and manage priorities in fast-paced environments

Fluent in Portuguese and English

Nice-to-have:

Prior experience in a pharmaceutical or biotech organization

Experience with data catalogue, metadata registry, or schema registry tooling

Data engineering fundamentals: pipeline architecture, schema-driven automation, validation frameworks

Hands-on experience with ML frameworks and model lifecycle (build, deploy, monitor)

Hands-on experience with Gen AI models and tools, such as text generation APIs, chatbots, and enterprise search solutions.

Track record of leading schema governance: versioning, changelogs, tagged releases, downstream impact assessment

Experience designing FAIR data products and leading data harmonisation efforts across multiple source systems

Knowledge graph experience: RDF, OWL, SPARQL, and property graph modelling (Neo4j/GraphDB)

Experience with LinkML or equivalent schema modelling frameworks (classes, slots, ranges, constraints, cardinality, ontology bindings)

Strong command of biomedical ontologies and controlled vocabularies (e.g. Ensembl, UniProt, Gene Ontology), including judgment on when/how to extend or map them

WHAT YOU'LL LOVE ABOUT WORKING HERE

Join a multicultural and inclusive team environment.

Enjoy a supportive atmosphere promoting work-life balance.

Engage in exciting national and international projects.

Hybrid work.

Your career growth is central to our mission. Our array of career growth programs and diverse professionals are crafted to support you in exploring a world of opportunities.

Training and certifications programs.

Health and life insurance.

Referral program with bonuses for talent recommendations.

Great office locations.

ABOUT CAPGEMINI

Capgemini is an AI-powered global business and technology transformation partner, delivering tangible business value. We imagine the future of organizations and make it real with AI, technology, and people. With our strong heritage of nearly 60 years, we are a responsible and diverse group of 420,000 team members in more than 50 countries. We deliver end-to-end services and solutions with our deep industry expertise and strong partner ecosystem, leveraging our capabilities across strategy, technology, design, engineering and business operations.

Questions about this role

Click "Apply with AI Applyd" above. We auto-fill the application from your resume and answer screening questions in seconds. No copy and paste, no juggling tabs.

Compensation for Data Scientist roles in Brazil varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our Data Scientist hub for Brazil medians across recent openings.

Most applications complete in under 90 seconds. You can track the status in your dashboard and watch the screenshot proof land the moment the application submits.

AI Applyd supports Greenhouse, Lever, Ashby, Workday, iCIMS, SmartRecruiters, Personio, Teamtailor and other major ATS platforms. If we can submit through the platform, we do.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.