Evaluation Engineer, Applied AI

Mistral AI

Paris, FRonsitePosted Jul 22, 2026
Posting intelligenceActively listed

Skills

huggingfacepytorchpythonopenaillmml

About the role

Location

Paris

Employment Type

Full time

Department

Solutions

About Mistral

Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems - across high-stakes industries like finance, manufacturing, defense, healthcare, and the public sector - co-creating customized AI systems that they can run on their terms.

We are a dynamic, collaborative team passionate about AI and its potential to transform society. Our diverse workforce thrives in competitive environments and is committed to driving innovation. Our teams are distributed between Europe, North America, Asia and the Middle East. We are creative, low-ego and team-spirited.

The Role

As an Evaluation Engineer on the Applied AI team, you will shape how enterprise customers measure and trust AI solutions. You will work closely with clients and internal stakeholders to deliver evaluation systems that define when an AI model is ready for real business impact. The team's mission is to ensure every project - no matter how ambitious - moves from idea to production with clarity and accountability. Your work will sit at the intersection of research, engineering, and solutions, directly influencing how models are improved and deployed across varied domains.

What You Will Do

Design and implement comprehensive frameworks to evaluate large language model (LLM) performance across customer use cases

Build and maintain scalable evaluation infrastructure and pipelines for rapid, reproducible assessment

Develop new evaluation methods for sector-specific or emerging model capabilities

Collaborate with customers to create custom evaluation suites that fit their unique needs and criteria

Partner with research teams to turn evaluation insights into actionable model improvements

Work with product teams to refine evaluation tooling based on user feedback

Define and communicate clear success metrics for "production-ready" AI models

What We're Looking For

At least 3 years of experience evaluating ML, LLM, or agentic systems

Proven background in implementing AI or machine learning products, especially with APIs or back-end systems

Strong Python coding skills

Deep understanding of machine learning concepts and algorithms, especially as they pertain to LLMs

Experience communicating technical topics clearly to both technical and non-technical audiences

Familiarity with evaluation frameworks like LM Eval Harness or OpenAI Evals, or contributions to related open-source projects

Experience working with ML libraries (e.g., PyTorch, HuggingFace Transformers)

Collaborative, direct communicator who values outcomes and low-ego teamwork

What We Offer

We offer a comprehensive benefits package designed to support your well-being, growth, and work-life balance. Benefits vary by country and may include healthcare coverage, parental leave, retirement plans, relocation support, wellness programs, meal and transportation allowances, and other location-specific perks.

For the most up-to-date details on benefits available in your location, please refer to our Benefits page.

Questions about this role

Click "Apply with AI Applyd" above and you are done. Your resume is rewritten for this advert, the screening questions are answered, and it is submitted on Mistral AI's own hiring system. No retyping your history, no fourteen tabs, no evening lost.

Compensation for Software Engineer roles in France varies widely by seniority, employer size, and remote vs onsite arrangement. Check the salary range on this listing when published, or browse our Software Engineer hub for France medians across recent openings.

You never touch the form - the application is filled and submitted for you on Mistral AI's own hiring system. It is not marked sent when we press submit. It is marked sent when a confirmation from their system arrives at the address we apply with, and your dashboard shows which stage each application is at until then.

Twelve applicant tracking systems have a real apply path: Workday, Greenhouse, Lever, Ashby, Workable, iCIMS, Personio, Recruitee, Teamtailor, Rippling, Breezy and SmartRecruiters. Your application goes in on the employer's own hiring system, never into an aggregator queue.

Want AI Applyd to auto-apply to roles like this?

We tailor your resume per posting, fill the forms, and track replies for you.