AI Agent Evaluation Engineer
Skills
About the role
About LetitbexAI:
LetitbexAI is a fast-growing AI-driven technology company focused on building intelligent, scalable, and enterprise-grade solutions. We work at the intersection of AI, data engineering, cloud, and business transformation, helping organizations unlock real value from artificial intelligence.
Position: AI Agent Evaluation Engineer
Experience: 8+ Years
Notice Period: Can be considered up 15 Days
Job Position (Title) : AI Agent Evaluation Engineer
Experience: 8+ Years
Job Summary ::
We are seeking a highly motivated and technically proficient AI Agent Evaluation Engineer to join our growing AI team. This crucial role will be responsible for defining, developing, and executing robust Agent evaluation frameworks and test strategies, with a significant focus on Responsible AI and Safety Evals, for our agents built using the Google Agent Development Kit (ADK). The ideal candidate will bridge the gap between AI development and reliable deployment, ensuring our agents are safe, ethical, effective, and meet high-quality performance standards. The role will be of 70% Automation and 30% Manual Testing
Required Skills & Qualifications:
Experience: 8+ years in Software QA, with at least 2-3 years focused on testing or evaluating AI/ML systems, conversational agents, or Large Language Models (LLMs).
Safety Evals Expertise (Mandatory): Direct experience in designing and executing safety evaluations (red teaming, adversarial testing), bias detection, and measuring toxicity/harmful content in generative AI models.
Agent/LLM Evals: Proven experience developing and running general evaluations (Evals) for LLM-powered applications knowing libraries like PyTest (Must)
Google ADK Familiarity (Mandatory): Direct or strong conceptual understanding of the Google Agent Development Kit (ADK) and its components.
Programming: Strong proficiency in Python is mandatory for script development, data processing, and automation.
Cloud & MLOps: Familiarity with Google Cloud Platform (GCP) services relevant to AI/ML (e.g., Vertex AI) and integrating testing into MLOps workflows.
Communication: Excellent analytical, problem-solving, and communication skills to articulate complex agent behaviors, risks, and safety findings clearly.
Tools and Libraries: Langsmith, DeepEval, Ragas, Giskard, Hugging face.
Questions about this role
Want AI Applyd to auto-apply to roles like this?
We tailor your resume per posting, fill the forms, and track replies for you.