Quality
EXL · Noida, Uttar Pradesh, India
Free to search · AI fit score against your CV · tailor your résumé in one click
EXL · Noida, Uttar Pradesh, India
About the Role: We are looking for an experienced Agentic AI Testing Lead to build and lead an AI Quality Engineering team responsible for validating LLM-based, multi-agent, RAG, and autonomous AI workflows. The role involves defining AI testing strategies, evaluation frameworks, automation pipelines, quality KPIs, and governance practices to ensure reliable, safe, accurate, and scalable AI solutions. Key Responsibilities 1. Agentic AI Testing, Evaluation & Automation • Define and execute testing strategies for LLM-based, multi-agent, RAG, and Agentic AI systems. • Validate autonomous agent behavior, reasoning, memory, tool usage, API/database integrations, and end-to-end workflows. • Evaluate AI outputs for accuracy, relevance, groundedness, consistency, completeness, toxicity, bias, hallucination risk, and guardrail compliance. • Define AI quality KPIs such as hallucination rate, groundedness score, agent success rate, task completion rate, response relevancy, latency, cost efficiency, and user satisfaction. • Build automated evaluation pipelines, quality scoring mechanisms, dashboards, and CI/CD-integrated quality gates. • Develop reusable test harnesses, simulators, and benchmarking frameworks to compare models, prompts, and agent configurations. 2. Team Leadership & Capability Building • Build and lead a team of Agentic AI Quality Engineers. • Define team structure, testing standards, best practices, and governance models. • Mentor QA engineers in AI testing methodologies, evaluation techniques, and automation frameworks. • Drive innovation and adoption of emerging AI testing tools and technologies. • Collaborate with Product, Engineering, Data Science, and AI Research teams to improve overall AI quality. 3. Reporting & Stakeholder Management • Provide quality assessments and recommendations to leadership and stakeholders. • Present testing outcomes, risk assessments, KPI trends, and model evaluation reports. • Drive quality governance for Agentic AI initiatives across the organization. • Ensure traceability of testing activities, evaluation criteria, and quality benchmarks. Required Skills & Experience Technical Skills • 4 to 8 years of experience in Software Testing, Quality Engineering, or Test Automation. • Minimum 2+ years of hands-on experience in GenAI, LLM Testing, Agentic AI Testing, or AI Quality Engineering. • Strong understanding of LLMs, AI agents, RAG, prompt validation, tool calling, agent memory, MCP, and multi-agent orchestration. • Experience defining AI quality metrics, evaluation methodologies, benchmarking frameworks, and model comparison approaches. • Hands-on automation experience with Python, Playwright, Pytest, API automation, test framework development, and CI/CD quality gates. • Experience with AI evaluation frameworks such as DeepEval, Ragas, LangSmith, OpenAI Evals, or equivalent tools. • Exposure to cloud platforms such as Azure, AWS, or GCP.