AI Model Governance Specialist (Risk)
ICICI Securities · Navi Mumbai, Maharashtra, India
Free to search · AI fit score against your CV · tailor your résumé in one click
ICICI Securities · Navi Mumbai, Maharashtra, India
Job Summary The role will be responsible for independently evaluating AI/LLM systems, identifying potential risk vectors, designing robust evaluation frameworks, and establishing controls to support responsible and secure enterprise AI deployment. Roles & Responsibilities • Conduct RAG & GenAI risk assessments across end-to-end RAG pipelines, including knowledge-base ingestion, chunking, embeddings, vector retrieval, context grounding, and generation. • Identify AI/LLM risks such as data leakage, prompt injection, context contamination, retrieval failures, and hallucinations. • Design and execute LLM evaluation and benchmarking frameworks to assess hallucination rates, toxicity, robustness, model calibration, fairness, and bias. • Curate golden/reference datasets and design evaluation test cases with reference answers. • Implement and validate LLM-as-a-Judge evaluation pipelines. • Establish AI governance and risk mitigation controls covering data preprocessing, pre-deployment validation gates, and post-deployment observability. • Independently inspect, evaluate, and validate AI/LLM model outputs. • Conduct technical audits of AI/LLM systems and assess model performance against defined evaluation criteria. • Support the development of enterprise AI control frameworks covering data provenance, model sign-off, rate-limiting, and continuous drift monitoring. • Evaluate automated AI assessment and hallucination detection tooling, including Ragas, TruLens, and DeepEval. Qualifications • B.Tech / M.Tech in Computer Science, Quantitative disciplines, or MCA. • Strong quantitative and/or Computer Science foundation. • Thorough understanding of AI/ML taxonomy and modern Generative AI workflows. • Deep understanding of RAG architecture and associated risks, including vector search, embeddings, relevance scoring, and context grounding/grounding boundaries. • Hands-on familiarity with reference-based testing, LLM-as-a-Judge frameworks, red-teaming fundamentals, and semantic similarity evaluation. Experience & Skills 9 to 12 years of relevant experience Required Technical Skills: • Strong expertise in Generative AI / GenAI and Large Language Models (LLMs). • Strong understanding of Retrieval-Augmented Generation (RAG) pipelines. • Experience in LLM evaluation, model validation, AI risk assessment, and AI governance. • Understanding of AI risk vectors including prompt injection, hallucination, data leakage, and retrieval failures. • Familiarity with enterprise LLM orchestration and deployment platforms such as Amazon Bedrock, Azure OpenAI Service, and Vertex AI. Preferred / Good-to-Have Skills: • Working knowledge of Python for querying APIs, parsing JSON outputs, and independently auditing model evaluation scripts. • Practical experience with NLP evaluation metrics such as ROUGE, BLEU, BERTScore, and semantic distance metrics. • Ability to construct enterprise AI control frameworks covering data provenance, model sign-off, rate-limiting, and continuous drift monitoring. • Familiarity with automated hallucination detection and AI evaluation tools such as Ragas, TruLens, and DeepEval.