AIML SME
Tata Consultancy Services · Chennai, Tamil Nadu, India - Gandhinagar, Gujarat, India - Indore, Madhya Pradesh, India
Free to search · AI fit score against your CV · tailor your résumé in one click
Tata Consultancy Services · Chennai, Tamil Nadu, India - Gandhinagar, Gujarat, India - Indore, Madhya Pradesh, India
Job Description: • Advanced proficiency in Python/SQL, deep expertise in ML/DL frameworks (PyTorch, TensorFlow), and architecting complex data pre-processing/feature engineering pipelines. Deep Subject Matter Expertise in Core AI/ML algorithms, Deep Learning architectures, and Generative AI (LLMs, RAG, Transformer models). Enterprise-level Machine Learning Operations (MLOps), including CI/CD for ML, automated retraining workflows, and model governance. Deep hands-on expertise with the broader AI/ML Ecosystem Tools (Kubeflow, MLflow). Expert-level understanding of GCP compute, storage, networking, IAM, and the entire Vertex AI suite (Workbench, Pipelines, Feature Store, Model Registry, Endpoints, Agent Builder). • Extensive experience in provisioning, configuring, and optimizing distributed GPU and Google Cloud TPU (v4/v5) environments for training and inference. Ability to analyze distributed logs, profile CUDA/memory bottlenecks, and diagnose deep-seated infrastructure/pipeline failures. Expert knowledge of designing, integrating, and troubleshooting enterprise RESTful and gRPC AI APIs and microservices. Leading Root Cause Analysis (RCA), diagnosing systemic issues, and implementing permanent architectural resolutions for complex technical problems. Providing high-level technical advisory, architectural consulting, and escalation support to internal engineering teams and enterprise clients. Extensive experience architecting, fine-tuning, and optimizing Conversational Agents, Large Language Models, and Dialogflow CX systems. Exceptional ability to communicate intricate technical information and architectural trade-offs clearly and concisely to C-level stakeholders, architects, and engineering teams. Support with 24x7 operations (Rotational Shifts). English language (verbal and written) proficiency is a must. Key Responsibilities\\* 1. Troubleshooting Model & Pipeline Issues API & Integration Support: Diagnosing, optimizing, and resolving complex RESTful and gRPC API integrations, latency bottlenecks, and payload serialization issues between client enterprise applications and Vertex AI endpoints. Inference Failures: Leading root-cause investigations for critical inference failures, resolving memory overflows (OOM), optimizing GPU/TPU resource allocation, and fine-tuning serving runtimes (e.g., Triton, vLLM, TensorRT-LLM). Environment Configuration: Architecting and troubleshooting custom Docker containers, GKE/Kubernetes clusters, and cloud-native GCP ML environments with strict enterprise networking and VPC Service Controls (VPC-SC). 2. Data & Performance Monitoring Data Quality Checks: Designing and implementing automated validation frameworks to identify data quality anomalies, schema mismatches, and pipeline corruptions across BigQuery, Dataflow, and Vertex AI Feature Store. Monitoring Drift: Architecting enterprise-grade monitoring solutions using Vertex AI Model Monitoring to detect data drift and concept drift, establishing automated alerting and retraining triggers. • Accuracy Inquiries: Providing deep technical analysis on model predictions, bias, and reliability using advanced interpretability and explainability frameworks. 3. Product Education & Technical Documentation Knowledge Base Authoring: Authoring enterprise-grade reference architectures, technical blueprints, and definitive best-practice guides on topics like "Distributed Training on TPUs," "Production RAG Architecture," and "LLM Fine-Tuning on Vertex AI." Customer Onboarding: Leading architectural reviews, technical onboarding for enterprise engineering and data science teams. Translating Documentation: Synthesizing complex GCP product roadmaps, cutting-edge AI research, and core engineering release notes into actionable implementation strategies for technical leadership and IT teams. 4. The "Feedback Bridge" to Engineering Bug Reporting: Identifying, reproducing, and isolating complex platform defects, performing core-level debugging, and collaborating directly with Google Cloud / Product Engineering teams to drive fixes. Feature Requests: Aggregating enterprise-level capability gaps, creating detailed technical RFCs, and partnering with Product Management to influence the GCP AI/ML product roadmap. Edge Case Discovery: Documenting unique edge cases and failure modes where AI models/infrastructure fail under load, designing guardrails to improve future model resilience and system stability.