AI SRE/ AI Site Reliability Engineer
Tata Consultancy Services · Bengaluru, Karnataka, India
Tata Consultancy Services · Bengaluru, Karnataka, India
**Key Responsibilities** - Manage and support infrastructure powering AI/ML and Generative AI applications. - Design and implement scalable, highly available, and secure platform solutions. - Build automation to reduce operational toil and improve platform reliability. - Develop and maintain Infrastructure as Code (Terraform, Helm, Ansible, CloudFormation). - Operate Kubernetes-based environments, container platforms, and cloud services. - Establish and monitor SLIs, SLOs, SLAs, observability dashboards, and alerting. - Lead incident response, root cause analysis (RCA), and reliability improvement initiatives. - Perform capacity planning, performance optimization, and cost management. - Support GPU-based compute environments, AI model serving, and data pipelines. - Implement disaster recovery, backup, resiliency, and security controls. - Collaborate with engineering and security teams to deploy and operate AI services safely. - Maintain operational documentation, runbooks, and best practices. - Participate in on-call support and production incident management. **Required Skills & Experience** - 5+ years of experience in **Site Reliability Engineering, DevOps, Infrastructure Engineering, or Platform Engineering**. - Strong programming/scripting skills in **Python, Go, Java, or similar languages**. - Hands-on experience with **Docker, Kubernetes, and container orchestration**. - Experience with **AWS, Azure, or Google Cloud Platform (GCP)**. - Strong knowledge of **Infrastructure as Code (Terraform, Helm, Ansible, CloudFormation)**. - Experience with **Observability and Monitoring tools** such as Grafana, Prometheus, Loki, ELK/EFK, Datadog, OpenTelemetry. - Knowledge of networking fundamentals: TCP/IP, DNS, Load Balancing, Routing. - Experience in incident management, RCA, performance tuning, and system scaling. - Strong understanding of security, compliance, governance, and production support. - Excellent communication and cross-functional collaboration skills. **Preferred Skills** - Experience supporting **Generative AI, LLM, MLOps, ModelOps, or AI Platforms**. - Knowledge of **GPU clusters, HPC environments, Slurm, Kubernetes GPU scheduling**. - Experience with **Kafka, Spark, Flink**, and distributed data processing frameworks. - Familiarity with databases such as **Snowflake, Redis, SQL, PostgreSQL**. - Understanding of **Embeddings, Fine-Tuning, RAG, Vector Databases, Model Serving**. - Experience with **Canary Deployments, Blue-Green Deployments, Chaos Engineering**. - Exposure to financial services or highly regulated environments.