Lead SRE / DevOps Architect
Tata Elxsi · Bengaluru, Karnataka, India
Free to search · AI fit score against your CV · tailor your résumé in one click
Tata Elxsi · Bengaluru, Karnataka, India
Role Overview We are seeking an experienced Infrastructure Architect / DevOps Lead / SRE to design and operate highly available, scalable production platforms across AWS, Kubernetes, Linux, and Media/OTT environments. The role requires strong hands-on expertise in cloud infrastructure, automation, container orchestration, monitoring, and Day2 operations for large-scale systems. Key Responsibilities Infrastructure & Cloud • Design, deploy, and manage production infrastructure across AWS and onprem • Operate AWS services: EC2, VPC, Load Balancers, Auto Scaling, S3, EBS/EFS, RDS, CloudFront • Ensure availability, scalability, security, and cost optimization DevOps & Automation • Automate deployments and operations using Shell scripting and Ansible • Support and improve CI/CD pipelines • Standardize deployment, rollback, and release processes Containers & Platform Operations • Manage Kubernetes clusters, deployments, scaling, upgrades, and troubleshooting • Operate Docker-based production workloads • Perform configuration and performance tuning Monitoring, Reliability & SRE • Implement monitoring and alerting with Prometheus, Grafana, Alertmanager, Kibana • Build dashboards and actionable alerts • Drive SRE practices to meet uptime, performance, and SLA goals Incident & Change Management • Lead production incidents, RCA, and preventive actions • Execute change management with validation and rollback • Ensure operational best practices Leadership & Stakeholder Management • Lead and mentor operations teams • Coordinate with internal teams, vendors, and customers • Handle documentation, KT, and operational readiness reviews Core Skills (MustHave) • Linux (RHEL, CentOS, Ubuntu) • AWS (production experience) • Kubernetes & Docker • Shell scripting (advanced) • Monitoring & Logging (Prometheus, Grafana, Kibana) • Incident Management & RCA • Networking fundamentals (TCP/IP, DNS, firewalls) Customer & Operations Governance • Act as primary operations contact for customers • Lead ops reviews, incident walkthroughs, and SLA reporting • Ensure clear, proactive communication Team & Workforce Management • Manage Day2 operations teams • Handle shifts, oncall schedules, and coverage • Mentor team members and ensure skill continuity Estimation & Onboarding • Provide effort and staffing estimates • Prepare Day2 operations models (L1/L2/L3) • Support transition from Build to Operate GoodtoHave • Ansible, CI/CD tools, Git • VMware (vCenter / ESXi) • Oracle / MS SQL Server • Basic Python or Java Behavioral Expectations • Strong ownership and problem-solving mindset • Calm and structured in production incidents • Experience with global, SLA-driven customers • Clear communication with technical and nontechnical stakeholders