Lead SRE
Baazi Games · Delhi, Delhi, India
Free to search · AI fit score against your CV · tailor your résumé in one click
Baazi Games · Delhi, Delhi, India
Lead Site Reliability Engineer (SRE) Experience: 7+ Years Location: Delhi NCR | Hybrid Employment Type: Full-time About the Role We are looking for a hands-on Lead SRE who can own production reliability end-to-end, beyond just managing DevOps tools and infrastructure. The ideal candidate should have strong experience with AWS, Kubernetes/EKS, CI/CD, Infrastructure as Code, observability, incident management, scalability, and distributed systems in high-traffic or high-transaction production environments. This is a technical leadership role where you will remain hands-on while mentoring the SRE/DevOps team and driving reliability practices across engineering. Key Responsibilities • Lead and mentor the DevOps/SRE team and establish engineering best practices. • Own platform reliability, availability, scalability, and operational excellence . • Design and operate highly available cloud infrastructure. • Build and manage production-grade Kubernetes/EKS environments. • Drive Infrastructure as Code using Terraform or similar tools. • Establish observability across monitoring, logging, tracing, and alerting . • Define and monitor SLIs, SLOs, and error budgets for critical services. • Lead production incidents, RCA, and postmortems. • Improve resilience through automation, capacity planning, disaster recovery, and performance engineering. • Partner with engineering teams to improve application reliability and operational readiness. • Drive cloud cost optimisation without compromising reliability. • Ensure infrastructure security, compliance, and operational governance. • Champion DevSecOps practices. • Standardise deployment, release management, and infrastructure governance. • Evaluate and adopt modern cloud-native and platform engineering technologies. Required Skills & Experience • 7+ years of relevant engineering experience with strong production infrastructure ownership. • Strong hands-on experience with AWS and Kubernetes/EKS . • Experience with Terraform or similar IaC tools . • Strong understanding of CI/CD and deployment automation . • Experience with observability tools such as Prometheus, Grafana, ELK/OpenSearch, Datadog, or New Relic . • Strong scripting/programming skills in Python, Bash, or Go . • Experience managing production incidents and implementing SRE practices. • Understanding of cloud security, IAM, secrets management, and infrastructure security . • Understanding of database reliability, backups, replication, and disaster recovery. • Strong troubleshooting and problem-solving skills. • Experience working with high-traffic or high-transaction production systems . Good to Have • Experience building or scaling an SRE function. • Service mesh experience with Istio/Linkerd . • Exposure to Kafka, SQS, Redis, Elasticsearch/OpenSearch and distributed messaging systems. • Experience with Platform Engineering / Internal Developer Platforms (IDPs) . • Knowledge of FinOps and cloud cost optimisation . • AWS, Kubernetes, or Terraform certifications. • Experience in exchange, brokerage, FinTech, gaming, or other low-latency/high-availability environments . Leadership Expectations • Remain hands-on and lead by example. • Mentor SRE/DevOps and engineering teams. • Drive an automation-first approach. • Work closely with Engineering, Security, QA, and Product teams. • Build a culture of ownership, reliability, continuous improvement, and operational excellence .