Site Reliability Engineer (SRE)Manager - IT
Apollo Hospitals · Chennai, Tamil Nadu, India
Apollo Hospitals · Chennai, Tamil Nadu, India
**Site Reliability Engineering (SRE) Manager - Information Technology** *Function: Software Performance Monitoring & Operations (24x7)* **The Apollo Hospitals Family** Set-up in 1983 by Dr. Prathap C. Reddy, renowned as the architect of modern healthcare in India. As the nations first corporate hospital, Apollo Hospitals is acclaimed for pioneering the private healthcare revolution in the country. Apollo Hospitals has emerged as Asia’s foremost integrated healthcare services provider and has a robust presence across the healthcare ecosystem, including Hospitals, Pharmacies, Primary Care & Diagnostic Clinics and several retail health models. The cornerstones of Apollo’s legacy are its unstinting focus on clinical excellence, affordable costs, modern technology and forward-looking research & academics. **Role Description** The SRE Manager leads Apollo's Software Performance Monitoring and Operations function, accountable for the reliability, performance, resilience and self-healing posture of all clinical, revenue and enterprise systems that run 24x7 across 50+ hospitals. This role owns the observability maturity roadmap — progressing from foundational visibility to SLI/SLO-driven operations and ultimately to autonomous, anti-fragile systems powered by AIOps. The SRE Manager establishes error budget policies, leads incident command for critical production issues, and drives blameless postmortems and permanent-fix cultures across product engineering. **Detailed Job Responsibility & Accountabilities** - Own the end-to-end SRE charter: SLIs, SLOs, error budgets, alerting standards, incident response, change risk reviews and production readiness reviews. - Drive the observability maturity plan: centralized logging, golden signals (latency, traffic, errors, saturation), distributed tracing, trace-ID propagation and AIOps-led auto-remediation. - Lead 24x7 production incident command for Sev-1 and Sev-2 events impacting EMR, HIS, billing, and patient-facing systems; coordinate across clinical, business and vendor teams. - Establish and govern the error budget policy and change freeze protocols; publish monthly reliability scorecards to engineering leadership. - Mentor and grow a team of SRE engineers operating in three-shift plus reliever coverage model; build rotation schedules, runbooks and escalation matrices. - Partner with Product Engineering, DevOps, Security, Infrastructure and Cloud teams to instrument new services and retrofit observability into legacy MedMantra modules. - Drive chaos engineering, game-day simulations and disaster recovery rehearsals across critical clinical platforms. - Lead post-incident reviews with blameless root cause analysis, track corrective actions to closure and publish learnings across engineering. - Own the SRE tooling landscape — Dynatrace, Azure Monitor, Log Analytics, ManageEngine, PagerDuty/Opsgenie, and emerging AIOps platforms. - Define production readiness checklists and gatekeep releases from a reliability and capacity standpoint. - Report reliability KPIs to the CIO and represent IT in clinical business continuity forums. **Desired Skills & Competencies** - Must have 12+ years in software engineering / operations with 5+ years in a dedicated SRE leadership role. - Must have deep expertise in SLI/SLO engineering, error budget practice and Google SRE principles. - Must have hands-on proficiency with Dynatrace, Prometheus/Grafana, ELK, Azure Monitor, distributed tracing (OpenTelemetry/Jaeger). - Must have strong scripting and automation skills (Python, Go, Bash) and familiarity with Terraform/Ansible. - Must have proven experience running 24x7 production operations for mission-critical consumer-scale or regulated systems. - Must have experience with incident command (ICS), postmortem practice and reliability engineering metrics. - Preferred exposure to AIOps/ML-based anomaly detection and self-healing platforms. - Healthcare or life-critical systems experience strongly preferred. **Job Performance Parameters** - Availability SLO attainment for clinical and revenue systems (target 99.95%+) - Mean Time To Detect (MTTD) and Mean Time To Recover (MTTR) - % Sev-1/Sev-2 incidents detected before user impact - Error budget burn rate adherence across product lines - % critical transactions with end-to-end distributed tracing coverage - % incidents auto-resolved through self-healing runbooks **Managerial and Behavioral Competencies** - Crisis leadership and calm under pressure - Systems thinking and root cause rigour - Effective Communication with technical and business audiences - Team Building, Coaching and Roster Management for 24x7 operations - Long term effective decision making