Senior Manager - IT Systems Engineering
Oliver Wyman · Gurugram, Haryana, India - Mumbai, Maharashtra, India
Oliver Wyman · Gurugram, Haryana, India - Mumbai, Maharashtra, India
Description: We are seeking a talented individual to join our GIS at Marsh. This role will be based in Mumbai. This is a hybrid role that has a requirement of working at least three days a week in the office. **Senior Manager Site Reliability Engineering (SRE) / Reliability Operations** **Role Overview** We are looking for a Senior Manager SRE / Reliability Operations to drive reliability engineering practices across critical services and platforms. In this role, you will establish and mature SLI/SLO/SLA and error budget governance, lead operational excellence (ITIL-aligned incident/problem/change), and partner closely with engineering, architecture, and DevOps teams to reduce toil through automation and improve resilience, observability, and release reliability. **We will count on you to:** - **SRE fitness: SLI/SLO/SLA error budget management** - Define, implement, and continuously improve SLIs and SLOs for services/products in collaboration with product and engineering teams - Align SLAs with business expectations and ensure operational commitments are measurable and reportable - Create and manage error budgets, including: - Error budget policies and burn-rate thresholds - Release gating recommendations based on error budget status - Executive reporting on reliability posture and trade-offs - Build SLO dashboards and reliability scorecards for leadership and stakeholders - Conduct periodic SLO reviews and drive corrective actions when reliability trends degrade - **Software engineering, automation, and toil reduction** - Identify operational toil and repetitive manual work; design and deliver automation to reduce or eliminate it - Develop tooling/scripts/services using standard languages (e.g., Python, Go, Java, PowerShell, or similar) aligned to the team s stack - Implement self-healing patterns (automated remediation, auto-rollback, auto-scaling, safe retries) - Standardize operational runbooks and embed automation into runbooks wherever possible - Improve operational efficiency through event-driven workflows and platform capabilities - **System design: reliability, resiliency, and stability** - Partner with architecture and engineering teams to design systems for: - High availability, fault tolerance, and graceful degradation - Scalability and performance under load - Robust dependency management and timeout/retry/circuit-breaker strategies - Lead or contribute to resiliency reviews, failure-mode analysis, and capacity planning - Design and facilitate resiliency testing (dependency failure simulations, controlled fault injection where appropriate) - Define and validate RTO/RPO targets and align with disaster recovery strategies - **Cloud adoption reliability enablement** - Support cloud adoption/migration initiatives (AWS/Azure/GCP as applicable), ensuring reliability-by-design - Establish operational patterns for cloud-native services including: - Auto-scaling, health checks, multi-AZ/region strategies - Infrastructure-as-Code (IaC) practices in partnership with platform/DevOps teams - Contribute to secure, compliant, and standardized cloud operations (guardrails, baselines, tagging/monitoring standards) - **Operational management (ITIL-aligned)** - Own and/or drive improvements in: - **Incident Management:** triage, mitigation, communications, major incident handling, post-incident reviews - **Problem Management:** root cause analysis, trend analysis, corrective/preventive actions, known error database hygiene - **Change Management:** change risk assessment, change validation, release readiness, change success metrics - Facilitate and standardize post-incident reviews (PIRs) focusing on blameless learning, actionable follow-ups, and prevention - Improve operational governance and documentation quality (runbooks, SOPs, service ownership, on-call readiness) - **Observability monitoring (Datadog)** - Build and maintain observability practices using Datadog, including metrics, logs, traces (APM), synthetics, and RUM (as applicable) - Establish dashboard standards for service health, SLO compliance, and operational KPIs - Design alert strategies focused on actionable alerts, noise reduction, correlation, and routing - Implement monitoring for golden signals (latency, traffic, errors, saturation) and service-specific signals - Tune alerts based on incident learnings and error budget burn rates; reduce false positives - **DevOps CI/CD reliability** - Collaborate with DevOps teams to improve CI/CD pipelines for reliability and speed, including: - Automated testing strategy integration (unit/integration/smoke) - Deployment strategies (blue/green, canary, rolling, feature flags) - Automated rollback and deployment verification - Improve release reliability and change failure rate through quality gates and operational readiness checks - Promote shift-left reliability: testing, observability instrumentation, and operational requirements embedded early in the SDLC **Deliverables / Outcomes (What success looks like)** - Defined SLIs/SLOs and error budgets for critical services with leadership-ready reporting - Measurable reduction in toil and faster operational response through automation - Improved reliability KPIs (availability, MTTR, incident volume/severity, change failure rate) - Mature Datadog dashboards and alerting with reduced noise and improved signal quality - Standardized incident/problem/change practices with consistent PIR execution and follow-through - Increased platform resiliency validated via reviews and targeted resiliency testing **What you need to have:** - 7+ years (or appropriate level) experience in SRE / Production Operations / Platform Engineering / DevOps supporting enterprise applications - Strong knowledge of: - SLA/SLO/SLI concepts and practical error budget implementation - Incident response, troubleshooting, and root cause analys