Senior Manager - IT Systems Engineering
Oliver Wyman · Gurugram, Haryana, India - Mumbai, Maharashtra, India
Free to search · AI fit score against your CV · tailor your résumé in one click
Oliver Wyman · Gurugram, Haryana, India - Mumbai, Maharashtra, India
Description: We are seeking a talented individual to join our GIS at Marsh. This role will be based in Mumbai. This is a hybrid role that has a requirement of working at least three days a week in the office. Senior Manager Site Reliability Engineering (SRE) / Reliability Operations Role Overview We are looking for a Senior Manager SRE / Reliability Operations to drive reliability engineering practices across critical services and platforms. In this role, you will establish and mature SLI/SLO/SLA and error budget governance, lead operational excellence (ITIL-aligned incident/problem/change), and partner closely with engineering, architecture, and DevOps teams to reduce toil through automation and improve resilience, observability, and release reliability. We will count on you to: • SRE fitness: SLI/SLO/SLA error budget management • Define, implement, and continuously improve SLIs and SLOs for services/products in collaboration with product and engineering teams • Align SLAs with business expectations and ensure operational commitments are measurable and reportable • Create and manage error budgets, including: • Error budget policies and burn-rate thresholds • Release gating recommendations based on error budget status • Executive reporting on reliability posture and trade-offs • Build SLO dashboards and reliability scorecards for leadership and stakeholders • Conduct periodic SLO reviews and drive corrective actions when reliability trends degrade • Software engineering, automation, and toil reduction • Identify operational toil and repetitive manual work; design and deliver automation to reduce or eliminate it • Develop tooling/scripts/services using standard languages (e.g., Python, Go, Java, PowerShell, or similar) aligned to the team s stack • Implement self-healing patterns (automated remediation, auto-rollback, auto-scaling, safe retries) • Standardize operational runbooks and embed automation into runbooks wherever possible • Improve operational efficiency through event-driven workflows and platform capabilities • System design: reliability, resiliency, and stability • Partner with architecture and engineering teams to design systems for: • High availability, fault tolerance, and graceful degradation • Scalability and performance under load • Robust dependency management and timeout/retry/circuit-breaker strategies • Lead or contribute to resiliency reviews, failure-mode analysis, and capacity planning • Design and facilitate resiliency testing (dependency failure simulations, controlled fault injection where appropriate) • Define and validate RTO/RPO targets and align with disaster recovery strategies • Cloud adoption reliability enablement • Support cloud adoption/migration initiatives (AWS/Azure/GCP as applicable), ensuring reliability-by-design • Establish operational patterns for cloud-native services including: • Auto-scaling, health checks, multi-AZ/region strategies • Infrastructure-as-Code (IaC) practices in partnership with platform/DevOps teams • Contribute to secure, compliant, and standardized cloud operations (guardrails, baselines, tagging/monitoring standards) • Operational management (ITIL-aligned) • Own and/or drive improvements in: • Incident Management: triage, mitigation, communications, major incident handling, post-incident reviews • Problem Management: root cause analysis, trend analysis, corrective/preventive actions, known error database hygiene • Change Management: change risk assessment, change validation, release readiness, change success metrics • Facilitate and standardize post-incident reviews (PIRs) focusing on blameless learning, actionable follow-ups, and prevention • Improve operational governance and documentation quality (runbooks, SOPs, service ownership, on-call readiness) • Observability monitoring (Datadog) • Build and maintain observability practices using Datadog, including metrics, logs, traces (APM), synthetics, and RUM (as applicable) • Establish dashboard standards for service health, SLO compliance, and operational KPIs • Design alert strategies focused on actionable alerts, noise reduction, correlation, and routing • Implement monitoring for golden signals (latency, traffic, errors, saturation) and service-specific signals • Tune alerts based on incident learnings and error budget burn rates; reduce false positives • DevOps CI/CD reliability • Collaborate with DevOps teams to improve CI/CD pipelines for reliability and speed, including: • Automated testing strategy integration (unit/integration/smoke) • Deployment strategies (blue/green, canary, rolling, feature flags) • Automated rollback and deployment verification • Improve release reliability and change failure rate through quality gates and operational readiness checks • Promote shift-left reliability: testing, observability instrumentation, and operational requirements embedded early in the SDLC Deliverables / Outcomes (What success looks like) • Defined SLIs/SLOs and error budgets for critical services with leadership-ready reporting • Measurable reduction in toil and faster operational response through automation • Improved reliability KPIs (availability, MTTR, incident volume/severity, change failure rate) • Mature Datadog dashboards and alerting with reduced noise and improved signal quality • Standardized incident/problem/change practices with consistent PIR execution and follow-through • Increased platform resiliency validated via reviews and targeted resiliency testing What you need to have: • 7+ years (or appropriate level) experience in SRE / Production Operations / Platform Engineering / DevOps supporting enterprise applications • Strong knowledge of: • SLA/SLO/SLI concepts and practical error budget implementation • Incident response, troubleshooting, and root cause analys