Site Recovery Manager
NCR Voyix · Chennai, Tamil Nadu, India
Free to search · AI fit score against your CV · tailor your résumé in one click
NCR Voyix · Chennai, Tamil Nadu, India
Reliability & Service Availability • Monitor enterprise applications, infrastructure, cloud platforms, and customer-facing services. • Ensure platform availability, performance, and service health within established SLAs and SLOs. • Identify service degradation trends and proactively address reliability risks. • Drive operational improvements to reduce incidents and improve system resiliency. • Establish and track reliability metrics, KPIs, and service health indicators. • Develop preventative measures to reduce future service disruptions. Observability & Monitoring • Configure and maintain monitoring platforms including: • AppDynamics • Dynatrace • Datadog • Splunk • Azure Monitor • Google Cloud Operations • ServiceNow Event Management • Develop dashboards, alerts, and health monitoring solutions. • Reduce alert fatigue through alert tuning and optimization. Automation & Engineering • Automate operational processes and repetitive tasks. • Develop scripts and tooling using: • PowerShell • Python • Bash • APIs • Improve operational efficiency through self-healing and automated remediation capabilities. • Support infrastructure-as-code and reliability engineering initiatives. Operational Readiness • Participate in Operational Readiness Reviews (ORR). • Validate monitoring, alerting, runbooks, and support procedures prior to production go-live. • Ensure escalation paths and support models are documented and operational. • Review application deployments for supportability and operational risks. Cloud & Infrastructure Support • Support hybrid environments across: • Azure • Google Cloud Platform (GCP) • AWS • VMware • Analyze application, network, database, and infrastructure performance issues. • Work with engineering teams to optimize platform stability and scalability. Governance & Reporting • Produce incident reports, service health updates, operational reviews, and executive summaries. • Maintain operational documentation, runbooks, and knowledge articles. • Track service performance metrics and reliability improvements. • Support audit and compliance initiatives as required. Required Qualifications • 6+ years of experience in: • Site Reliability Engineering • Production Support • Systems Engineering • DevOps • Network Operations Center (NOC) • Command Center Operations • Experience supporting mission-critical production environments. • Strong understanding of: • Incident Management • Problem Management • Change Management • Service Level Management • Experience with ServiceNow or similar ITSM platforms. Technical Skills Operating Systems • Windows Server • Linux/Unix Cloud Platforms • Microsoft Azure • Google Cloud Platform (GCP) • Amazon Web Services (AWS) Monitoring & Observability • AppDynamics • Datadog • Dynatrace • New Relic • Splunk • Azure Monitor • ServiceNow Event Management Preferred Qualifications • Experience working in a Global Command Center environment. • AWS, Azure, or GCP certifications. • Experience supporting retail, hospitality, payments, or enterprise SaaS platforms. • Experience with CI/CD pipelines and DevOps practices. • Knowledge of SRE concepts including: • SLI/SLO/SLA management • Error budgets • Chaos testing • Resiliency engineering Key Competencies • Critical Incident Leadership • Technical Troubleshooting • Problem Solving • Customer Focus • Operational Excellence • Communication Skills • Executive Presence • Collaboration • Continuous Improvement • Decision Making Under Pressure Success Metrics The GCC Site Reliability Engineer will be measured on: • Service availability and uptime • Incident response times • Mean Time to Detect (MTTD) • Mean Time to Restore (MTTR) • Reduction in recurring incidents • Monitoring effectiveness • Automation adoption • Operational readiness compliance • Customer impact reduction • Service reliability improvements Work Environment • 24x7 operational support organization. • Participation in on-call and major incident rotations. • Collaboration with global teams across multiple regions. • Hybrid cloud and enterprise production environments. • Fast-paced, mission-critical operational setting.