Senior Site Reliability Engineer
NCR Voyix · Hyderabad, Telangana, India
NCR Voyix · Hyderabad, Telangana, India
Key Responsibilities Design and implement enterprise observability solutions across Azure, GCP, Kubernetes, and hybrid environments. Develop monitoring, logging, tracing, and telemetry standards using industry best practices. Build enterprise dashboards that provide real-time infrastructure, application, and customer health visibility. Define and implement SLIs, SLOs, error budgets, and operational health metrics. Improve proactive detection, alert quality, and incident response through automation and intelligent alerting. Integrate observability with ServiceNow, automation platforms, and operational workflows. Partner with Product Engineering, Infrastructure, Security, and Operations teams to improve platform reliability and operational readiness. Support enterprise initiatives involving AI-driven observability, event correlation, and operational analytics. Required Skills Site Reliability Engineering (SRE) Kubernetes (AKS/GKE) Azure and Google Cloud Platform Grafana, Datadog, Prometheus, OpenTelemetry (or similar) Monitoring, logging, distributed tracing, and telemetry Infrastructure as Code (Terraform) Python, Go, or PowerShell automation CI/CD and DevOps practices