Search 100,000+ live jobs across India

Free to search · AI fit score against your CV · tailor your résumé in one click

Job description

Job Title: PagerDuty Administrator / Incident Management Engineer Experience: 6-10Years Interview Mode: Virtual Interview Date: Work Location: Chennai We are hiring an experienced PagerDuty Administrator / Incident Management Engineer with strong expertise in incident management, alerting platforms, monitoring integrations, and IT operations. The ideal candidate will be responsible for configuring and managing PagerDuty, optimizing alert workflows, driving incident response processes, and collaborating with cross-functional teams to ensure high service availability and operational excellence. Must Have: • strong hands-on experience with PagerDuty administration and configuration. • Expertise in Incident Management processes and ITIL-based operations. • Experience integrating monitoring and observability platforms. • Good understanding of alert management, escalation workflows, and operational best practices. • Experience with REST APIs, Webhooks, and third-party tool integrations. • Knowledge of Linux and Windows operating systems. • Experience working in cloud environments including AWS, Azure, or GCP. Good-to-Have Skills • Experience in SRE or DevOps environments. • Scripting knowledge in Python, Shell Scripting, or PowerShell. • Understanding of CI/CD pipelines and deployment automation. • Experience with Infrastructure as Code tools such as Terraform and Ansible. • Exposure to ChatOps integrations using Slack and Microsoft Teams. Key Responsibilities • Configure, manage, and maintain the PagerDuty platform for enterprise incident management. • Design and administer escalation policies, on-call schedules, notification rules, and event routing. • Integrate PagerDuty with monitoring and observability tools such as Dynatrace, Splunk, Datadog, Prometheus, CloudWatch, and SevOne. • Reduce alert fatigue through alert deduplication, suppression, threshold tuning, and event correlation. • Manage the complete incident lifecycle from detection, response, resolution, post-incident review, and closure. • Collaborate with DevOps, SRE, NOC, Infrastructure, and Application teams during major incidents. • Create and maintain runbooks, operational playbooks, and incident response procedures. • Conduct Post Incident Reviews (PIR), Root Cause Analysis (RCA), and continuous service improvement activities. Mandatory Skills • PagerDuty Administration • Incident Management • IT Operations • Alerting & Escalation Management • Monitoring Tools (Dynatrace, Splunk, Nagios, Datadog, Prometheus, CloudWatch) • REST API Integrations • Webhooks

More jobs at Tata Consultancy Services

All Tata Consultancy Services jobs (10,413)

Operations jobs in Chennai

Operations jobs in Chennai (1,412)

Other Operations jobs in India

All Operations jobs in India (18,226)