T

Site Relibility Enginneer (SRE)

Tata Consultancy Services · Bengaluru, Karnataka, India

3–9 yrs experiencefull_timePosted 1w ago

Job description

**Job Location** **: Bangalore,Hyderabad** **Experience - 6+ Only** **Job Requirements** - Drive **operational stability** across ERP and other finance applications - Implement and enhance **automation and tooling** to reduce manual effort and improve efficiency - Own and execute **Disaster Recovery (DR) planning and testing** to ensure business continuity - Lead **service design and service transition** activities for new and existing systems - Manage **incident, problem, and change processes** aligned with SRE and ITIL practices - Establish effective **service communication frameworks** for incidents and outages - Drive **continual service improvement (CSI)** initiatives across finance systems - Define and manage **service metrics, SLAs, SLIs, and reporting dashboards** - Collaborate with engineering teams to improve **system reliability, observability, and performance** - Ensure smooth onboarding and ownership of **satellite applications within the finance ecosystem** - Proactively identify risks and implement preventive measures to minimize production issues **Preferred Skills and Experience** - Strong experience in Site Reliability Engineering / Production Support / DevOps roles - Experience supporting Oracle Cloud ERP or similar enterprise SaaS platforms - Knowledge of monitoring, alerting, and observability tools (e.g., Splunk, Grafana, OCI monitoring) - Experience in incident management, RCA, and problem management - Exposure to automation frameworks and scripting (e.g., Python, Shell, Terraform) - Understanding of cloud platforms (OCI/AWS/Azure) and distributed systems - Familiarity with ITIL processes and service management frameworks - Experience working in global, distributed teams - Exposure to financial systems and processes is an advantage **Key Responsibilities:** - Operational Stability & SRE Practices - Maintain high system availability through proactive monitoring and incident prevention - Define and track SLIs/SLOs to measure service health and reliability - Lead root cause analysis (RCA) and implement preventive fixes - Automation & Tooling - Build and maintain automation for repetitive operational tasks - Improve deployment pipelines and operational workflows - Develop scripts/tools to enhance productivity and reduce human intervention - Service Design & Transition - Ensure new services are designed with reliability, scalability, and supportability in mind - Lead service transition activities, including documentation, readiness, and handover - Disaster Recovery & Resilience - Develop and execute DR strategies and testing plans - Ensure systems meet recovery objectives (RTO/RPO) - Service Management & Communication - Manage incident, change, and problem processes - Ensure clear and timely communication during production incidents - Collaborate with stakeholders during outages and recovery - Continual Service Improvement (CSI) - Identify and implement improvements to systems and processes - Drive reliability engineering best practices across teams - Service Metrics & Reporting - Define and maintain dashboards for system health and performance - Provide insights and reporting to leadership for decision-making - Satellite Application Ownership - Own end-to-end support for finance-related satellite applications - Ensure integration stability between ERP and downstream/upstream systems