Site Reliability Engineer (SRE)
Persistent Systems · State of Mahārāshtra, India
Persistent Systems · State of Mahārāshtra, India
**Job Description** **About Persistent** We are an AI-led, platform-driven Digital Engineering and Enterprise Modernization partner, combining deep technical expertise and industry experience to help our clients anticipate what’s next. Our offerings and proven solutions create a unique competitive advantage for our clients by giving them the power to see beyond and rise above. We work with many industry-leading organizations across the world, including 20 Fortune 50 companies and 4 of the 5 top banks in both the US and India, and numerous innovators across the healthcare ecosystem. Our disruptor’s mindset, commitment to client success, and agility to thrive in the dynamic environment have enabled us to sustain our growth momentum. Persistent has been recognized across top industry platforms for innovation, leadership, and inclusion. We reported $1,654.4M FY26 revenue with 17.4% Y-o-Y growth. We have delivered 24 sequential quarters of growth with $436.0M in Q4 FY26 revenue, up 3.2% Q-o-Q and 16.2% Y-o-Y growth. Our 27,500+ global team members, located in 18 countries, have been instrumental in helping the market leaders transform their industries. We have been recognized as the Fastest Growing IT Services Brand Globally in the 2026 Brand Finance IT Services 25 Report. We named a Leader in the Everest Group Private Equity (PE) Services PEAK Matrix*®* Assessment 2026 and Software Product Engineering PEAK Matrix*®* Assessment 2026. **About Position:** We are seeking highly motivated Site Reliability Engineers (SREs) to ensure the availability, reliability, scalability, and performance of our SaaS production environments. The ideal candidate will have hands-on experience with cloud platforms, Kubernetes, Infrastructures Code (IaC), automation, monitoring, and incident management. As part of the SRE team, you will act as the first responder for production incidents, execute runbooks, support deployments, automate operational tasks, and contribute to continuous improvements in platform reliability and operational excellence. - **Role: SRE Engineer** - **Location: All Persistent Location** - **Experience: 3 to 7 years** - **Job Type: Full-Time Employment** **What You'll Do:** - Monitor production environments using Datadog, PagerDuty, Grafana, and Prometheus. - Act as the first responder for alerts and incidents. - Acknowledge Priority-1 (P1) alerts within 5 minutes and initiate response within 15 minutes. - Execute incident response runbooks and follow escalation procedures. - Participate in on-call rotation and 24x7 operational support. - Perform root cause analysis (RCA) and contribute to post-incident reviews. - Provision, manage, and optimize cloud infrastructure on AWS and/or GCP. - Manage Kubernetes clusters (EKS/GKE) and related cloud-native services. - Configure networking components, security policies, IAM roles, DNS, load balancers, and storage services. - Monitor cloud utilization and support cost optimization initiatives. - Execute SaaS application deployments using CI/CD pipelines. - Manage Kubernetes deployments using Helm Charts and GitOps practices. - Validate release quality and support production rollouts. - Collaborate with development teams to ensure smooth application releases. - Develop and maintain Terraform modules and infrastructure automation. - Implement configuration management using Ansible. - Build automation scripts using Python and Bash to eliminate repetitive operational tasks. - Improve operational efficiency through tooling and process automation. - Configure dashboards, alerting rules, and service monitoring. - Implement metrics, logs, and traces using observability platforms. - Maintain monitoring standards aligned with SLIs, SLOs, and SLAs. - Support proactive capacity planning and performance tuning. - Execute and validate backup and recovery processes using Veeam, AWS Backup, or GCP Snapshots. - Support disaster recovery testing and business continuity initiatives. - Ensure platform reliability, availability, and recovery readiness. - Create and maintain runbooks, standard operating procedures (SOPs), and knowledge base articles. - Contribute at least one knowledge article or operational improvement document per month. - Participate in shift handover meetings and operational reviews. **Expertise You'll Bring:** - Strong troubleshooting and analytical skills. - Experience working in production support and on-call environments. - Good understanding of SLI, SLO, and SLA concepts. - Ability to work under pressure during critical incidents. - Strong communication and collaboration skills. - Self-driven with a continuous improvement mindset. - Experience working in Agile, DevOps, or SRE teams. **Benefits:** - Competitive salary and benefits package - Culture focused on talent development with quarterly growth opportunities and company-sponsored higher education and certifications - Opportunity to work with cutting-edge technologies - Employee engagement initiatives such as project parties, flexible work hours, and Long Service awards - Annual health check-ups - Insurance coverage: group term life, personal accident, and Mediclaim hospitalization for self, spouse, two children, and parents **Values-Driven, People-Centric & Inclusive Work Environment:** Persistent is dedicated to fostering diversity and inclusion in the workplace. We invite applications from all qualified individuals, including those with disabilities, and regardless of gender or gender preference. We welcome diverse candidates from all backgrounds. - We support hybrid work and flexible hours to fit diverse lifestyles. - Our office is accessibility-friendly, with ergonomic setups and assistive technologies to support employees with physical disabilities. - If you are a person with disabilities and have specific requirements, please inform us during the application process or at any time during your employment **Let’s unleash your full potential at Persistent -** *