Senior SRE Engineer
EPAM Systems · Chennai, Tamil Nadu, India
EPAM Systems · Chennai, Tamil Nadu, India
We are looking for a **Senior SRE Engineer** to join our collaborative team. Our team is building the platform fueling digital transformation. We take a cloud-first approach to deliver customer-centric experiences and platform web services that are used by product teams and partners. As a Senior SRE Engineer, you'll be a vital part of the Engagement Team for Digital Transformation. This team enables and aids product teams and partners by adopting and integrating cloud services with a customer-centric ideology always in mind. You will be an expert in cloud services, building, proving out, and communicating the value of the platform that is enabling Digital Transformation. **Responsibilities** - Design and implement high-availability and disaster recovery strategies for critical workloads - Build, maintain, and evolve cloud infrastructure using Infrastructure-as-Code practices - Develop and manage CI/CD pipelines to enable reliable and efficient delivery - Implement and maintain monitoring, observability, and alerting across the platform - Lead cross-functional reliability initiatives and platform-wide automation projects - Ensure networking, security, and identity/access management best practices in the cloud - Collaborate with product teams and partners to adopt and integrate cloud services - Optimize infrastructure costs by applying FinOps practices and cost optimization frameworks - Troubleshoot production issues, perform root cause analysis, and drive improvements - Mentor engineers and influence reliability standards across teams **Requirements** - 5+ years of experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles - Deep AWS expertise (EC2, S3, RDS, IAM, VPC, Lambda, CloudFormation/Terraform, etc.) - Strong knowledge of Infrastructure-as-Code (IaC) using Terraform, AWS CDK, or CloudFormation - Proven experience with CI/CD tools such as Jenkins, GitHub Actions, GitLab CI, or similar - Proficiency in containerization and orchestration with Docker, Kubernetes, and ECS or EKS - Expertise in monitoring and observability tools (Datadog, New Relic, Prometheus, Grafana, ELK, CloudWatch, etc.) - Strong scripting or programming competency in Python, Bash, or Go - Sound understanding of networking, security, and identity/access management in the cloud - Experience designing high-availability and disaster recovery strategies for critical workloads - Excellent communication, problem-solving, and leadership skills with the ability to influence across teams **Nice to have** - AWS or other Cloud Certification (Solutions Architect, DevOps Engineer, etc.) - Experience with AIOps, Serverless Architectures, and event-driven systems - Familiarity with FinOps practices and cost optimization frameworks - Experience with SaaS monitoring tools such as Sumo Logic, PagerDuty, or similar - Exposure to Atlassian tools (Jira, Confluence, Bitbucket) - Familiarity with SQL/NoSQL databases