Site Reliability Engineer
Ascendion · Chennai, Tamil Nadu, India
Free to search · AI fit score against your CV · tailor your résumé in one click
Ascendion · Chennai, Tamil Nadu, India
Key Responsibilities • Design, build, and operate highly available cloud infrastructure on AWS and/or Azure. • Drive SRE practices including SLIs, SLOs, SLAs, Error Budgets, Incident Management, and RCA. • Manage and optimize Kubernetes platforms (EKS/AKS) in production environments. • Implement Infrastructure as Code (Terraform, CloudFormation, ARM/Bicep). • Build automation and self-healing solutions using Python, Go, Bash, or PowerShell. • Establish enterprise observability using Prometheus, Grafana, ELK, Datadog, Splunk, OpenTelemetry, etc. • Design and support CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI/CD, Azure DevOps, ArgoCD. • Perform capacity planning, performance tuning, disaster recovery, and cloud cost optimization. • Partner with Development, Security, and Platform teams to improve reliability and operational excellence. • Participate in on-call support, major incident management, and production troubleshooting. Mandatory Skills • Strong hands-on experience with AWS and/or Azure Cloud • Expertise in Kubernetes (EKS/AKS), Docker, Helm • Strong knowledge of Terraform and Infrastructure as Code • Experience with CI/CD, GitOps, and Automation • Hands-on with Monitoring, Logging, and Observability tools • Strong Linux, Networking, and Distributed Systems knowledge • Scripting/Programming experience in Python, Go, Bash, or PowerShell • Experience supporting large-scale, mission-critical production environments