Forward Deployment Engineer (SRE)
PwC Acceleration Centers · Hyderabad, Telangana, India
Free to search · AI fit score against your CV · tailor your résumé in one click
PwC Acceleration Centers · Hyderabad, Telangana, India
Senior Associate – Forward Deployment Engineer (SRE) Site Reliability Engineering | Forward Deployed Engineering Location: Bangalore / Hyderabad Experience Required 6–9 years. Job Summary A senior, customer-facing SRE who owns the reliability of a client's production systems end to end. Embedded as the main technical point of contact, you will design reliable infrastructure, lead incidents, apply AI-assisted operations, and mentor the team. Key Responsibilities • Own the reliability of production systems for one or more enterprise customers on AWS. • Define and manage SLIs, SLOs, and error budgets, and ensure monitoring provides meaningful insight. • Lead the response to serious (P0/P1) incidents and run blameless post-mortems that lead to fixes. • Lead within the on-call rotation. • Operate and optimise Kubernetes clusters and AWS services as workloads grow. • Introduce AI SRE tooling and AIOps where it speeds up triage and resolution, with appropriate guardrails. • Advise customers on improving their reliability practices, and mentor associate engineers. • Share field learnings with product and engineering teams. Required Qualifications • Substantial SRE experience with real ownership of production reliability on AWS. • Experience running Kubernetes and AWS services at scale. • A track record of defining and managing SLIs, SLOs, and error budgets. • Strong observability practice. • Proven incident leadership and post-mortem facilitation. • Strong automation skills (Python, Go, or Bash). • Hands-on understanding of AI-assisted operations, including introducing AI SRE tooling with sensible guardrails. • AWS certification at Associate level as a minimum. Preferred Qualifications • Owning CI/CD pipelines and release automation. • Writing and maintaining Terraform modules. • GitOps workflows and Helm. • Designing telemetry collection across services. • Incident-management tooling such as PagerDuty. • Hands-on experience integrating AIOps or AI SRE tooling into production operations. • Familiarity with observability or evaluation of AI systems (e.g., Langfuse). • Experience delivering AI-based work or AI/ML-driven initiatives in production. • AWS Certified Solutions Architect – Professional and/or AWS Certified DevOps Engineer – Professional. Technical Skills & Tools • Cloud (AWS): EC2, EKS, ECS, Lambda, S3, RDS, VPC, IAM, CloudWatch, ELB/ALB, Route 53 • Containers & orchestration: Docker, Kubernetes (EKS), Helm • Observability & monitoring: Prometheus, Grafana, Datadog, OpenTelemetry, Loki, Tempo/Jaeger (tracing), ELK/Elastic • Reliability practices: SLIs/SLOs, error budgets, incident command, capacity planning, failure-mode analysis • Incident & on-call: PagerDuty, Opsgenie, blameless post-mortems • Automation & scripting: Python, Go, Bash, Git • AI-assisted operations: AI SRE tooling, AIOps, anomaly detection, automated RCA • DevOps (good to have): CI/CD (GitHub Actions, GitLab CI), Terraform modules, GitOps (ArgoCD), Helm Soft Skills & Competencies • Confident, customer-facing communication. • Calm, decisive incident leadership. • Mentors and raises standards across the team. • Sound technical judgement.