Site Reliability Engineering Lead - BPL (based in Pune)
Barclays · Pune Division, Maharashtra, India
Barclays · Pune Division, Maharashtra, India
**Role Purpose** Barclaycard Payments Limited (BPL) is establishing a disciplined Technology Operations capability to support launch, early scale and future separation of new payments products and platforms in a modern, greenfield fintech environment. The Site Reliability Engineering Lead will be responsible for building the central SRE enablement capability from the ground up within Technology Operations, strengthening reliability, operability and resilience across launch-critical and business-critical services, while avoiding the creation of a separate operational support silo. This role can either be based in Pune, India or London, UK. The role will define, implement and embed reliability engineering standards, practices and operating rhythms from first principles, including observability, alert quality, service-level objectives, error-budget management, automation, toil reduction and post-incident learning. The role will operate in close partnership with Engineering, Product, ITSM, Incident Command, Controls and supplier teams to ensure that reliability accountability remains with the teams that build and operate the services, while Technology Operations provides the standards, assurance, insight and continuous-improvement mechanisms required to operate safely at scale. **Scope** The role provides SRE enablement across Technology Operations, with an initial emphasis on Q3 launch readiness for priority services, including Next Generation Gateway and Service Portal. The scope includes designing and establishing the foundational SRE capability, reliability standards, service telemetry, SLO / SLI definition, alert hygiene, operational readiness, runbook quality, automation opportunities, incident learning and reliability improvement planning across both internally managed and provider-delivered services. The role is not accountable for feature delivery, application support ownership or the operation of a traditional production support function. Engineering will remain accountable for service ownership and recovery under the “you build it, you run it” model; the SRE Lead will enable consistent reliability practice, assure readiness and drive systemic service improvement. **Key Responsibilities** **SRE Operating Model & Standards** - Design and establish the SRE enablement model for Technology Operations from the ground up, including role boundaries, engagement model, service onboarding criteria and interfaces with Engineering, ITSM, Incident Command and Controls. - Establish minimum reliability standards for a modern fintech operating model, including observability, alerting, ownership, runbooks, escalation paths, recovery patterns and service health reporting. - Establish SLO / SLI guidance, including how reliability metrics should be defined, measured, reviewed and incorporated into service decision-making. - Create scalable, repeatable ways of working that can mature from Day 0 launch readiness into an enduring SRE capability as the organisation, product estate and service volumes grow. **Observability, Alerting & Service Health** - Promote consistent adoption of golden signals, telemetry, dashboards and alert-quality standards across priority services. - Partner with Engineering and platform teams to reduce alert noise, improve signal quality and ensure actionable alerts are routed to accountable service owners. - Support the development of a consolidated operational view of service health across internal teams and external providers. **Incident Learning & Reliability Improvement** - Partner with Incident Command and Problem Management to ensure major incident reviews produce clear technical learning, root-cause themes and defined reliability improvement actions. - Identify recurring failure patterns and translate them into a prioritised reliability backlog, owned jointly with Engineering and service owners. - Embed evidence-based learning and measurable reduction in repeat incidents, failure demand and avoidable operational toil. **Automation, AIOps & Toil Reduction** - Identify automation opportunities across detection, diagnosis, response, recovery, evidence capture and operational reporting. - Work with tooling, platform and engineering teams to define AIOps use cases that improve signal correlation, reduce manual triage and accelerate controlled recovery. - Ensure automation is designed to support control, auditability and operational resilience requirements, without introducing unmanaged operational risk. **Operational Readiness & Launch Support** - Support readiness assessment for launch-critical services, confirming that ownership, telemetry, runbooks, recovery patterns and escalation paths are fit for go-live. - Provide targeted technical investigation and reliability support during launch, hypercare and early-life service operation. - Define Day 1+ maturity steps to transition the SRE capability from targeted launch support into a scalable, repeatable reliability model appropriate for a growing fintech organisation. **Key Deliverables** - SRE operating model and engagement approach for Technology Operations, including service onboarding criteria and boundaries with Engineering. - Greenfield SRE capability roadmap, setting out the foundational Day 0 controls, Day 1+ operating practices and longer-term maturity path for reliability engineering across the fintech platform estate. - Minimum reliability standards for launch-critical services, covering SLOs / SLIs, observability, alerting, runbooks, ownership and recovery expectations. - Initial service-health and reliability dashboard requirements, aligned to wider Tech Ops reporting and service review cadence. - Reliability improvement backlog for priority services, linked to incident themes, problem management and engineering remediation activity. - Automation and AIOps opportunity assessment, prioritised by operational risk reduction, toil reduction and recoverability benefit. - Day 0 / Day 1