Software Engineering Specialist
BT Group · State of Karnataka, India
BT Group · State of Karnataka, India
Req ID: 61049 Job Function: Software Engineering Posting Start Date: 02/08/2026 Posting End Date: 21/08/2026 Division: Digital Job Location: IND-Bengaluru-RMZ Ecoworld Advertised Salary: Competitive Job Req ID: 61049 Posting Date: 21-August-2026 Function: SRE Location: Bengaluru Salary: Competitive ## **About the role** - The Physical Network Inventory team are responsible for maintaining the GIS inventory systems that are used to plan and build the Openreach network. - These systems have been fundamental in supporting the Openreach rollout of full fibre to the UK. - Site Reliability Engineering (SRE) role demands software engineering applied to infrastructure operations, making sure platforms remain fast, reliable, and scalable - UX role demands human-centred design of products, making sure interfaces are intuitive, accessible, and efficient for the end use ## **What you’ll be doing** We are seeking a highly motivated hands-on SRE & UX Lead who can bridge the gap between platform reliability and user experience. This role requires a unique combination of engineering, operational excellence, and user-centred thinking. The candidate will ensure enterprise platforms are reliable, scalable, observable, secure, intuitive, efficient, accessible, and easy to operate for users, support teams, engineers, and business stakeholders. The role will contribute across SRE practices, production operations, automation, observability, incident management, UX strategy, workflow optimisation, dashboards, design systems, and continuous improvement of user journeys. **Service Reliability, Availability & Operational Excellence** - Lead reliability engineering practices across applications, platforms, and infrastructure. - Define, implement, and monitor SLIs, SLOs, error budgets, uptime targets, and operational KPIs. - Own production stability, performance, resilience, service availability, and operational maturity. - Drive production readiness reviews, reliability assessments, capacity planning, and resilience engineering. - Drive proactive service improvement initiatives to improve availability, performance, and reliability. - Provide expert third-line technical support for critical business applications and platforms. - Lead Major Incident resolution, ensuring rapid service restoration and minimal business impact. - Conduct blameless post-incident reviews, root cause analysis, and corrective/preventive actions. - Identify recurring issues and drive permanent remediation through engineering solutions. In-Life Service Management **DevOps & CI/CD Engineering** - Design, implement, and maintain CI/CD pipelines for reliable and automated software delivery. - Improve deployment quality through automated testing, validation, release controls, and configuration management. - Drive Infrastructure as Code practices and promote DevOps best practices across development and support teams. **Observability & Monitoring** - Implement and maintain enterprise observability solutions, dashboards, alerting, synthetic monitoring, and performance analytics. - Utilize platforms such as Dynatrace, Grafana, Prometheus, ELK, Splunk, or equivalent tools. - Analyse performance metrics, reliability trends, KPIs, and service health data to proactively prevent incidents. **Continuous Improvement** - Champion SRE, automation, observability, reliability engineering, and platform engineering principles. - Contribute to reliability roadmaps, architectural improvements, and operational maturity uplift. - Mentor team members on DevOps and SRE best practices and foster a culture of continuous learning and innovation. **UX Strategy for Enterprise Platforms** - Lead UX strategy for complex enterprise platforms, operational tools, dashboards, and workflow-based systems. - Translate user needs, business goals, and technical constraints into intuitive user experiences. - Own user journeys, task flows, wireframes, prototypes, workflow maps, and usability improvements. **Design Systems, Accessibility & UI Collaboration** - Contribute to design systems, reusable UI patterns, dashboard standards, and component libraries. - Work closely with UI designers and frontend teams to ensure consistent, accessible implementation across monitoring tools, admin portals, operational dashboards, and user-facing applications. - Provide design specifications, interaction guidelines, acceptance criteria, and usability recommendations. ## **Essential Skills / Experience** **Essential skills** - 10+ years of experience across SRE, DevOps, platform engineering, application support, production support, product engineering, UX design, enterprise application design, or operational tooling. - Experience managing mission-critical enterprise applications and services in complex production environments. - Strong understanding of SRE principles, including SLIs, SLOs, error budgets, observability, security, production readiness, and operational support processes. - Hands-on experience with cloud platforms such as AWS, Azure, or GCP. - Strong experience with Linux/Unix systems, networking concepts, DNS, load balancers, certificates, firewalls, API gateways, middleware, APIs, databases, and cloud-native services. - Experience with Kubernetes, Docker, CI/CD platforms, infrastructure as code, configuration management, and automation tooling. - Proficiency with monitoring and observability tools such as Dynatrace, Prometheus, Grafana, ELK/EFK, Splunk, Datadog, AppDynamics, New Relic, or equivalent. - Strong scripting or programming experience in Python, Shell, PowerShell, Java, JavaScript, React/Angular, CSS, or similar technologies. - Experience with production troubleshooting, RCA, incident, problem, change, release management, and service improvement practices. - Strong UX capability in user research, journey mapping, wireframing, prototyping, usability testing, information architecture, workflow design, and service design. - Experien