Job description

**About The Team** TDO need to own any issues related to site availability and drive through the resolution and have the accountability to make necessary changes to fix issues and bring the sites backup online and make sure customer experience is seamless, execute recovery levers during outages, configuration mishaps, and DR situations. leverage technical experience to keep critical systems running through any event. The Technical Duty Officer (TDO) is responsible for the availability and performance of our global sites. The TDO will take command and control of Major Incidents focusing on restoration by identifying and coordinating with appropriate resources through all the phases of triage, restoration and validation. Technically you will understand the full end to end stack and use this knowledge to detect and lead a team through incident response. Excellent judgement is crucial as you will provide final approval on site changes and hold critical switches for functionality of the site. You will ensure all documentation surrounding the Major Incidents are accurate and communication with the leadership team is clear and complete. Your ability to continuously challenge yourself and develop a strong network with peers and stakeholders cross functionally will see you exceed in this role. Our goal is to protect the customer experience and deliver outstanding levels of availability. What You’ll Do Architecture Acumen: Requires knowledge of: Architectural principles; Systems and environment behavior; Architectural Styles, Patterns and plans; Architectural standards; Non-functional System performance parameters; Technology Strategy. Defect Management and Troubleshooting: Requires knowledge of: Defect life-cycle process, defect tracking tools and methodologies; Defect reporting; Regression testing; Root cause analysis; Root cause corrective action. To conduct root cause analysis (RCA) and root cause corrective action (RCCA) to identify the origin of defects/ performance gaps and prevent them from recurring. Track registered issues for the product/solution and prioritize them for resolution. Measure usability of the product/solution as per customer/business requirement after defect fixing and plugging test gaps. Analyze the issues and plan a series of steps which potentially includes reconfiguration, integration, removal or addition of application components to enhance the application's functionality, usability and security. DevOps Orientation: Requires knowledge of: Different operating systems; Software maintenance tools and techniques; Application monitoring tools and techniques; Debugging tools; Mock screen; Pseudocodes; Reverse Engineering; Traceability matrix; System performance, security, integration; Data migration and accessibility; Design Methodologies. Agentic AI framework—requires a shift from passive monitoring to active, autonomous oversight of digitalisation. The core requirement is managing "human-in-the-loop" systems where AI agents plan and execute tasks, while the TDO ensures alignment with business goals, security, and safety. Prompt Engineering and Evaluation: Proficiency in designing prompts that enable agents to act reliably and setting up evaluation frameworks for monitoring performance. Requirement And Scoping Analysis: Requires knowledge of: Traceability matrix; Risk analysis methodologies; Cost Analysis; Business objectives; Classification of requirements; User stories To explore relevant products/solutions from an existing repertoire, that can address business/technical needs. Strong and demonstrable incident management skills with relevant experience in an enterprise organization. Methodical and systematic problem solving approach, combined with a solid awareness of ownership, initiative and drive. Experience investigating, analysing and troubleshooting large scale enterprise systems. Understanding of Unix/Linux systems from kernel to shell and beyond, taking in system libraries, file systems, and client-server protocols along the way. Experience working with and developing enterprise monitoring/tooling solutions like Grafana, Prometheus, Kibana, Splunk, Graphite, Dynatrace, catchpoint. Working knowledge of one or more cloud technologies such as AZURE, GCP and OpenStack. Expert verbal and written communication skills. Demonstrate excellent judgement in decision making. Strong focus on collecting and inferring metrics. Excellent communication skills What You’ll bring 10-14 years in an infrastructure, systems, engineering or development environment delivering operational excellence to highly complex distributed systems. Bachelor's Degree in Computer Science or a related field, or relevant work experience of 10+ years. Experience and exposure working in a 24/7 operations support environment. Working and technical expertise in K8 and microservice architectures. Experience administering Unix/Linux in a production environment. Ability to supervise the Site Reliability Operations team, mentor and provide guidance. Working knowledge of BASH, Python, AI or other scripting languages Utilize AI-powered monitoring and anomaly detection tools to predict potential failures and resource bottlenecks before they impact users. Ensure the reliability, performance, and scalability of infrastructure specifically designed for AI/ML workloads, Networking knowledge and understanding of network concepts, such as different protocols (TCP/IP, UDP, ICMP, etc.), MAC addresses, IP packets, DNS, OSI layers, and load balancing). **About Walmart Global Tech** Imagine working in an environment where one line of code can make life easier for hundreds of millions of people. That’s what we do at Walmart Global Tech. We’re a team of software engineers, data scientists, cybersecurity expert's and service professionals within