Search 100,000+ live jobs across India

Free to search · AI fit score against your CV · tailor your résumé in one click

A

Senior App/Prod Support (Tier 3 Site Reliability Engineer (SRE) / Platform Engineer)

AT&T · Hyderabad, Telangana, India

Est. ~₹22L (est.)6–12 yrs experiencefull_timePosted 2w ago

Job description

Key Responsibilities • Own platform reliability practices for availability, resilience, latency, and operational efficiency. • Drive DevOps and automation initiatives including Golden Image improvements and support automation use cases. • Implement and maintain GitHub Actions pipelines and CI/CD reliability standards. • Lead JFROG Helm chart automation and JFROG images/ACR migration work. • Support microservices deployment enablement and platform/tooling upgrades. • Own and optimize monitoring, alerting, observability, and logging stack components: Prometheus, AlertManager, Grafana, Azure Monitor, Thanos, OpenSearch, FluentBit, and related tools. • Support health-check frameworks including Airflow health-check requirements. • Provide troubleshooting support to Tier 1 and Tier 2 for high-complexity incidents. • Collaborate with architecture and delivery teams on reliability and scalability patterns. • Lead cloud infrastructure creation, maintenance, governance, and access controls. • Drive capacity planning, DR planning/exercises, and platform best-practice documentation. • Support cost management, role enforcement, and license management governance. • Maintain SOP documentation for established alerts and incident patterns. Required Qualifications / Must-Have Skills • 6+ years of experience in SRE, platform engineering, DevOps, or advanced production support roles. • Strong hands-on expertise with Kubernetes, especially Azure Kubernetes Service (AKS), and cloud-native platform operations. • Advanced experience with CI/CD engineering and GitHub Actions. • Deep observability experience with Prometheus/Grafana/AlertManager and logging stacks. • Strong Python automation scripting skills for reliability engineering, platform tooling, and operational toil reduction. • End-user proficiency with AI-assisted productivity and operations tools for incident analysis, troubleshooting acceleration, and documentation support (AI/ML model development is not required). • Familiarity with Java, React, and Spring Boot based services for production troubleshooting and stability improvements (not a feature-development role). • Strong hands-on experience with the mandated streaming stack, including enterprise operational depth in Confluent Kafka, Confluent Cloud, and Azure Event Hub: Confluent Kafka, Confluent Cloud, Azure Event Hub, AWS-MSK, and Apache Flink. • Experience in governance controls: access management, role enforcement, and separation of duties. • Proven high-severity incident leadership and post-incident reliability improvement execution. Good-to-Have / Nice-to-Have • Postgres performance and reliability operations. • Telecom-scale high-availability systems experience. Experience Level Senior to Lead IC (typically 10 to 17 years) Location / Work Mode Onsite (Hyderabad / Bangalore or designated AT&T location) What We Offer • Opportunity to define and scale platform reliability standards. • High technical ownership and strong cross-functional influence. • Enterprise-scale impact across observability, automation, and resilience engineering. Weekly Hours: 40 Time Type: Regular Location: IND:AP:Hyderabad / Argus Bldg 4f & 5f, Sattva, Knowledge City- Adm: Argus Building, Sattva, Knowledge City, IND:KA:Bangalore / Intl Tech Park, Navigator Bldg, Whitefield Road: Whitefield Road:Intl Tech Park, Navigator Bldg It is the policy of AT&T to provide equal employment opportunity (EEO) to all persons regardless of age, color, national origin, citizenship status, physical or mental disability, race, religion, creed, gender, sex, sexual orientation, gender identity and/or expression, genetic information, marital status, status with regard to public assistance, veteran status, or any other characteristic protected by federal, state or local law. In addition, AT&T will provide reasonable accommodations for qualified individuals with disabilities. AT&T is a fair chance employer and does not initiate a background check until an offer is made.

More jobs at AT&T

All AT&T jobs (155)

Engineering jobs in Hyderabad

Engineering jobs in Hyderabad (7,602)

Other Engineering jobs in India

All Engineering jobs in India (44,514)