SRE Reliability Engineer
NTT DATA BUSINESS SOLUTIONS · Bengaluru, Karnataka, India
Free to search · AI fit score against your CV · tailor your résumé in one click
NTT DATA BUSINESS SOLUTIONS · Bengaluru, Karnataka, India
Job Summary We are looking for an experienced Site Reliability Engineer (SRE) with a strong background in Kubernetes, production support, observability, SQL, and Java-based applications. The role will focus on ensuring the availability, reliability, scalability, and performance of business-critical production systems. We are currently seeking a SRE Reliability Engineer to join our team in Bangalore, Karnataka, India. The ideal candidate will combine strong application troubleshooting skills with SRE and DevOps practices, using Datadog and/or Prometheus for monitoring and observability and Kubernetes for managing containerized workloads. A strong understanding of Java applications and relational databases/SQL is essential for diagnosing issues across application, infrastructure, and data layers. Key Responsibilities • Own the reliability, availability, and operational health of business-critical production applications and services. • Provide L2/L3 production support, including incident triage, troubleshooting, resolution, and stakeholder communication. • Monitor and support applications deployed on Kubernetes, including pods, deployments, services, ingress, resource utilization, scaling, and cluster-related issues. • Implement and maintain application and infrastructure monitoring using Datadog, Prometheus, dashboards, metrics, logs, and alerts. • Define and monitor SLIs, SLOs, SLAs, error budgets, and service-health indicators for critical applications. • Troubleshoot production issues across Java applications, APIs, microservices, Kubernetes, databases, and infrastructure. • Analyze Java application logs, exceptions, JVM performance, memory utilization, thread behavior, and garbage collection to identify performance and reliability issues. • Use SQL to investigate production incidents, validate data, identify data-related issues, and perform application-level troubleshooting. • Participate in incident management, major incident calls, root-cause analysis (RCA), and post-incident reviews. • Identify recurring production issues and drive permanent remediation through automation and engineering improvements. • Build and enhance monitoring dashboards, alerting mechanisms, and operational runbooks to improve early detection and reduce recovery time. • Drive improvements in MTTR, availability, performance, capacity, and production stability. • Automate repetitive operational activities using scripting and appropriate DevOps/SRE tooling. • Support application releases, production deployments, rollback activities, and post-deployment validation. • Work closely with Development, Infrastructure, DevOps, Database, Security, and Business teams to ensure production readiness. • Participate in on-call and production support rotations as required. Required Skills & Experience • 5+ years of overall IT experience, with significant experience in SRE, Production Support, Application Support, or DevOps roles. • Strong hands-on experience with Kubernetes and containerized applications. • Experience with Datadog and/or Prometheus for monitoring, alerting, metrics, and observability. • Strong experience supporting Java/J2EE or Java-based microservices applications in production. • Good understanding of JVM troubleshooting, application logs, memory, threads, garbage collection, and performance issues. • Strong SQL skills with experience troubleshooting relational databases and application data issues. • Experience managing P1/P2 production incidents, including incident coordination, RCA, and problem management. • Good understanding of REST APIs, microservices, distributed systems, and application integration patterns. • Experience with Linux/Unix environments and shell scripting. • Understanding of CI/CD pipelines, release management, and deployment practices. • Strong analytical and troubleshooting skills with the ability to diagnose issues across multiple technology layers. Preferred Skills • Experience with cloud platforms such as AWS/ Azure. • Familiarity with Docker, Helm, GitLab/GitHub Actions, or similar DevOps tooling. • Experience with centralized logging platforms such as ELK/OpenSearch or Splunk. • Exposure to Infrastructure as Code tools such as Ansible/Terraform. • Understanding of load balancing, networking, DNS, certificates, and application security. • Experience implementing automation to reduce manual operational effort and production toil. • Familiarity with ITIL processes including Incident, Problem, and Change Management. Key SRE Competencies • Production Reliability: Ability to maintain highly available and resilient production services. • Observability: Strong understanding of metrics, logs, traces, dashboards, and actionable alerting. • Incident Management: Ability to rapidly diagnose and restore services during critical incidents. • Problem Management: Strong RCA skills with focus on permanent remediation rather than repeated tactical fixes. • Automation: Ability to identify and automate repetitive production-support activities. • Performance Engineering: Ability to identify bottlenecks across Java applications, Kubernetes, and databases. • Stakeholder Management: Ability to communicate clearly during incidents and work effectively across engineering and business teams.