Site Reliability Engineer (SRE) - OpenTelemetry & Kubernetes
Larsen & Toubro Infotech · Bengaluru/Bangalore, Karnataka
Larsen & Toubro Infotech · Bengaluru/Bangalore, Karnataka
Specialist - System Management Role Title Kubernetes OpenTelemetry OTel Engineer SRE We are seeking an OpenTelemetry OTel Engineer with a focus on Site Reliability Engineering SRE to support our observability platform running on Kubernetes This is an SREfocused rolethe primary objective is ensuring the reliability availability and performance of the OpenTelemetry collection and pipeline infrastructure A working understanding of OpenTelemetry and Kubernetes is essential to operate monitor and troubleshoot the telemetry platform effectively Key Responsibilities SRE Operational Reliability Primary Focus Monitor the health throughput and performance of OpenTelemetry collectors and pipelines proactively address bottlenecks Respond to incidents INCs troubleshoot telemetry dataflow issues and drive timely resolution Participate in oncall rotation and support major incident management Define and track SLOsSLIs for telemetry pipeline reliability and availability Contribute to root cause analysis RCA and implement preventive measures Automate routine operational tasks to reduce manual effort scriptingconfig management Support capacity planning and forecasting for telemetry data growth Maintain resilience through failover recovery and pipeline redundancy procedures OpenTelemetry Platform Support Deploy configure and maintain OpenTelemetry Collectors receivers processors exporters on IKPKubernetes Support instrumentation of applications for traces metrics and logs Manage telemetry pipeline configuration sampling batching and data routing to backends Operate and maintain OTel components running on Kubernetes pods deployments config maps secrets Ensure reliable data delivery to observability backends eg Prometheus Grafana Splunk Jaeger Collaboration Continuous Improvement Work with development infrastructure and platform teams to onboard new services into observability Document runbooks standard operating procedures and known error solutions Identify and implement improvements to enhance stability resilience and efficiency Required Skills Experience Solid understanding of SRE principles reliability monitoring incident management automation Working knowledge of OpenTelemetry concepts traces metrics logs collectors exporters Experience operating workloads on Kubernetes IKP or equivalent Familiarity with LinuxUnix systems and commandline operations Scripting skills Bash Python for automation and operational tooling Experience with monitoringobservability tools Prometheus Grafana Splunk Awareness of ITIL processes incident problem and change management Strong troubleshooting and problemsolving abilities Good communication and documentation skills.