Site Reliability Engineer - Kafka Streaming Platform
Larsen & Toubro Infotech · Pune, Maharashtra
Larsen & Toubro Infotech · Pune, Maharashtra
Specialist - Cloud \& Infra Management Role Title Site Reliability Engineer Kafka Streaming Platform We are seeking a Site Reliability Engineer SRE to ensure the reliability availability and performance of our Kafkabased streaming platform This is an SREfocused rolethe primary objective is operational stability incident response monitoring and automation of the streaming environment rather than building streaming applications A working understanding of Kafka is essential to support and troubleshoot the platform effectively Key Responsibilities SRE Operational Reliability Primary Focus Monitor platform health throughput consumer lag and performance proactively address bottlenecks Respond to incidents INCs troubleshoot streaming issues and drive timely resolution Participate in oncall rotation and support major incident management Define and track SLOsSLIs for streaming reliability and availability Contribute to root cause analysis RCA and implement preventive measures Automate routine operational tasks to reduce manual effort scriptingconfig management Support capacity planning and forecasting for dataflow growth Maintain resilience through backup failover and recovery procedures Platform Support Maintenance Support and maintain Kafka stream applications from an operational standpoint Stream app routine health checks Collaboration Continuous Improvement Work with development infrastructure and data teams to onboard and support streaming use cases Document runbooks standard operating procedures and known error solutions Identify and implement improvements to enhance stability resilience and efficiency Required Skills Experience Solid understanding of SRE principles reliability monitoring incident management automation Working knowledge of Kafka topics partitions offsets consumer groups replication from an operational perspective Familiarity with LinuxUnix systems and commandline operations Scripting skills Bash Python for automation and operational tooling Experience with monitoringobservability tools Prometheus Grafana Splunk Awareness of ITIL processes incident problem and change management Strong troubleshooting and problemsolving abilities Good communication and documentation skills Desirable Nice to Have Experience operating Confluent Platform Schema Registry Connect ksqlDB Control Center Experience with CICD pipelines and infrastructureascode Ansible Jenkins Container orchestration understanding Kubernetes Docker Understanding of security concepts SASL TLSSSL ACLs