Senior Reliability Engineer
LSEG · IND-Hyderabad-CapitaLand
LSEG · IND-Hyderabad-CapitaLand
Cloud Reliability Engineer 📍 Hyderabad, India | 🕒 24×7 Reliability Engineering Model Build Reliability at Scale. Own Critical Cloud Systems. We’re hiring a Cloud SRE Engineer to join our Operations Reliability Engineering (ORE) team—focused on keeping mission-critical systems highly available, scalable, and resilient. This role goes beyond traditional support. You’ll operate at the intersection of cloud engineering, reliability, and automation, owning the health of distributed systems running on AWS while driving operational excellence. If you enjoy deep system visibility, proactive reliability engineering, and solving production-scale challenges, this role is for you. What You’ll Own 🚀 Cloud Reliability & Platform Operations • Own and manage reliability of AWS-based infrastructure (EC2, S3, RDS/Aurora) • Drive best practices for resource management, tagging, and optimization • Perform lifecycle operations on cloud infrastructure aligned to business needs • Manage AMI strategies, backup readiness, and system recovery posture 📈 Observability, Monitoring & Incident Engineering • Leverage Datadog to proactively detect and resolve issues • Build strong situational awareness across:• Infrastructure health (CPU, memory, storage) • Database performance and replication lag • Act on alerts with a focus on root cause identification, not just resolution • Improve alert quality, reduce noise, and enhance system observability 🧠 Database Reliability Engineering • Ensure high availability and performance of Aurora/RDS clusters • Monitor and optimize:• Replication and mirroring health • Query performance and database load • Partner with engineering teams on performance tuning and resilience improvements 💾 Data Protection, Backup & Recovery Engineering • Own the integrity and reliability of enterprise backup systems (Commvault) • Ensure backups meet strict SLA and compliance requirements • Perform and validate restore operations (critical recovery scenarios) • Continually improve backup validation, reporting, and audit readiness 🔐 Security, Access & Compliance • Manage certificates, access provisioning, and credential governance • Support SOX compliance and audit requirements with precision • Ensure all operations meet security and regulatory standards ⚙️ Operational Excellence & Automation Opportunities • Drive consistency through runbooks, SOPs, and process improvements • Identify opportunities to automate repetitive operational tasks • Contribute to evolving the team from reactive support → proactive SRE culture What Makes You a Strong Fit 💻 Core Skills • Strong hands-on experience with core AWS services, including EC2, S3, RDS/Aurora, and CloudWatch • Deep expertise across AWS service domains: • • Compute: EC2, Auto Scaling, AWS Lambda • Storage: S3, EBS, EFS • Databases: RDS, Aurora, DynamoDB • Networking: VPC, Subnets, Load Balancers (ALB/NLB), Route 53 • Proven ability to design and implement highly available, scalable, and fault-tolerant architectures, including multi-AZ deployments and disaster recovery (DR) strategies • Strong experience with observability and monitoring tools (e.g., Datadog preferred, CloudWatch, Prometheus/Grafana) • Hands-on experience with backup and recovery solutions (e.g., Commvault), including DR planning and testing • Solid understanding of: • • Distributed systems and reliability engineering principles • Database operations, including replication, backup, restore, and performance tuning • Experience working in production environments with high availability and critical uptime requirements 🛠 Engineering Mindset • You think beyond tasks—focused on system reliability and resilience • Strong troubleshooting skills with a bias for root cause analysis • Comfort working in mission-critical environments with real-time impact 🔄 Ops + SRE Balance • Experience in 24×7 production environments • Ability to operate calmly under pressure while maintaining precision • Interest in moving towards automation, efficiency, and reliability engineering practices ➕ Nice to Have • Scripting experience (Python, Shell) for automation • Exposure to ITIL, incident management frameworks • Experience in regulated environments (SOX compliance) ✅ Recommended Certifications (Preferred) • AWS Certified Cloud Practitioner (CLF-C02) • AWS Certified Solutions Architect – Associate Why This Role Stands Out • 🌍 Work on business-critical, high-scale cloud systems • ⚙️ Be part of a team evolving towards true SRE practices—not just operations • 📊 Gain deep exposure to AWS, observability, and database reliability at scale • 🚀 High ownership, real impact, and strong growth in Cloud + SRE + DevOps 👉 If you’re passionate about reliability, enjoy solving real production challenges, and want to operate at scale—this is your role. Apply now. We're proud to have been recognised as a Great P