Application Support Engineer III
Swiss Re · Hyderabad, Telangana, India
Swiss Re · Hyderabad, Telangana, India
### Application Support Engineer Are you passionate about keeping mission-critical systems running at peak performance while driving the future of IT operations through AI and automationJoin us and be part of a world-class application support organization at one of the world's leading reinsurance companies. ### About the Role We are looking for an experienced Application Support Engineer to play a pivotal role in ensuring the stability, resilience, and efficiency of our application landscape. This is not your average support role; it sits at the intersection of engineering excellence, intelligent automation, and strategic IT operations. You will combine deep technical expertise with a forward-thinking mindset, focusing on observability, end-to-end monitoring, infrastructure provisioning, security hardening, disaster recovery, and cloud cost optimization. You will work collaboratively with application teams, infrastructure teams, DevOps engineers, and business stakeholders to proactively monitor system health, resolve production issues, improve operational efficiency, and ensure business continuity all while championing AI-driven operations (AIOps, GenAI) and continuous service improvement. ### Key Responsibilities - **Enhance Application Stability** by establishing and improving observability practices across applications using logs, metrics, and traces with tools such as Azure Monitor, Application Insights, Log Analytics, Splunk, and Dynatrace implementing proactive monitoring, intelligent alerting, and event correlation - **Champion Site Reliability Engineering (SRE) principles** to improve the reliability, scalability, and resilience of production platforms defining, implementing, and monitoring Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets - **Drive IT Resilience initiatives** including Disaster Recovery (DR), High Availability (HA), and validation of RTO/RPO requirements to ensure business continuity - **Lead and participate in major incident response**, recovery exercises, and operational readiness reviews, ensuring swift resolution and thorough post-incident analysis - **Automate operational tasks** and implement self-healing solutions to improve reliability, reduce manual effort, and drive efficiency across the operational landscape - **Collaborate cross-functionally** with application, infrastructure, DevOps, and business teams to proactively identify risks, resolve production issues, and continuously improve service quality - **Contribute to cloud cost optimization** efforts by identifying inefficiencies and recommending improvements across the Azure environment - **Support the adoption of AIOps and GenAI-driven operations** to transform and modernize the IT operations model ### About the Team The Digital Technology Re (DT Re) organization is at the forefront of driving digital transformation for our reinsurance business units across PC Re, LH Re, and Solutions. We strive to build an inspiring environment where people harness data and technology to create a sustainable and strategic competitive advantage. We develop innovative analytics, data science capabilities, and robust data foundations that generate data-driven insights at the heart of Swiss Re business. Working hand-in-hand with our business counterparts in Property Casualty and Life Health, we deliver differentiating insights, elevate underwriting excellence, and effectively manage risk pools. Our team is a vibrant, international workforce spanning multiple locations and serving a truly global customer base. ### About You You are a technically strong, curious, and collaborative professional who thrives in fast-paced, high-stakes environments. You bring an automation-first mindset, sharp analytical skills, and a genuine passion for operational excellence. You communicate clearly and confidently across all levels of seniority from engineers to senior stakeholders and you take pride in driving continuous improvement, not just maintaining the status quo. You are energized by the opportunity to shape the next generation of IT Operations through autonomous, AI-driven transformation. We are looking for candidates who meet these requirements: - 8+ years of experience in Production Support, Site Reliability Engineering (SRE), DevOps, Platform Operations, or Application Support in complex, enterprise-scale environments - Strong expertise in Observability and Monitoring with hands-on experience using platforms such as Dynatrace, Splunk, Grafana, Prometheus, ELK Stack, Datadog, New Relic, or OpenTelemetry including dashboards, alerts, distributed tracing, log analytics, synthetic monitoring, Real User Monitoring (RUM), and Application Performance Monitoring (APM) - Proven experience with Disaster Recovery and High Availability including backup and restore, replication technologies, RTO/RPO planning, and Business Continuity management - Proficiency in scripting and automation using Python, PowerShell, Bash, or similar languages, along with Infrastructure as Code (Terraform, Ansible) and CI/CD tools such as Azure DevOps or GitHub Copilot - Hands-on experience with Microsoft Azure and working knowledge of relational databases including Oracle, SQL Server, PostgreSQL, and MySQL ### These are additional nice to haves: - Masters or graduate degree with a Computer Science background - Experience with Kubernetes, microservices, and cloud-native environments - Knowledge of AIOps and automated remediation solutions - Relevant certifications in SRE, Azure, Elastic, Grafana, or ITIL - Experience supporting large-scale, distributed, mission-critical systems - Knowledge of security, compliance, and risk frameworks in enterprise IT - Reinsurance industry experience is highly valued - Excellent organizational skills with the ability to manage multiple priorities simultaneously - Strong dedication to quality and a client-focused mindset - A passion for learning and continuous improvement for your