P

Sr Staff Site Reliability Engineer (Wildfire) - NetSec - Bangalore

Palo Alto Networks · Office - India - Bangalore Bagmane Tech Park

12–20 yrs experiencePosted 1w ago

Job description

Our Mission At Palo Alto Networks®, we’re united by a shared mission—to protect our digital way of life. We thrive at the intersection of innovation and impact, solving real-world problems with cutting-edge technology and bold thinking. Here, everyone has a voice, and every idea counts. If you’re ready to do the most meaningful work of your career alongside people who are just as passionate as you are, you’re in the right place. Who We Are In order to be the cybersecurity partner of choice, we must trailblaze the path and shape the future of our industry. This is something our employees work at each day and is defined by our values: Disruption, Collaboration, Execution, Integrity, and Inclusion. We weave AI into the fabric of everything we do and use it to augment the impact every individual can have. If you are passionate about solving real-world problems and ideating beside the best and the brightest, we invite you to join us! We believe collaboration thrives in person. That’s why most of our teams work from the office full time, with flexibility when it’s needed. This model supports real-time problem-solving, stronger relationships, and the kind of precision that drives great outcomes. Job Summary The Team Palo Alto Networks’ Cloud-Delivered Security Services (CDSS) is the intelligence engine of our Next-Generation Security platform. We provide a suite of AI-driven, subscription-based services—including Advanced Wildfire, Advanced Threat Prevention, DNS Security, URL Filtering, and IoT Security—integrated natively into our firewalls. Our infrastructure processes trillions of events daily, delivering real-time protection to over 85,000 global enterprises. Working in CDSS means building the backbone of global cybersecurity at a scale few companies in the world ever reach. Job Summary We are seeking an ambitious, technically sharp Senior Staff Site Reliability Engineer to drive the reliability, operational scalability, and infrastructure engineering for Palo Alto Networks’ WildFire malware analysis platform. In this high-impact role, you will take technical ownership of system resilience across WildFire services and appliance platforms. You will bridge the gap between threat analysis pipelines, low-level system execution, and high-throughput malware sandboxing infrastructure - ensuring sub-second detection telemetry and high availability for enterprise deployments worldwide. Key Responsibilities • Hybrid & Cloud Infrastructure Resilience: Architect, scale, and maintain operational reliability across WildFire’s multi-tenant public clouds (AWS/GCP/Azure/OCI), private cloud appliances, and hybrid inspection pipelines.  • Infrastructure Resilience: Architect, scale, and maintain the overarching operational reliability for WildFire’s cloud analysis engines, virtualized sandboxes, and distributed appliance infrastructure.  • Autonomous Operations & IaC: Spearhead the transition to fully automated operational workflows using Terraform, Ansible, and GitOps (ArgoCD) to eliminate operational toil across multi-tenant and edge environments.  • SLO & Error Budget Governance: Define and enforce SLIs, SLOs, and SLAs across malware inspection pipelines. Partner with security engineering leads on Error Budget strategies to balance rapid threat signature deployment with platform stability.  • Observability & Threat Telemetry: Architect end-to-end observability stacks (Prometheus, Grafana, OpenTelemetry, Datadog/ELK) optimized for low-latency tracing, high-concurrency sample processing monitoring, and rapid MTTR.  • Cloud & Platform Release Engineering: Build and scale enterprise CI/CD automation (GitHub Actions / GitLab CI) empowering engineering teams to safely deploy cloud microservices, threat analysis engines, and platform firmware updates.  • On-Call & Incident Response: Lead production on-call rotations, establishing escalation paths, automated alerting, and incident response procedures to ensure 24/7 reliability for mission-critical WildFire services.  • Incident Leadership & RCA: Lead critical incident response for high-severity platform outages. Conduct blameless post-mortems and implement systemic prevention measures across OS, container, and network layers.  • Performance Tuning & Sandboxing Efficiency: Perform capacity modeling, Linux kernel tuning, and resource optimization for virtualized sandboxing environments and high-volume malware sample processing pipelines.  • Capacity Planning & Performance Tuning: Perform capacity modeling, cost optimization, Linux kernel tuning, and resource allocation for cloud compute clusters and virtualized sandboxing environments.  • Technical Mentorship: Mentor engineers across teams, championing SRE and DevSecOps best practices while conducting rigorous operational reviews. Qualifications Required Qualifications • Experience: 4 - 9 years of professional experience in SRE, DevOps, Platform Engineering, or Infrastructure Engineering roles. • Cloud & Hybrid Systems: Strong hands-on experience architecting and managing production workloads in major Cloud Platforms (AWS, GCP, Azure, or