Backup & Recovery Engineer (Rubrik)
Mizuho · Pune City, Maharashtra, India
Mizuho · Pune City, Maharashtra, India
**Why Mizuho** At Mizuho, we provide the stability of an international industry leader with the career trajectory of a growing business. Our steady, strategic growth gives our people at all levels rewarding degrees of responsibility and richer work experience than a boutique firm or an established giant could offer alone. It’s the local expertise of our employees that makes our global network so powerful. By collaborating with colleagues and clients who share the same ambition and drive, you can amplify your sphere of influence and base of knowledge as part of one of the largest banks in the world. **Position Overview:** The Backup & Recovery Engineer (Pune) is responsible for the day-to-day operation, monitoring, reporting, and continuous improvement of the enterprise backup and recovery platform. **This is a backup-focused engineering role with Rubrik as the primary technology area and enterprise storage as a secondary responsibility.** This role is accountable for maintaining backup reliability, monitoring platform health, reviewing failed jobs, supporting restore readiness, configuring and validating SLA domains, and producing operational reports for service health, compliance, and management visibility. The engineer will work closely with infrastructure, virtualization, storage, application, and operations teams to ensure workloads are protected and recoverable. Storage responsibilities are secondary and primarily support backup-related dependencies, recovery activities, capacity visibility, and operational coordination across platforms such as NetApp, Pure Storage, Dell PowerStore, IBM Storwize / FlashSystem, and SAN/NAS environments. **Key Responsibilities:** **Backup Platform Operations - Primary Focus** • Administer and operate Rubrik Security Cloud and associated enterprise backup infrastructure • Monitor backup platform health, service availability, capacity status, replication status, and operational alerts • Review daily backup activity and remediate failed backups, missed jobs, missed SLAs, policy exceptions, and workload protection gaps • Perform backup job troubleshooting, log review, evidence gathering, vendor case coordination, and escalation where needed • Support restore requests, recovery validation, and periodic restore testing for business and infrastructure teams • Maintain backup operational runbooks, escalation procedures, health-check documentation, and platform support procedures **Rubrik SLA Configuration & Policy Management** • Configure, maintain, and validate Rubrik SLA domains, retention policies, archival settings, replication behavior, and protection assignments • Ensure workloads are assigned to the appropriate SLA domain based on business requirements, recovery expectations, and retention needs • Review unprotected workloads, orphaned objects, policy conflicts, and SLA compliance exceptions • Partner with application and infrastructure owners to confirm protection requirements and resolve backup coverage gaps • Support periodic audits of policy adherence, retention alignment, workload coverage, and recoverability posture **Reporting, Analytics & Operational Visibility** • Produce daily, weekly, and monthly reports for backup success rates, SLA compliance, failed jobs, unprotected workloads, capacity trends, and restore activity • Build and maintain operational dashboards that provide visibility into platform health, backup reliability, protection coverage, and risk areas • Analyze backup trends and recurring failures to identify systemic issues and improvement opportunities • Provide reporting support for audit, risk, compliance, management reviews, and operational governance activities • Use Rubrik reporting, APIs, GraphQL, PowerShell, Python, or other tools to automate recurring reports and health checks where appropriate **Monitoring & Incident Remediation** • Serve as a first-line operational owner for backup platform monitoring and alert response • Investigate backup failures, replication issues, retention issues, SLA misses, capacity alerts, performance concerns, and recovery failures • Participate in incident response activities involving backup availability, recovery readiness, or protected workload impact • Coordinate with storage, compute, virtualization, network, application, and vendor teams during troubleshooting and remediation • Document corrective actions and contribute to reducing repeat incidents through process improvements and alert tuning **Backup Lifecycle, Upgrades & Platform Maintenance** • Support Rubrik software upgrades, maintenance activities, health checks, and post-upgrade validation • Assist with lifecycle tracking for backup platform components, known issues, support status, and required remediation activities • Coordinate with vendors and senior engineers for upgrade planning, implementation support, and issue resolution • Follow documented change procedures including pre-checks, implementation steps, validation tasks, and rollback planning • Maintain awareness of product releases, security advisories, operational defects, and platform improvement opportunities **Secondary Storage Support** • Provide secondary operational support for enterprise storage platforms that support protected workloads and recovery activities • Assist with basic storage administration and troubleshooting across NetApp, Pure Storage, Dell PowerStore, IBM Storwize / FlashSystem, and SAN/NAS environments • Support storage capacity reporting, backup target capacity review, and remediation of storage-related backup failures • Collaborate with storage engineers on replication, volume, LUN, share, pathing, and connectivity issues that impact backup and recovery services • Participate in storage incident support when backup services, recovery operations, or protected workload availability are affected **Documentation, Governance & Continuous Improvement** • Maintain accurate documentation for SLA domains,