AI Support
Mizuho · Pune City, Maharashtra, India
Mizuho · Pune City, Maharashtra, India
**Why Mizuho** At Mizuho, we provide the stability of an international industry leader with the career trajectory of a growing business. Our steady, strategic growth gives our people at all levels rewarding degrees of responsibility and richer work experience than a boutique firm or an established giant could offer alone. It’s the local expertise of our employees that makes our global network so powerful. By collaborating with colleagues and clients who share the same ambition and drive, you can amplify your sphere of influence and base of knowledge as part of one of the largest banks in the world. The AI Support role is responsible for deep technical investigation, diagnosis and remediation of complex issues impacting production AI applications. This role serves as the final escalation point for technical issues and works in close partnership with the AI Support Lead to restore system stability and reduce recurrence of high impact incidents. This role requires strong technical depth combined with prior production support experience. Candidates whose background is limited to development without exposure to live production support are not suitable for this role. **Key Responsibilities:** **Deep Technical Incident Resolution** - Provide support for complex production issues impacting AI applications across Production, UAT and Development. - Perform advanced root cause analysis across application code, data pipelines, model behavior and infrastructure dependencies. - Design and implement corrective fixes, patches and workarounds to restore system stability. - Validate and monitor fixes post deployment to ensure issues are fully resolved. **Escalation Support & Collaboration** - Act as the technical escalation point during high severity or complex incidents. - Partner closely with the AI Support Lead during Sev 1 and Sev 2 incidents to support decision making and remediation strategy. - Provide clear technical assessments and risk evaluations to inform business facing decisions. - Ensure escalations are handled efficiently without unnecessary rework or delays. **Operational Readiness & Stability** - Identify systemic weaknesses in AI applications and supporting infrastructure that increase operational risk. - Improve observability, logging, alerting and diagnostics to reduce time to resolution. - Support production releases by validating readiness, rollback plans and post release stability. - Participate in planned maintenance, upgrades and platform changes with a focus on operational impact. **Documentation & Knowledge Transfer** - Document root causes, fixes and known failure patterns to reduce repeat incidents. - Contribute to runbooks, troubleshooting guides and operational documentation. - Share knowledge proactively to raise the overall technical maturity of the support team. **Continuous Improvement** - Participate in post incident reviews and technical retrospectives. - Identify opportunities to reduce incident frequency, severity and recovery time. - Recommend architectural or design changes to improve long term stability and supportability. **Required Experience:** **Mandatory:** - Strong hands on experience (4+ years) supporting complex production systems in an support role. - Demonstrated ability to diagnose and resolve complex issues under production pressure. - Experience working in SLA driven support environments. - Ability to communicate complex technical issues clearly to non technical stakeholders. **Strongly Preferred:** - Experience supporting AI, analytics or data driven platforms in production. - Strong understanding of data pipelines, model lifecycle management and application monitoring. - Experience supporting cloud based architectures and distributed systems. **Technical Requirements:** - Strong hands on expertise in Python for production AI applications including debugging complex failures and implementing production safe fixes. - Strong understanding of AI model lifecycle, data pipelines, orchestration layers and runtime behavior. - Strong working knowledge of SQL and data storage patterns used by AI applications. - Experience performing deep root cause analysis across application code, data processing logic and infrastructure dependencies. - Experience implementing patches, fixes and workarounds in live production environments. - Strong familiarity with observability practices including logging, metrics and alerting. - Experience supporting cloud based AI systems and distributed architectures in production. **What This Role Is Not** - Not a general development role - Not a support lead role - Not a passive escalation resource. **This role owns technical resolution and stability.** **Company Overview:** Mizuho Pune is an integral part of Mizuho Financial Group, one of the world’s leading financial institutions with a strong global presence across the Americas, EMEA, and Asia. Based in India, Mizuho Pune supports Mizuho’s international businesses by delivering high-quality, scalable, and resilient services across multiple functions. Mizuho Pune plays a critical role in driving operational excellence, standardization, and innovation for Mizuho Americas. By combining deep domain expertise with strong process, technology, and analytical capabilities, it partners closely with regional and global teams to support corporate and investment banking, capital markets, and corporate services functions, while adhering to the highest standards of risk management, regulatory compliance, and control. Mizuho Pune offers competitive compensation and benefits package aligned with industry standards and local market practices. Mizuho Pune is an equal opportunity employer and is committed to fostering an inclusive and diverse workplace. Employment is subject to applicable background verification checks in accordance with Indian laws and company policies. **https://www.mizuhogroup.com/asia-pacific/mizuho-global-services/careers**