R

Senior Director- Infrastructure, Operations & App Support

Randstad · Hyderabad, Telangana, India

15–25 yrs experiencefull_timePosted 2 days ago
Apply now →

Job description

**Role & responsibilities** **Cloud & Data & AI/ML Platform Operations** - Own the operational health, availability, and performance of enterprise cloud platforms (AWS, Azure, GCP) and AI/ML infrastructure ensuring production-grade reliability for models, pipelines, data  & analytics platforms, and cloud-native applications. - Establish and enforce MLOps and AIOps operational standards model monitoring, drift detection, automated retraining pipelines, inference infrastructure management, and incident response for AI/ML workloads. **Site Reliability Engineering (SRE)** - Build, lead, and mature the enterprise SRE function embedding reliability engineering principles (SLOs, SLIs, error budgets, chaos engineering) across critical digital platforms and services. - Lead the post-incident review (PIR) and blameless retrospective culture, ensuring every significant incident drives lasting systemic improvements rather than short-term fixes. **Managed Service Provider (MSP) Governance** - Serve as the executive owner of all MSP and third-party operational vendor relationships governing contracts, SLAs, performance metrics, and strategic alignment across managed infrastructure, cloud, support, and security services. - Lead structured QBRs, performance reviews, and executive-level escalations with MSP partners holding providers accountable to contractual commitments while fostering collaborative, long-term partnerships. **Enterprise IT Support & Service Management** - Oversee the enterprise IT support function Tier 1/2/3 support, service desk operations, and application support ensuring exceptional end-user experience and first-contact resolution metrics. - Lead the continuous maturation of ITSM processes (Incident, Problem, Change, Release, and Configuration Management) in alignment with ITIL best practices and enterprise risk controls. **Executive Visibility, Reporting & Stakeholder Engagement** **People Leadership, Mentoring & Team Culture** **Technical Competencies** - Deep expertise in cloud platform operations (AWS, Azure, GCP) architecture patterns, operational tooling, FinOps, and multi-cloud governance. - Strong grounding in AI/ML operations MLOps pipelines, model monitoring, inference infrastructure, and AIOps platform tooling. - SRE mastery — observability stacks (Datadog, Dynatrace, Prometheus/Grafana), chaos engineering, SLO frameworks, and incident management platforms. - ITSM fluency — ServiceNow or equivalent, ITIL v4 processes, and enterprise support operations at scale. - MSP governance and vendor management — SLA construction, performance metrics, contract lifecycle, and strategic sourcing principles. - Security and compliance operations awareness — vulnerability management, cloud security posture, and regulatory compliance frameworks.