Senior Manager, Systems Engineering
Oracle · Bengaluru, Karnataka, India
Oracle · Bengaluru, Karnataka, India
**Job Description** This role reports into senior leadership responsible for global AI/GPU host management strategy, fleet readiness, AIOps adoption, intelligent automation, service reliability, and operational tooling. You will translate that strategy into disciplined execution across day-to-day operations, service queues, triage, incident response, patching, upgrades, runbook maturity, telemetry improvements, and operational readiness for OCI's AI/GPU host fleet. You will be a hands-on leader during operational escalations, driving crisp communications, rapid diagnosis, and durable corrective actions. You will partner cross-functionally with geographically distributed OCI engineering, platform, networking, data center, compliance, and operations teams to deliver measurable reliability, efficiency, and execution outcomes. **Responsibilities** **Key Responsibilities** **Organizational Leadership & Talent Development** - Lead, grow, and develop India-based teams of operators, developers, and technical leads in a 24x7 operational environment; establish clear ownership boundaries, on-call expectations, and accountability mechanisms. - Recruit, hire, coach, and retain high-performing talent; set goals, manage performance, develop successors, and raise the technical and operational bar across the team. - Create a culture of operational excellence, ownership, continuous learning, pragmatic simplification, and disciplined execution while supporting rapid AI/GPU infrastructure growth. **Service Ownership: AI/GPU Host Operations at Cloud Scale** - Own operational outcomes for assigned AI/GPU host management services, including host health, fleet readiness, hardware/software triage, service queues, reliability, and operational readiness. - Guide teams that diagnose AI compute host issues across hardware, Linux, services, networking, automation, and monitoring layers; ensure issues are resolved quickly and with durable prevention mechanisms. - Drive operational execution for service patching, upgrades, staged rollouts, change controls, and readiness reviews using metrics-driven planning and governance. - Partner with global OCI teams to ensure India operations align with broader host management roadmaps, standards, escalation practices, and fleet reliability goals. **Operational Excellence, Metrics, and Governance** - Define and mature operational KPIs and reporting for service health, incident performance, ticket aging, ticket resolution quality, queue backlog, change execution, and operational readiness. - Lead high-severity incidents and escalations: coordinate rapid triage, communicate status clearly, drive high-quality post-incident reviews, and follow through on corrective actions. - Improve operating mechanisms for risk management, audit/compliance alignment, change management, handoffs, escalation hygiene, and cross-team execution tracking. - Use data and trend analysis to identify repeat issues, operational bottlenecks, staffing gaps, process defects, and opportunities for reliability improvement. **Engineering Enablement, AIOps, and Automation** - Drive adoption of AIOps and intelligent automation to reduce manual toil, improve alert quality, accelerate event correlation, standardize remediation, and improve triage accuracy. - Partner with engineering and platform teams to prioritize operational tooling, telemetry improvements, workflow enablement, runbook automation, and self-healing opportunities. - Establish measurable automation outcomes such as reduced manual handling, improved MTTR, increased auto-triage coverage, fewer repeat issues, and better operator/engineer effectiveness. - Manage focused software engineering work aligned to operations outcomes, including scripts, dashboards, workflow tooling, triage aids, monitoring enhancements, and reliability automation. **Cross-Functional & Stakeholder Engagement** - Work closely with senior leaders, peer managers, technical leads, and globally distributed teams to deliver predictable execution across AI/GPU host operations. - Translate complex technical and operational situations into accurate narratives, decisions, risks, and action plans for senior stakeholders. - Represent the team in operational reviews, readiness discussions, incident forums, and cross-functional planning sessions with clarity and ownership. **Qualifications / Experience** - BS or MS in Computer Science, Engineering, or a related technical field, or equivalent practical experience. - 8+ years of experience in software engineering, infrastructure operations, cloud operations, site reliability, production operations, or related technical areas. - 5+ years of people management and/or technical leadership experience, including experience leading operators, engineers, senior technical contributors, or multiple operational workstreams. - Experience building and scaling teams, including recruiting, hiring, coaching, performance management, goal setting, and leadership development. - Strong operational background with incident management, service ownership, queue management, operational readiness, process improvement, and post-incident corrective actions. - Experience operating or supporting Linux-based infrastructure at scale, including hardware/software troubleshooting and service lifecycle execution. - Working familiarity with scripting and automation ecosystems such as Python, Bash, or similar tools sufficient to sponsor, review, and guide operational tooling direction. - Understanding of distributed systems fundamentals and the ability to reason across hardware, operating systems, networking, services, monitoring, automation, and customer impact. - Familiarity with networking protocols such as TCP/IP and HTTP and with standard cloud infrastructure architectures. - Strong organizational and planning skills, including prioritization, scheduling, execution tracking, and operational governance. - Strong written and verbal communication skills, includ