N

Sr. Operations Manager

Novartis · Hyderabad (Office)

8–15 yrs experiencePosted Yesterday
Apply now →

Job description

Job Description Summary Role Summary The Senior Operations Manager is responsible for the operations, reliability, security, and user enablement of the Novartis Biomedical Research Data Science Environments (DSE), MLOps capabilities, and platforms. The role ensures scientists, data scientists, and AI/ML engineers have a stable, compliant, and high-performing environment across a converged portfolio spanning interactive analytics platforms (Plenty, Posit/RStudio/RSConnect, Python/JupyterHub), model lifecycle and MLOps capabilities, and emerging agentic AI platforms (e.g., Biomni). The role works in close partnership with the High Performance Computing teams, Novartis infrastructure, security, and third-party providers. Location: Hyderabad, India #LI-Hybrid  Job Description Key Responsibilities  • Own operations across the DSE, MLOps, and Agentic Ops portfolio; coordinate small teams of external application supporters and manage operational priorities.  • Partner with product co-leads (Product Manager & Tech Lead) to align operational activities and outcomes with product roadmaps and business priorities.  • Provide user support, incident management, and issue escalation, coordinating with Novartis teams and third-party providers for timely resolution.  • Manage user community communications — proactive notifications on platform issues, model/agent updates, maintenance, and scheduled downtimes.  • Own user training, onboarding, and offboarding coordination across analytics platforms, MLOps tooling, and agentic frameworks.  • Perform platform, model, and agent performance monitoring, storage hygiene, and vulnerability management; manage incident detection, response, and resolution.  • Lead access, authentication, and authorization management, and Responsible-AI guardrails (prompt security, data-leakage, audit trails), in alignment with Novartis security, compliance, and AI governance standards.  • Test key functionalities after upgrades, patches, and framework releases, in collaboration with internal Novartis teams or third-party providers.  • Collaborate with the DSE Product Team, CSE/HPC teams, AI/ML engineering, and BR Compute Management for direction, alignment, and prioritization.  • Deliver regular usage, license, model, and agent utilization reports and status updates to product leads and the BR Compute Management Team.  • Drive adoption of best practices, automation, FinOps discipline (compute/GPU/token cost), and AI-assisted operations to enhance team efficiency, user experience, and service quality.  • Maintain and extend user documentation, wikis, publishing guides, setup instructions, internal SOPs, and best-practice guidance. Platform-Specific Responsibilities • Administer the Plenty platform — reliability, performance, secure operations, user lifecycle, and model lifecycle management as a core MLOps capability.  • Manage the lifecycle of the Posit Suite (RStudio / RSConnect) — maintenance, configuration, upgrades, patching, and backup validation.  • Support RSConnect publishing needs — dashboards, model/API endpoint deployment, and user access management.  • Manage the lifecycle of Python & JupyterHub environments and kernels, coordinating with HPC evolution cycles for aligned user experiences.  • Support users with HPC/GPU access, ML training workflows, and notebook-based applications (e.g., Dash, Streamlit) and their promotion to RSConnect.  • Drive DSE platform improvements — containerization, IDE integration, and MLOps-friendly workflows such as experiment tracking and reproducible environments.  • Enable model lifecycle capabilities across the DSE portfolio — versioning, deployment, monitoring, drift detection, and retirement — and help industrialize experimental workflows into governed MLOps pipelines.  • Support the early-stage operations and enablement of agentic AI platforms (e.g., Biomni) as this capability evolves — helping onboard agent frameworks, LLM backends, tool integrations, and vector/RAG components.  • Contribute to access, guardrail, and audit-trail practices for agentic AI, and track agent execution, token/compute consumption, and cost.  • Coordinate integration with HPC, DGX, and cloud compute back-ends used by data science, ML, and agentic workloads.  Required Qualifications & Experience   • Bachelor's or Master's degree in Computer Science, Data Science, Engineering, Computational Sciences, or a related discipline.  • 7+ years of hands-on experience operating data science / analytics / ML platforms in a research or enterprise environment, with at least 2+ years in a lead/co-lead operations role spanning DSE and/or MLOps.  • Strong Linux system administration and scripting skills (Bash,