Senior Staff Consulting Architect - AI Infrastructure
Nutanix · Pune Division, Maharashtra, India
Nutanix · Pune Division, Maharashtra, India
Role Overview Nutanix is seeking a Sr. Staff Consulting Architect – AI to provide senior technical leadership in the architecture, design, and transformation of enterprise AI platforms based on the Nutanix AI Platform (NAI) and equivalent enterprise AI/ML platforms. The role is intended for a highly experienced AI platform architect who can lead complex enterprise AI transformation initiatives spanning AI infrastructure, GPU-enabled Kubernetes, MLOps/LLMOps, Generative AI, RAG, model serving, data services, automation, security, and hybrid/multi-cloud environments. The Sr. Staff Consulting Architect will act as a trusted technical advisor to senior customer stakeholders, helping organizations define their AI platform strategy, target architecture, adoption roadmap, and operating model. The role requires a combination of deep hands-on expertise and strategic architecture leadership across AI/ML platforms, Kubernetes, virtualization, GPU infrastructure, cloud-native technologies, data platforms, and automation. The architect will also provide technical leadership to Staff Consulting Architects and Consultants, establish reference architectures and reusable AI solutions, and collaborate with Product Management, Engineering, and other technical organizations to influence the evolution and adoption of Nutanix AI solutions. Key Responsibilities - Enterprise AI Strategy & Architecture - Lead architecture strategy for enterprise AI/ML and Generative AI platform transformation programs. - Define target-state AI platform architectures aligned to customer business, data, application, security, and operational requirements. - Architect enterprise AI platforms across: - Nutanix AI Platform (NAI) - Kubernetes-based AI platforms - Nutanix AHV - Bare-metal GPU infrastructure - Hybrid and multi-cloud environments - Public cloud AI/ML platforms where applicable - Develop phased AI adoption roadmaps covering: - AI infrastructure - Model development - Model training and fine-tuning - Inference - RAG - MLOps/LLMOps - Application integration - Governance and operations - Advise customer executives, enterprise architects, data scientists, ML engineers, platform teams, and application teams on enterprise AI adoption. - Translate business use cases into scalable, secure, resilient, and operationally sustainable AI architectures. - Evaluate architecture trade-offs across performance, scalability, GPU utilization, cost, security, data locality, and operational complexity. - AI Infrastructure & GPU Platform Architecture - Design enterprise AI infrastructure supporting: - GPU-enabled Kubernetes - AI/ML workloads - Model training - Fine-tuning - Inference - Vector search - RAG workloads - Define GPU architecture considering: - GPU selection and sizing - GPU scheduling - GPU partitioning and sharing - CPU/GPU topology - NUMA - Memory requirements - Network bandwidth - Storage throughput - Architect high-performance AI infrastructure across virtualized and bare-metal environments. - Design GPU-enabled Kubernetes platforms using technologies such as: - NVIDIA GPU Operator - NVIDIA CUDA - NVIDIA Container Toolkit - NVIDIA Triton Inference Server - KServe - Evaluate and architect high-performance networking for AI workloads, including where applicable: - RDMA - InfiniBand - RoCE - High-speed Ethernet - NVIDIA networking technologies - Optimize infrastructure for GPU utilization, workload density, performance, availability, and cost. - Define architectures for scalable AI compute pools and shared enterprise AI platforms. - Kubernetes & AI Platform Engineering - Architect Kubernetes platforms optimized for enterprise AI workloads. - Define AI platform architecture across: - Kubernetes control plane - GPU worker nodes - Storage - Networking - Ingress - Security - Observability - Design multi-cluster and multi-tenant AI environments. - Define workload isolation across: - Cluster - Namespace - Node - GPU - Architect integration with enterprise infrastructure and services including: - Identity and access management - LDAP/Active Directory - DNS - PKI/certificates - Load balancers - Storage - Backup and DR - Monitoring and logging - Drive platform engineering and self-service capabilities for data scientists and AI developers. - Define standardized patterns for AI workspace provisioning and lifecycle management. - Generative AI, RAG & AI Application Architecture - Lead architecture for enterprise Generative AI solutions and platforms. - Design end-to-end RAG architectures covering: - Data ingestion - Document processing - Chunking - Embedding generation - Vector databases - Retrieval - Re-ranking - Prompt orchestration - LLM inference - Architect enterprise AI application patterns using technologies such as: - Large Language Models (LLMs) - Embedding models - Vector databases - LangChain - LlamaIndex - NVIDIA NeMo - NVIDIA NIM - Triton - KServe - Define architectures for: - LLM inference - Model serving - Fine-tuning - RAG - AI agents - Enterprise copilots - Evaluate open-source and commercial AI models and frameworks based on customer requirements. - Design model deployment strategies considering: - Performance - Latency - Throughput - GPU utilization - Model size - Security - Data privacy - Cost - Guide customers in moving AI prototypes and proof-of-concepts into production-grade platforms. - MLOps, LLMOps & Automation - Define enterprise MLOps and LLMOps architectures covering the complete AI lifecycle. - Design workflows for: - Data preparation - Model training - Fine-tuning - Model evaluation - Model registration - Model deployment - Model monitoring - Model retirement - Integrate AI platforms with CI/CD and GitOps practices. - Design automated AI platform provisioning using: - Terraform - Ansible - Kubernetes operators - APIs - Establish automation for: - GPU infrastructure provisioning - Kubernetes clusters - AI workspaces - Model deployment - Infer