AI and HPC Systems Performance Engineer
Hewlett Packard Enterprise · Bengaluru, Karnataka, India
Free to search · AI fit score against your CV · tailor your résumé in one click
Hewlett Packard Enterprise · Bengaluru, Karnataka, India
AI and HPC Systems Performance Engineer This role has been designed as Hybrid with a requirement that you will work on average 2 days per week from an HPE office. High Performance Computing, AI and Labs is a critical element of HPE. We are focused on delivering innovative solutions that accelerate our customers digital transformation, enabling them to tackle their complex, and data-intensive workloads. Combining deep expertise and the development of the world s most cutting-edge, high-performance supercomputers, is defining the next era of computing delivering valuable insight innovation. Join us and redefine what s next for you. We are looking for an experienced AI Performance Engineer with expertise in tuning GPU server performance for a variety of Artificial Intelligence (AI) training and inference workloads running on Linux platforms. The ideal candidate will be a senior or principal-level engineer with demonstrated experience installing, configuring, characterizing, and optimizing industry-standard server infrastructure including compute, storage, networking, and accelerator technologies for AI workloads. The individual in this role will investigate workload behavior by capturing and analyzing system telemetry, profiling data, traces, and performance metrics to characterize workload execution and identify opportunities for optimization through software, firmware, and hardware configuration changes. They will work closely with customers, partners, and internal engineering teams to optimize performance and scalability of AI solutions deployed on HPE platforms. This role requires strong research, analytical, and problem-solving skills, along with experience building, deploying, optimizing, and maintaining containerized AI/ML environments and workloads. The engineer will collaborate with software development teams to capture workload telemetry, improve observability, and optimize AI software stacks running on HPE infrastructure. They should understand the performance and capacity characteristics of modern AI training and inference workloads and be comfortable working independently to evaluate emerging technologies, author technical papers, and develop AI reference architectures. Experience troubleshooting complex, multi-tier software systems and distributed AI environments is highly desirable. Experience with one or more of PyTorch, JAX, Hugging Face Transformers, DeepSpeed, Megatron-LM, Ray, vLLM, SGLang, TensorRT-LLM, Dynamo, ONNX Runtime, Kubernetes, Redis, Vector Databases, Retrieval-Augmented Generation (RAG) architectures, distributed training and inference, and large-scale AI/LLM workloads is highly desired. The ideal candidate will also have experience characterizing and optimizing performance across multi-GPU and distributed AI environments utilizing modern GPU interconnect, networking, and storage technologies. Strong written and verbal communication skills are required. What you ll do: • Install, configure, and optimize complex AI infrastructure components including GPU servers, storage systems, high-speed networking, and AI software stacks. • Develop automation scripts, deployment frameworks, and Infrastructure-as-Code solutions to streamline AI platform provisioning and workload execution. • Perform system-level performance characterization and optimization of AI training and inference workloads on HPE platforms utilizing GPU accelerators and distributed computing technologies. • Design, execute, and analyze performance benchmarks for AI/ML workloads, including large language models (LLMs), multimodal models, Retrieval-Augmented Generation (RAG) pipelines, and distributed training environments. • Characterize and optimize performance across multi-GPU and distributed AI environments using modern interconnect, storage, and networking technologies such as InfiniBand, Ethernet fabric, GPUDirect, and RDMA. • Capture, analyze, and interpret system telemetry, performance metrics, logs, traces, and profiling data to identify bottlenecks and optimization opportunities. • Develop tools, software, and automation frameworks to improve AI workload observability, performance analysis, scalability testing, and benchmark execution. • Collaborate with customers, partners, and internal engineering organizations to characterize, troubleshoot, and optimize AI solutions deployed on HPE infrastructure. • Work closely with ISV, IHV, GPU vendor, and open-source ecosystem partners to evaluate, optimize, and validate AI software and hardware solutions. • Evaluate emerging AI frameworks, models, accelerators, and infrastructure technologies; provide technical recommendations and performance guidance. • Author technical reports, white papers, reference architectures, benchmark studies, and best-practice guidance for AI performance optimization and solution design. • Document findings, performance issues, and optimization recommendations, and communicate technical results to engineering teams, customers, and management. • Provide technical leadership, mentoring, and guidance to junior engineers and contribute to the development of performance engineering best practices. • Communicate project status, technical risks, and performance findings to management and stakeholders in a timely manner. What you need to bring: • Typically 8+ years of experience • Strong experience with Linux system administration and command-line environments across multiple enterprise Linux distributions. • Experience with modern AI/ML frameworks and ecosystems including PyTorch, JAX, Hugging Face Transformers, and related technologies. • Experience with AI model training, inference, benchmarking, performance characterization, and optimization. • Experience with data analysis, statistical methods, experiment design, and performance modeling techniques. • Experience conducting technical research and evaluating emerging AI technologies, frameworks, and hardware platforms. • Exper