Sr. Principal Infrastr Eng (HPC)
Mphasis · Bengaluru, Karnataka, India
Free to search · AI fit score against your CV · tailor your résumé in one click
Mphasis · Bengaluru, Karnataka, India
Skill - HPC, Slurm, ClearML, Linux Location - Bengaluru Experience -7-10yrs Early joiners are preferred Job Summary: We are seeking a highly skilled Senior Principal Infrastructure Engineer to administer and optimize our ClearML Server and associated infrastructure. The ideal candidate will have a strong background in MLOps platforms, HPC execution environments, and containerized solutions, with a focus on supporting GPU-based AI/ML workloads. This role requires a proactive approach to managing resources, ensuring security, and automating processes to enhance operational efficiency. Responsibilities: • Administer ClearML Server, including management of agents, execution queues, projects, users, roles, experiment tracking, pipelines, datasets, artifacts, and model registry. • Configure ClearML Agents on CPU and GPU worker nodes, integrating with HPC execution platforms such as Slurm, PBS Professional, or Kubernetes. • Support GPU-based AI/ML workloads utilizing NVIDIA drivers, CUDA, NCCL, UCX, and containerized environments. • Maintain secure container execution using technologies like Apptainer/Singularity, Docker, Enroot, Pyxis, or Kubernetes. • Implement confidential-computing controls leveraging AMD SEV-SNP, Intel TDX, NVIDIA Confidential Computing, Secure Boot, TPM, and remote attestation. • Integrate authentication, RBAC, TLS certificates, secrets management, and audit controls for ClearML and confidential workloads. • Monitor ClearML services, agents, queues, GPU utilization, task failures, scheduler integration, and overall platform health. • Automate deployment, configuration, monitoring, and troubleshooting processes using Python, Bash, Ansible, and Git. Mandatory Skills: • Strong Linux administration skills, particularly with RHEL, SLES, Rocky Linux, or Ubuntu. • Hands-on experience with ClearML administration or a comparable MLOps platform. • Familiarity with HPC schedulers such as Slurm, PBS Professional, or LSF. • Knowledge of GPU platforms, including NVIDIA drivers, CUDA, and distributed training basics. • Proficiency in container technologies: Apptainer/Singularity, Docker, Enroot, Pyxis, or Kubernetes. • Strong scripting and automation skills in Python, Bash, Ansible, and Git. • Understanding of security fundamentals, including RBAC, IAM, TLS, certificates, secrets management, secure boot, and audit logging. • Familiarity with confidential-computing concepts such as TEE, encrypted memory, TPM, remote attestation, AMD SEV-SNP, Intel TDX, or equivalent technologies.