Search 100,000+ live jobs across India

Free to search · AI fit score against your CV · tailor your résumé in one click

Job description

Skill - HPC, Slurm, ClearML, Linux Location - Bengaluru Experience -7-10yrs Early joiners are preferred Job Summary: We are seeking a highly skilled Senior Principal Infrastructure Engineer to administer and optimize our ClearML Server and associated infrastructure. The ideal candidate will have a strong background in MLOps platforms, HPC execution environments, and containerized solutions, with a focus on supporting GPU-based AI/ML workloads. This role requires a proactive approach to managing resources, ensuring security, and automating processes to enhance operational efficiency. Responsibilities: • Administer ClearML Server, including management of agents, execution queues, projects, users, roles, experiment tracking, pipelines, datasets, artifacts, and model registry. • Configure ClearML Agents on CPU and GPU worker nodes, integrating with HPC execution platforms such as Slurm, PBS Professional, or Kubernetes. • Support GPU-based AI/ML workloads utilizing NVIDIA drivers, CUDA, NCCL, UCX, and containerized environments. • Maintain secure container execution using technologies like Apptainer/Singularity, Docker, Enroot, Pyxis, or Kubernetes. • Implement confidential-computing controls leveraging AMD SEV-SNP, Intel TDX, NVIDIA Confidential Computing, Secure Boot, TPM, and remote attestation. • Integrate authentication, RBAC, TLS certificates, secrets management, and audit controls for ClearML and confidential workloads. • Monitor ClearML services, agents, queues, GPU utilization, task failures, scheduler integration, and overall platform health. • Automate deployment, configuration, monitoring, and troubleshooting processes using Python, Bash, Ansible, and Git. Mandatory Skills: • Strong Linux administration skills, particularly with RHEL, SLES, Rocky Linux, or Ubuntu. • Hands-on experience with ClearML administration or a comparable MLOps platform. • Familiarity with HPC schedulers such as Slurm, PBS Professional, or LSF. • Knowledge of GPU platforms, including NVIDIA drivers, CUDA, and distributed training basics. • Proficiency in container technologies: Apptainer/Singularity, Docker, Enroot, Pyxis, or Kubernetes. • Strong scripting and automation skills in Python, Bash, Ansible, and Git. • Understanding of security fundamentals, including RBAC, IAM, TLS, certificates, secrets management, secure boot, and audit logging. • Familiarity with confidential-computing concepts such as TEE, encrypted memory, TPM, remote attestation, AMD SEV-SNP, Intel TDX, or equivalent technologies.

More jobs at Mphasis

All Mphasis jobs (185)

Engineering jobs in Bengaluru

Engineering jobs in Bengaluru (14,549)

Other Engineering jobs in India

All Engineering jobs in India (44,397)