Senior Platform Engineer
Standard Chartered India · Chennai, Tamil Nadu
Standard Chartered India · Chennai, Tamil Nadu
**Job Summary** To own the end-to-end operational health, reliability, and performance of our AI/ML platform running across Azure, AWS, and On-Premises GPU environments. The ideal candidate will have deep expertise in maintaining GPU-accelerated infrastructure, AI/ML application stacks, and hybrid-cloud deployments while ensuring maximum uptime, scalability, and security. You will play a critical role in supporting production AI workloads, troubleshooting complex platform issues, and continuously optimizing the infrastructure that powers our AI applications (LLMs, ML models, inference services, RAG pipelines, etc.). **Key Responsibilities** **Platform Operations \& Maintenance** * Manage day-to-day operations, monitoring, and maintenance of the AI Platform across Azure, AWS, and on-premises GPU clusters. * Ensure high availability, reliability, and performance of AI/ML applications and services (training, inference, fine-tuning, RAG pipelines). * Perform proactive health checks, capacity planning, patching, and lifecycle management of infrastructure components. * Handle incident management, root cause analysis (RCA), and problem resolution within defined SLAs. **GPU Infrastructure Management** * Maintain and troubleshoot GPU nodes (NVIDIA A100, H100, V100, L40S, etc.), including driver installations, CUDA updates. * Optimize GPU utilization, workload scheduling, and resource allocation across multi-tenant environments. * Manage GPU clusters using Kubernetes (with GPU operators) or similar orchestrators. **Cloud \& Hybrid Environment Management** * Administer AI/ML workloads on Azure ML, AWS SageMaker, EKS/AKS, EC2/Azure VMs, and on-prem Kubernetes/OpenShift clusters. * Manage container-based deployments using Docker, Kubernetes, Helm, and GPU-enabled runtimes. * Configure and maintain networking, storage (NFS, Blob, S3), and IAM across hybrid environments. **AI Application Support** * Support production AI applications including LLM serving, MLflow, Kubeflow, Ray, and similar platforms. * Deploy, upgrade, and maintain model registries, vector databases and orchestration tools. * Collaborate with Data Scientists and ML Engineers to onboard models, optimize inference, and troubleshoot performance bottlenecks. **Automation \& CI/CD** * Automate platform provisioning, configuration, and maintenance using Terraform, Ansible, ARM/Bicep, CloudFormation. * Build and maintain CI/CD pipelines for AI/ML workloads (GitHub Actions, Azure DevOps, Jenkins, GitLab). * Implement Infrastructure as Code (IaC) practices for reproducible environments. **Monitoring, Logging \& Observability** * Implement and manage observability stacks: Prometheus, Grafana, ELK, Azure Monitor, CloudWatch. * Set up GPU-specific monitoring and alerting for AI workloads. * Track model performance, drift, latency, and throughput. **Security \& Compliance** * Enforce security best practices (identity, network, secrets management, encryption). * Ensure compliance with organizational and regulatory standards * Manage RBAC, key vaults, and secure model artifact storage. **Strategy** * To Maintain a highly available, secure, scalable, and cost-optimized AI Platform across Azure, AWS, and On-Premises GPU environments that enables seamless development, deployment, and operation of AI/ML and GenAI workloads with enterprise-grade reliability. **Business** * AI Factory architecture and engineering leadership * T\&A and TTO management * Chief Data Office and central data governance * Software Engineering Platform Teams * Business technology teams across Trade, Cash, Financial Markets, Risk \& Compliance, and Retail Operational, Technology \& Cyber Risk; Group Internal Audit Compliance, Regulators, and the independent AI governance function **Regulatory \& Business Conduct** * Display exemplary conduct and live by the Group's Values and Code of Conduct. * Take personal responsibility for embedding the highest standards of ethics, including regulatory and business conduct, across Standard Chartered Bank. This includes understanding and ensuring compliance with, in letter and spirit, all applicable laws, regulations, guidelines and the Group Code of Conduct. * Effectively and collaboratively identify, escalate, mitigate and resolve risk, conduct and compliance matters. **Key stakeholders** * AI Factory architecture and engineering leadership * T\&A and TTO management * Chief Data Office and central data governance * Software Engineering Platform Teams * Business technology teams across Trade, Cash, Financial Markets, Risk \& Compliance, and Retail Operational, Technology \& Cyber Risk; Group Internal Audit Compliance, Regulators, and the independent AI governance function **Skills And Experience** * Azure Administrator * Databricks * Linux Administrator * Kubernetes * GPU * Python **Qualifications** * Bachelor's/Master's degree in Computer Science, Engineering, or related field. * 8 years of experience in Platform/Infrastructure/DevOps/SRE roles, with at least 3 years in AI/ML platform engineering. * Strong hands-on experience with Azure and AWS (compute, networking, storage, IAM, GPU instances). * Solid experience managing on-premises GPU infrastructure (NVIDIA DGX, HPE, Dell, Supermicro GPU servers). * Expertise in Kubernetes, GPU Operator, container runtimes, and Helm. * Proficiency in Linux system administration, shell scripting, and Python. * Experience with IaC tools (Terraform, Ansible) and CI/CD pipelines. * Familiarity with AI/ML frameworks: PyTorch, TensorFlow, Hugging Face, and inference servers (Triton, vLLM, TGI). * Strong troubleshooting skills across networking, storage, GPU, and application layers. * Experience with monitoring/observability tools (Prometheus, Grafana, DCGM). * Certifications: Azure Solutions Architect, AWS Solutions Architect/DevOps, CKA/CKAD, NVIDIA DLI. **About Standard Chartered** We're an international bank, nimble enough to act, big enough for imp