AI-MLOps SRE Lead Engineer
Randstad · Hyderabad, Telangana, India
Free to search · AI fit score against your CV · tailor your résumé in one click
Randstad · Hyderabad, Telangana, India
Job Title: AI-MLOps SRE Lead Engineer Location: Hyderabad. (Hybrid) Discover your role • Drive service reliability, availability, and performance across multi-cloud environments, establishing SLOs, SLIs, error budgets, and reliability standard methodologies. • Design, build, and scale enterprise ML platform infrastructure using technologies such as Dataiku, Amazon SageMaker AI, Databricks, and Google Vertex AI. • Develop AI-driven observability capabilities using anomaly detection, predictive analytics, and automated remediation solutions to proactively identify and resolve operational issues. • Lead the implementation and monitoring of LLM, SLM, RAG, and AI Agent platforms, ensuring performance, governance, operational efficiency, and scalability. • Design and implement Infrastructure as Code, CI/CD pipelines, self-healing systems, and platform automation capabilities to improve engineering productivity and operational resilience. • Architect enterprise ChatOps solutions that integrate operational events, AI workflows, observability platforms, and automated remediation capabilities. • Partner with Data Science, AI Engineering, and Platform teams to deliver secure, scalable, and production-ready AI/ML solutions. • Evaluate emerging AI-native operational technologies and integrate innovative solutions that enhance platform reliability, engineering efficiency, and business value. • Conduct technical debt assessments, identify architectural risks, and provide strategic recommendations to improve enterprise platform maturity. • Serve as a technical leader and trusted advisor, mentoring engineers and influencing reliability engineering, MLOps, and cloud platform strategy across the organization. • This role requires • Bachelor's degree in Computer Science, Information Technology, Engineering, Data Science, Artificial Intelligence, or a related subject area; Master's degree preferred. • 6-8 years of experience in Site Reliability Engineering, Platform Engineering, DevOps, or related technology fields with enterprise-scale delivery experience. • Strong hands-on experience operating across two or more major cloud platforms, including AWS, GCP, or Azure. • Deep expertise with ML platform technologies including Databricks, Amazon SageMaker AI, Dataiku, and Google Vertex AI. • Proven experience implementing end-to-end ML workflows including model training, deployment, experiment tracking, monitoring, and pipeline orchestration. • Advanced proficiency in Infrastructure as Code tools such as Terraform, Pulumi, AWS CDK, and modern CI/CD automation practices. • Strong programming and scripting skills in Python, Go, Bash, or similar languages. • Experience building enterprise observability solutions using Prometheus, Grafana, Datadog, OpenTelemetry, distributed tracing, metrics, and logging platforms. • Demonstrated expertise in anomaly detection, predictive analytics, automated remediation, and AI-assisted operational capabilities. • Proven experience designing and implementing enterprise ChatOps solutions and AI-enabled operational workflows. • Strong ability to assess technical debt, influence technical strategy, and drive platform modernization initiatives. • Experience using AI tools, LLM-powered assistants, and AI Agents to enhance engineering operations and productivity. • Experience with Kubernetes and container orchestration platforms such as EKS, GKE, or AKS preferred. • Familiarity with MLOps technologies including Kubeflow, Feast, MLflow, and model evaluation frameworks such as LangSmith, RAGAS, Evidently AI, or Weights & Biases preferred. • Knowledge of cloud cost optimization, policy-as-code, compliance automation, and multi-cloud governance practices preferred.