S

MLops Engineer

Straive · Bangalore Urban, Karnataka, India

3–9 yrs experiencefull_timePosted 2w ago

Job description

We are looking for a highly skilled **MLOps Engineer** to build, deploy, automate, and manage Machine Learning solutions in production. The ideal candidate should have hands-on experience in designing scalable ML pipelines, deploying models, automating workflows, monitoring model performance, and implementing CI/CD for machine learning applications. The role requires close collaboration with Data Scientists, Data Engineers, and DevOps teams to ensure reliable and efficient ML model lifecycle management. **Key Responsibilities** - Design, develop, and maintain end-to-end MLOps pipelines for training, testing, deployment, and monitoring of machine learning models. - Deploy machine learning models to cloud and on-premise environments using containerization and orchestration technologies. - Build and automate CI/CD pipelines for machine learning applications. - Develop reusable ML workflows for model training, validation, deployment, and versioning. - Implement model monitoring, performance tracking, drift detection, and automated retraining strategies. - Manage ML artifacts, datasets, feature stores, and model registries. - Collaborate with Data Scientists to operationalize machine learning models. - Optimize infrastructure for scalable and cost-effective model deployment. - Troubleshoot production issues and ensure high availability of ML services. - Maintain security, governance, and compliance standards across ML platforms. **Required Skills** - 5–8 years of IT experience with at least 3+ years of hands-on experience in MLOps. - Strong programming experience in Python. - Hands-on experience with one or more MLOps platforms: - MLflow - Kubeflow - Azure Machine Learning - AWS SageMaker - Google Vertex AI - Experience deploying machine learning models using: - Docker - Kubernetes - Strong knowledge of CI/CD tools: - Jenkins - Azure DevOps - GitHub Actions - GitLab CI/CD - Experience with version control using Git. - Strong understanding of machine learning lifecycle and model management. - Experience with REST APIs for model serving. - Knowledge of SQL and data engineering concepts. - Experience with Linux and shell scripting. **Cloud Platforms** Experience with at least one of the following: - Microsoft Azure - Amazon Web Services (AWS) - Google Cloud Platform (GCP) **Good to Have** - Experience with Apache Airflow or Prefect for workflow orchestration. - Knowledge of Terraform or Infrastructure as Code (IaC). - Experience with Spark or PySpark for large-scale data processing. - Familiarity with Databricks. - Experience with Kafka or other streaming platforms. - Knowledge of Feature Store implementation. - Experience with monitoring tools such as Prometheus, Grafana, ELK, or Azure Monitor. - Exposure to LLMOps, Generative AI, or Large Language Models is an added advantage.