M

SRE Engineer

Maersk · India, Bengaluru, 560064

~₹28L (est.)3–9 yrs experiencePosted 3 days ago
Apply now →

Job description

Site Reliability Engineer (SRE)  Ocean & Enablement Platform (O&E) – SRE Team    Role Overview  Join an exciting and cross-edge team shaping the future of container technology at Maersk. The O&E SRE Team ensures reliability, performance, and automation excellence across the Ocean &Enablement Platform.  As a Site Reliability Engineer, you will collaborate closely with O&E platform teams to enhance system resilience, optimize cost and performance, and enable zero-touch operations through automation.  You will gain hands-on experience in cloud technologies, incident management, and Python-based automation — while exploring AIOps and AI-driven automation to contribute to the reliability goals that power the Ocean & Enablement Platform ecosystem.  Key Responsibilities  • Support and improve reliability, availability, and performance across O&E applications and shared services.  • Participate in on-call rotations, handle incidents, perform RCA, and implement corrective actions.  • Develop and maintain automation tools and scripts using Python, Ansible, or shell scripting to reduce manual toil.  • Deploy, monitor, and manage workloads on AWS / Azure / GCP, ensuring cost-efficient and reliable operations.  • Configure and enhance observability (Prometheus, Grafana, ELK) for proactive detection and fast recovery.  • Support SRE principles by defining SLIs, SLOs, and error budgets for O&E services.  • Explore AIOps and AI-driven automation to reduce alert noise, accelerate triage, and enable intelligent remediation.  • Collaborate with product, infra, and observability teams to improve reliability and incident response.  • Maintain runbooks, SOPs, and automation playbooks for operational readiness.  Required Skills & Experience  • Bachelor’s degree in Computer Science, Engineering, or related field.  • 3–5 years of experience in SRE, DevOps, or Cloud Infrastructure roles.  • Strong programming skills in Python for automation and scripting.  • Familiarity with AWS, Azure, or GCP cloud environments.  • Working knowledge of Docker, Kubernetes, and Infrastructure as Code tools (Terraform / Ansible).  • Experience with observability tools (Prometheus, Grafana, ELK, Datadog, etc.).  • Exposure to incident response, troubleshooting, and root cause analysis.  • Ownership-driven mindset with passion for improving reliability through automation.  Nice to Have  • Exposure to AIOps or AI/ML-driven automation for reliability and operations.  • Interest in building agentic AI workflows and AI-assisted automation for SRE use cases.  • Familiarity with LLM platforms (e.g., Azure AI Foundry / Azure OpenAI) and prompt-based tooling.  • Familiarity with cloud cost optimization and reliability frameworks.  • Knowledge of CI/CD pipelines and AIOps or AI/ML automation.  • Experience working with or administering databases (Oracle DBA, PostgreSQL, or other relational databases).  • Experience developing or maintaining service health dashboards.  What You’ll Gain  • Hands-on experience in multi-cloud environments supporting global Ocean & Enablement platforms.  • Opportunity to contribute to automation, AIOps, and observability fram