SRE Engineer
Maersk · India, Bengaluru, 560064
Maersk · India, Bengaluru, 560064
Site Reliability Engineer (SRE) Ocean & Enablement Platform (O&E) – SRE Team Role Overview Join an exciting and cross-edge team shaping the future of container technology at Maersk. The O&E SRE Team ensures reliability, performance, and automation excellence across the Ocean &Enablement Platform. As a Site Reliability Engineer, you will collaborate closely with O&E platform teams to enhance system resilience, optimize cost and performance, and enable zero-touch operations through automation. You will gain hands-on experience in cloud technologies, incident management, and Python-based automation — while exploring AIOps and AI-driven automation to contribute to the reliability goals that power the Ocean & Enablement Platform ecosystem. Key Responsibilities • Support and improve reliability, availability, and performance across O&E applications and shared services. • Participate in on-call rotations, handle incidents, perform RCA, and implement corrective actions. • Develop and maintain automation tools and scripts using Python, Ansible, or shell scripting to reduce manual toil. • Deploy, monitor, and manage workloads on AWS / Azure / GCP, ensuring cost-efficient and reliable operations. • Configure and enhance observability (Prometheus, Grafana, ELK) for proactive detection and fast recovery. • Support SRE principles by defining SLIs, SLOs, and error budgets for O&E services. • Explore AIOps and AI-driven automation to reduce alert noise, accelerate triage, and enable intelligent remediation. • Collaborate with product, infra, and observability teams to improve reliability and incident response. • Maintain runbooks, SOPs, and automation playbooks for operational readiness. Required Skills & Experience • Bachelor’s degree in Computer Science, Engineering, or related field. • 3–5 years of experience in SRE, DevOps, or Cloud Infrastructure roles. • Strong programming skills in Python for automation and scripting. • Familiarity with AWS, Azure, or GCP cloud environments. • Working knowledge of Docker, Kubernetes, and Infrastructure as Code tools (Terraform / Ansible). • Experience with observability tools (Prometheus, Grafana, ELK, Datadog, etc.). • Exposure to incident response, troubleshooting, and root cause analysis. • Ownership-driven mindset with passion for improving reliability through automation. Nice to Have • Exposure to AIOps or AI/ML-driven automation for reliability and operations. • Interest in building agentic AI workflows and AI-assisted automation for SRE use cases. • Familiarity with LLM platforms (e.g., Azure AI Foundry / Azure OpenAI) and prompt-based tooling. • Familiarity with cloud cost optimization and reliability frameworks. • Knowledge of CI/CD pipelines and AIOps or AI/ML automation. • Experience working with or administering databases (Oracle DBA, PostgreSQL, or other relational databases). • Experience developing or maintaining service health dashboards. What You’ll Gain • Hands-on experience in multi-cloud environments supporting global Ocean & Enablement platforms. • Opportunity to contribute to automation, AIOps, and observability fram