14 Aug
|
APM Terminals
|
Bengaluru
14 Aug
APM Terminals
Bengaluru
Site Reliability Engineer (SRE)
Ocean Enablement Platform (OE) SRE Team
Role Overview
Join an exciting and cross-edge team shaping the future of container technology at Maersk. The OE SRE Team ensures reliability, performance, and automation excellence across the Ocean Enablement Platform.
As a Site Reliability Engineer, you will collaborate closely with OE platform teams to enhance system resilience, optimize cost and performance, and enable zero-touch operations through automation.
You will gain hands-on experience in cloud technologies, incident management, and Python-based automation while exploring AIOps and AI-driven automation to contribute to the reliability goals that power the Ocean Enablement Platform ecosystem.
Key Responsibilities
- Support and improve reliability, availability, and performance across OE applications and shared services.
- Participate in on-call rotations, handle incidents, perform RCA, and implement corrective actions.
- Develop and maintain automation tools and scripts using Python, Ansible, or shell scripting to reduce manual toil.
- Deploy, monitor, and manage workloads on AWS / Azure / GCP, ensuring cost-efficient and reliable operations.
- Configure and enhance observability (Prometheus, Grafana, ELK) for proactive detection and rapid recovery.
- Support SRE principles by defining SLIs, SLOs, and error budgets for OE services.
- Explore AIOps and AI-driven automation to reduce alert noise, accelerate triage, and enable intelligent remediation.
- Collaborate with product, infra,
and observability teams to improve reliability and incident response.
- Maintain runbooks, SOPs, and automation playbooks for operational readiness.
Required Skills Experience
- Bachelor s degree in Computer Science, Engineering, or related field.
- 3 5 years of experience in SRE, DevOps, or Cloud Infrastructure roles.
- Strong programming skills in Python for automation and scripting.
- Familiarity with AWS, Azure, or GCP cloud environments.
- Working knowledge of Docker, Kubernetes, and Infrastructure as Code tools (Terraform / Ansible).
- Experience with observability tools (Prometheus, Grafana, ELK, Datadog, etc.).
- Exposure to incident response, troubleshooting, and root cause analysis.
- Ownership-driven mindset with passion for improving reliability through automation.
Nice to Have
- Exposure to AIOps or AI/ML-driven automation for reliability and operations.
- Interest in building agentic AI workflows and AI-assisted automation for SRE use cases.
- Familiarity with LLM platforms (e.g., Azure AI Foundry / Azure OpenAI) and prompt-based tooling.
- Familiarity with cloud cost optimization and reliability frameworks.
- Knowledge of CI/CD pipelines and AIOps or AI/ML automation.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 SRE Engineer (Bengaluru)
🏢 APM Terminals
📍 Bengaluru