We are from the Technology Operations Platform, and our vision is to improve the user experience of business users and the engineering community by making technology operations simpler and more productive. Our ultimate goal is to develop a comprehensive suite of tools and solutions that empower Site Reliability Engineering (SRE) teams to seamlessly get started and effectively manage performance and reliability across the organization.
Team/Prospect:
As a Site Reliability Engineer at Maersk, you will play a critical role in ensuring the reliability, scalability, and performance of our global systems. You will work closely with development and operations teams to automate processes, build resilient infrastructure, and drive continuous improvement. This role demands solid expertise in SRE principles, with a focus on automation and observability. AI/ML knowledge is a valuable plus.
Key Focus Areas:
Status Page: Emphasis on maintaining a status page for transparency and incident communication.
Zero-Touch Automation:
Highlighting strategies to eliminate manual interventions.
Middleware & Microservices: Demonstrating architectural know-how for robust service interactions.
Observability: Utilize observability tools and practices to build scalable SRE and automation solutions for global platforms and users.
Collaborate effectively with the Observability team to enhance system insights without overlapping responsibilities.
Operational Excellence: Applying principles to improve reliability and team efficiency.
Key Responsibilities:
Design, implement, and maintain scalable and reliable infrastructure.
Develop automations to eliminate manual, redundant toil.
Collaborate with cross-functional teams to define SLIs, SLOs, and error budgets.
Monitor system performance and availability using tools like Prometheus and Grafana.
Conduct root cause analysis and postmortems for incidents.
Drive adoption of SRE best practices across engineering tea