The Reliability Operations team partners with infrastructure, platform, and product engineering teams to improve the reliability, observability, and operational efficiency of corporate applications. The team focuses on automation, monitoring, incident response, and cloud platform operations to ensure highly available and scalable services.
Key Responsibilities
- Develop Python-based automation solutions to reduce manual operational tasks and improve incident response.
- Build, maintain, and support AWS cloud infrastructure using:
- EC2
- ECS
- EKS
- ELB
- CloudFormation
- Containerize and deploy applications using Docker and Kubernetes.
- Implement and enhance observability capabilities, including:
- Work with SQL and NoSQL databases for reporting, automation, and operational needs.
- Troubleshoot production issues across applications, infrastructure,
and databases.
- Collaborate with engineering and operations teams to deliver scalable and reliable solutions.
- Ensure high-quality code and operational excellence through testing and automation.