- Design, implement, and maintain monitoring tools such as Datadog to ensure high availability of IT systems.
- Collaborate with cross-functional teams to identify and troubleshoot issues affecting system performance.
- Develop and execute incident response plans to minimize downtime and restore service levels quickly.
- Conduct regular health checks on monitored systems to identify potential issues before they impact users.
Job Requirements :
- 7-14 years of experience in SRE or related field (e.g., administration).
- Robust understanding of Datadog or similar monitoring tools.
- Experience with cloud-based technologies (AWS/Azure) is a plus.
- Proven track record in designing scalable architectures for large-scale applications.