The Disaster Recovery & Capacity Engineer builds out solutions to support Platform disaster response/crisis management activities in compliance with the Engineering and Customer requirements and helps provide and coordinate disaster preparedness with respect to the organisation s Platform, helping ensure business continuity.
They also ensure we have enough resources to meet current and future Platform demand efficiently, involving forecasting needs, capacity planning, monitoring performance (KPIs), managing risks (shortages/overloads), and developing strategies for optimisation.
Responsibilities
- Work with Engineering & Service Management to ensure that the disaster recovery and Capacity plans drive disaster recovery (DR) strategy and procedures both in Cloud and DC venues.
- Build out tooling that supports the DR plans and tracks progress and maturity against set KPI s and Metrics.
- Work with Engineering & Service Management to ensure that disaster recovery solutions are adequate, in place, maintained, and tested as part of the regular operational life cycle.
- Provide ongoing feedback for risk management, mitigation, and prevention.
- Regularly report Disaster Recovery activities.
- Develop and implement capacity planning tooling, frameworks, policies, and strategies.
- Provide capacity requirements and impact assessments for recent services or changes.
- Manage capacity deviations and implement improvement initiatives
- Collaborate with other Platform managers to deliver objectives on our platform evolution roadmap.
Your skills and experience
We work with a broad range of technologies, and we don t expect you to know everything on day one. You ll have time to learn our tools and grow into the role. Were looking for diverse experiences to help strengthen our team.
Essential
- 5+ years of experience in IT operations/Production Engineering.
- Experience of Linux administration will be a day-one skill
- Experience with Kubernetes administration
- Being comfortable in a scripting language suitable for automation tasks
- A degree in computing or science is helpful but not essential
- Understanding of current recovery solutions and high availability architectures for cloud and on prem is needed.
- Understanding of current Capacity Management & Planning scenarios and tooling.
- Experience with Agile principles and practices.
- Expertise in problem diagnosis across complex, distributed systems.
Desirable
- Experience supporting SaaS products.
- Experience with Incident Management, Post Mortems and related practices.
- Knowledge of observability and monitoring best practices.
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.