09 Aug
|
IntraEdge
|
Hyderabad
09 Aug
IntraEdge
Hyderabad
Job Title: Resiliency & Continuity Specialist (Cloud Technology Resilience)
Experience: 6+ Years
Location: Hyderabad
Employment Type: Full-Time
About the Role
We are looking for an experienced Resiliency & Continuity Specialist to join our Enterprise Technology Resilience team. In this role, you will serve as a subject matter expert responsible for planning, coordinating, governing, and validating cloud resilience exercises to ensure enterprise applications and platforms are recoverable, highly available, and compliant with organizational resilience standards.
You will partner with Engineering, SRE, Infrastructure, Cloud, Risk, and Governance teams to ensure cloud-hosted systems are resilient through comprehensive System Recovery Plans (SRPs), Disaster Recovery (DR) testing, Chaos Engineering exercises, audit-ready documentation, and continuous improvement initiatives .
The ideal candidate should have strong knowledge of cloud resilience, disaster recovery, operational resilience, cloud architecture, governance, and technology risk management in enterprise environments.
Key Responsibilities
Cloud Technology Resilience
- Lead and coordinate enterprise cloud resilience initiatives across critical applications and infrastructure.
- Plan, execute, and govern in-region and cross-region cloud resilience exercises.
- Validate that applications meet Recovery Time Objectives (RTO) and Recovery Point Objectives (RPO).
- Ensure cloud-hosted systems are resilient, recoverable, and aligned with enterprise resilience standards.
- Collaborate with Engineering, Infrastructure, Cloud Operations, and SRE teams to strengthen platform resilience.
Resilience Exercise Planning & Execution
- Coordinate end-to-end resilience testing activities for cloud applications.
- Prepare resilience test plans, execution schedules, roles, dependencies, and success criteria.
- Validate readiness before resilience testing begins.
- Support execution of:
- Disaster Recovery (DR) Testing
- Chaos Engineering Exercises
- Regional Failover Testing
- Cross-Region Recovery Testing
- Backup & Restore Validation
- Business Continuity Exercises
- Ensure testing is executed according to enterprise governance standards.
System Recovery Plan (SRP) Governance
- Review and validate System Recovery Plans (SRPs) for completeness and operational readiness.
- Verify recovery procedures, ownership, sequencing, dependencies, and execution timelines.
- Ensure SRPs follow approved enterprise templates and governance standards.
- Identify gaps and drive remediation with application owners.
- Maintain version control and governance for recovery documentation.
Evidence Validation & Audit Readiness
- Review resilience exercise evidence packages to ensure:
- End-to-end documentation
- Timestamped execution evidence
- Recovery metrics
- Pass/Fail criteria
- Recovery duration
- Validation of business objectives
- Ensure evidence aligns with enterprise audit requirements.
- Maintain audit-ready documentation supporting compliance and regulatory reviews.
- Identify missing artifacts and enforce quality standards.
Operational Resilience & Disaster Recovery
- Support enterprise Disaster Recovery (DR) planning and execution.
- Validate backup and restore procedures.
- Assess application recoverability and business continuity readiness.
- Ensure recovery capabilities align with organizational resilience objectives.
- Support annual and quarterly resilience testing cycles.
Chaos Engineering
- Support design and execution of Chaos Engineering scenarios.
- Validate resilience against:
- Infrastructure failures
- Network failures
- Service outages
- Cloud region failures
- Dependency failures
- Collaborate with engineering teams to improve system fault tolerance.
- Recommend improvements based on resilience testing outcomes.
Cloud Architecture & Availability
- Evaluate cloud architectures to ensure high availability and resilience.
- Review application deployment models for resilience best practices.
- Validate:
- Multi-Availability Zone deployments
- Cross-region replication
- Load balancing
- Auto Scaling
- Infrastructure as Code (IaC)
- Backup strategies
- Service dependencies
- Recommend architectural improvements to reduce outage risks.
Monitoring & Observability
- Support implementation of monitoring and observability standards.
- Review dashboards and monitoring configurations.
- Validate application health metrics and recovery indicators.
- Collaborate with SRE teams to improve operational visibility.
- Support incident analysis and post-incident reviews.
Governance & Risk Management
- Ensure technology resilience activities comply with enterprise governance policies.
- Support technology risk assessments related to resilience.
- Collaborate with Governance, Risk, and Compliance (GRC) teams.
- Maintain resilience assessment workbooks and governance documentation.
- Track remediation activities until closure.
Stakeholder Collaboration
- Partner with:
- Cloud Engineering Teams
- Site Reliability Engineering (SRE)
- Infrastructure Teams
- Enterprise Architecture
- Risk & Compliance
- Application Development Teams
- Business Continuity Teams
- Security Teams
- Provide guidance on resilience standards and recovery readiness.
- Facilitate resilience planning workshops and technical discussions.
Reporting & Continuous Improvement
- Prepare resilience dashboards and executive reports.
- Analyze resilience metrics and trends.
- Identify opportunities to improve recovery capabilities.
- Enhance resilience templates, standards,
and governance processes.
- Drive continuous improvement initiatives across enterprise resilience programs.
Required Qualifications
Experience
- 6+ years of experience in Technology Resilience, Disaster Recovery, Operational Resilience, SRE, Infrastructure Operations, Technology Risk, or IT Governance.
- Experience coordinating enterprise Disaster Recovery and resilience testing exercises.
- Experience working in large-scale cloud environments.
- Experience creating audit-ready technical documentation.
Technical Skills
Technology Resilience
- Operational Resilience
- Disaster Recovery Planning
- Business Continuity
- System Recovery Planning (SRP)
- Recovery Readiness
- Resilience Governance
Cloud Platforms Hands-on knowledge of:
- AWS
- Microsoft Azure
- Google Cloud Platform (GCP)
Strong understanding of:
- Regions & Availability Zones
- High Availability
- Multi-Region Architecture
- Backup & Recovery
- Auto Scaling
- Infrastructure as Code
- Cloud Networking
Chaos Engineering
Knowledge of
- Chaos Testing
- Fault Injection
- Failure Simulation
- Resilience Validation
- Recovery Testing
Infrastructure & Operations
Experience with
- Load Balancing
- Backup & Restore
- Replication
- Disaster Recovery Architecture
- Infrastructure Automation
Observability & Monitoring
Experience with
- Monitoring Tools
- Logging
- Metrics
- Dashboards
- Alerting
- Incident Management
- Observability Platforms
Governance Tools
Experience working with
- ServiceNow
- GRC Platforms
- Harness
- Risk Management Tools
Reporting Tools
Knowledge of
- Microsoft Excel
- Power BI
- Tableau
- Crystal Reports
- Microsoft Project
- Microsoft Visio
Documentation
Experience preparing
- Recovery Plans
- Audit Documentation
- Test Evidence
- Architecture Diagrams
- Operational Procedures
- Governance Documentation
Preferred Qualifications
- Experience in Banking, Financial Services, Insurance, or other highly regulated industries.
- Knowledge of Site Reliability Engineering (SRE) principles.
- Familiarity with DevOps and CI/CD environments.
- Experience working with cloud-native architectures and Kubernetes.
- Knowledge of ITIL and enterprise governance frameworks.
Certifications
Preferred certifications include
- AWS Certified Cloud Practitioner / Solutions Architect
- Microsoft Azure Certifications
- Google Cloud Certifications
- CBCP (Certified Business Continuity Professional)
- Disaster Recovery Institute (DRI) Certification
- Project Management Certification (PMP, PRINCE2, or equivalent)
- ITIL Foundation (Preferred)
Soft Skills
- Excellent analytical and problem-solving skills.
- Robust communication and stakeholder management abilities.
- Excellent documentation and organizational skills.
- Ability to coordinate cross-functional technical teams.
- Strong attention to detail and governance.
- Ability to work independently and manage multiple priorities.
📌 Resiliency & Continuity Specialist (Cloud Technology Resilience) (Hyderabad)
🏢 IntraEdge
📍 Hyderabad