27 Aug
|
Zorba AI
|
Bengaluru
27 Aug
Zorba AI
Bengaluru
We are looking for experienced SRE / Chaos Engineering / DevOps professionals with 5+ years of experience in cloud infrastructure, production reliability, observability, and resilience engineering. The ideal candidate will have hands-on experience with Azure, DevOps/CI-CD practices, monitoring, incident management, and failure/resilience testing.
Experience with Azure Chaos Studio is highly preferred. Candidates with strong SRE, Chaos Engineering, resilience engineering, or fault-injection experience and good Azure exposure will also be considered.
Key Responsibilities
- Design and execute resilience, failure, and chaos engineering experiments to validate system reliability.
- Work with Azure Chaos Studio to perform controlled fault-injection and resilience testing.
- Identify system weaknesses related to availability, failover, recovery, and fault tolerance.
- Develop and execute chaos experiments across cloud and containerized environments.
- Collaborate with development, DevOps, infrastructure, and SRE teams to improve system reliability.
- Monitor production environments and analyze system performance, availability, and reliability.
- Support incident management, troubleshooting, root-cause analysis, and production recovery.
- Implement and maintain CI/CD pipelines using GitHub Actions, Jenkins, or similar tools.
- Automate infrastructure and deployment activities using Terraform and Ansible.
- Work with Docker and Kubernetes/container orchestration platforms.
- Use observability and monitoring tools such as Grafana, Prometheus, CloudWatch, Nagios, Azure Monitor, or similar platforms.
- Participate in disaster recovery, failover, recovery, and business continuity testing.
- Define and improve operational resilience, reliability, and recovery practices.
- Document chaos experiments, findings, remediation actions,
and reliability improvements.
Required Skills
- 5+ years of experience in SRE, DevOps, Cloud Engineering, Reliability Engineering, or Production Engineering.
- Robust experience with Azure cloud environments.
- Knowledge or hands-on experience with Azure Chaos Studio is highly preferred.
- Strong understanding of Chaos Engineering and resilience engineering concepts.
- Experience with failover, recovery, fault tolerance, disaster recovery, and failure testing.
- Hands-on experience with CI/CD tools such as GitHub Actions or Jenkins.
- Experience with Terraform and/or Ansible.
- Experience with Docker and Kubernetes.
- Strong knowledge of monitoring and observability concepts.
- Experience with tools such as Prometheus, Grafana, Azure Monitor, CloudWatch, Nagios, or equivalent.
- Experience in production support and incident management.
- Strong troubleshooting and problem-solving skills.
Good to Have
- Azure Chaos Studio
- Chaos Mesh / LitmusChaos / Gremlin / AWS Fault Injection Simulator
- Azure Kubernetes Service (AKS)
- Azure Monitor / Application Insights
- SLO, SLA, SLI and error-budget concepts
- Disaster Recovery and Business Continuity
- GameDay / resilience testing
- Python, PowerShell, or Bash scripting
- Infrastructure as Code and automated reliability testing
Ideal Candidate Profile Candidates from the following backgrounds are encouraged to apply:
- Site Reliability Engineer (SRE)
- Chaos Engineer
- Cloud Reliability Engineer
- Reliability Engineer
- DevOps Engineer - SRE
- Platform Engineer
- Azure DevOps Engineer
- Cloud Operations Engineer
- Production Engineer
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Azure Chaos Studio_Expert (Bengaluru)
🏢 Zorba AI
📍 Bengaluru