- We are looking for an experienced Site Reliability Engineer SRE Production Engineer to build and operate highly available scalable and resilient production systems
- The ideal candidate will have solid expertise in cloud infrastructure automation observability incident management performance engineering and DevOps practices
- The role requires close collaboration with software engineering infrastructure security and platform teams to ensure operational excellence and improve system reliability at scale
Key Responsibilities:
- Reliability Engineering
- Design build and maintain highly available and fault tolerant production systems
- Define and monitor SLIs SLOs and SLAs for critical services
- Drive reliability improvements through automation and proactive engineering
- Conduct capacity planning and performance optimization activities
- Production Support Operations
- Manage production environments and ensure service uptime
- Lead incident response troubleshooting and root cause analysis RCA
- Develop runbooks operational playbooks and disaster recovery procedures
Technical Requirements:
- Cloud Infrastructure
- Deploy and manage cloud native infrastructure across AWS Azure or GCP
- Automate infrastructure provisioning using Infrastructure as Code IaC
- Implement scalable and secure infrastructure solutions
- Support Kubernetes based platforms and containerized workloads