- Work as a Site Reliability Engineer to maintain and support production applications and infrastructure on GCP.
- Monitor the health, performance, availability, and reliability of production systems.
- Handle production support, incident management, troubleshooting, and issue resolution.
- Develop and maintain DevOps/CI-CD pipelines for application deployment and automation.
- Use Python scripting to automate repetitive tasks and improve operational efficiency.
- Work with Kubernetes/GKE, Docker, and Terraform for containerization and infrastructure automation.
- Perform Root Cause Analysis (RCA) for production issues and implement permanent fixes.
- Work with development and infrastructure teams to improve system stability, monitoring, and deployment processes.