08 Aug
|
Zorba AI
|
Chennai
Job Summary
We are seeking an experienced Site Reliability Engineer (SRE) to join our engineering team. The ideal candidate will have strong expertise in Azure/GCP cloud platforms, DevOps practices, Kubernetes, Infrastructure as Code, CI/CD, monitoring, and production support. The role focuses on improving application reliability, automating operations, optimizing system performance, and ensuring high availability for enterprise applications in a microservices environment.
The candidate should possess strong troubleshooting skills, experience with modern cloud-native architectures, and a passion for automation and continuous improvement.
Key Responsibilities Site Reliability Engineering
- Implement Site Reliability Engineering (SRE) best practices to improve application availability, scalability, and reliability.
- Monitor production systems and ensure service health through proactive monitoring and alerting.
- Define and maintain SLI, SLO, and Error Budgets.
- Participate in production incident management, root cause analysis (RCA), and post-incident reviews.
- Provide L2/L3 production support and participate in on-call rotations.
Cloud Infrastructure
- Deploy, manage, and maintain cloud infrastructure on Microsoft Azure and/or Google Cloud Platform (GCP).
- Manage Kubernetes environments such as AKS and GKE.
- Implement Infrastructure as Code (IaC) using Terraform and Ansible.
- Automate cloud provisioning, deployments, and infrastructure management.
DevOps & CI/CD
- Design and maintain CI/CD pipelines using GitHub Actions.
- Implement automated build, testing, deployment, and release processes.
- Improve deployment reliability through automation and DevOps best practices.
- Collaborate with development teams to enhance release quality and deployment efficiency.
Monitoring & Reliability
- Configure and maintain monitoring, logging, and alerting solutions.
- Analyze application and infrastructure logs using Splunk, Dynatrace, Grafana, or similar tools.
- Develop dashboards and monitoring metrics to ensure system reliability.
- Improve monitoring through automation and preventive measures.
Automation
- Develop automation scripts using Python or C#.
- Automate operational tasks, deployments, housekeeping, and infrastructure management.
- Improve operational efficiency by reducing manual interventions.
Production Support
- Troubleshoot complex production issues across distributed systems.
- Perform code analysis, log analysis, and performance tuning.
- Collaborate with engineering teams to resolve critical production incidents.
- Maintain production stability while minimizing downtime.
Collaboration
- Work closely with cross-functional teams including Development, DevOps, QA, and Product teams.
- Participate in Agile ceremonies and SRE governance activities.
- Share operational best practices and contribute to continuous improvement initiatives.
Mandatory Skills
- 5+ years of experience in Site Reliability Engineering, DevOps, Production Support, or Platform Engineering.
- Strong experience with Microsoft Azure and/or Google Cloud Platform (GCP).
- Hands-on experience with Kubernetes (AKS/GKE).
- Solid knowledge of DevOps practices and CI/CD pipelines.
- Experience building CI/CD workflows using GitHub Actions.
- Infrastructure as Code using Terraform and/or Ansible.
- Experience with Splunk, Dynatrace, Grafana, or similar monitoring tools.
- Robust experience with microservices architecture.
- Experience with ServiceNow or other ITSM tools.
- Knowledge of ITIL processes.
- Strong troubleshooting and production support experience.
- Experience working with SLI, SLO, Error Budget, and reliability engineering practices.
- Strong understanding of API-based architectures.
- Experience with desktop and mobile application support.
Skills: sre,devops,azure,reliability engineering
📌 Site Reliability Engineer (SRE) (Chennai)
🏢 Zorba AI
📍 Chennai