03 Oct
|
Good Co India
|
India
03 Oct
Good Co India
India
Role & Responsibilities
- Lead the Site Reliability Engineering function and ensure the availability, reliability, scalability, and performance of critical applications and infrastructure.
- Define and implement SRE practices around SLIs, SLOs, SLAs, error budgets, and service reliability.
- Design and operate highly available, fault-tolerant, and scalable cloud-native systems.
- Lead incident response, troubleshooting, root-cause analysis, and post-incident reviews for critical production issues.
- Build and improve monitoring, observability, logging, alerting, and performance-management capabilities.
- Drive automation of infrastructure, deployment, operational, and repetitive engineering processes.
- Manage and optimize cloud infrastructure across AWS, Azure, GCP, or hybrid environments.
- Implement Infrastructure as Code using tools such as Terraform and Ansible.
- Manage containerized and orchestration platforms including Docker and Kubernetes.
- Establish and improve CI/CD pipelines, deployment strategies, release automation, and rollback mechanisms.
- Lead capacity planning, performance optimization, disaster recovery, backup, and business-continuity initiatives.
- Identify and reduce operational risks, technical debt, toil, and recurring production issues.
- Implement reliability and resilience engineering practices, including failure testing and chaos engineering where appropriate.
- Partner with development, infrastructure, security, and product teams to improve application reliability throughout the software lifecycle.
- Establish SRE standards, operational runbooks, documentation, and engineering best practices.
- Mentor SRE/DevOps engineers and provide technical leadership across reliability initiatives.
- Track infrastructure and cloud costs and identify opportunities for performance and cost optimization.
Preferred Candidate Profile
- 5 to 10 years of experience in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, or a related field.
- Strong hands-on experience managing production systems and large-scale, distributed applications.
- Strong proficiency in Linux and scripting/programming languages such as Python, Go, or Bash.
- Extensive experience with AWS, Azure, GCP, or hybrid-cloud environments.
- Strong experience with Kubernetes, Docker, Terraform, Ansible, and Infrastructure as Code.
- Strong understanding of distributed systems, microservices, networking, databases, and cloud architecture.
- Hands-on experience with observability platforms such as Prometheus, Grafana, ELK, Splunk, or OpenTelemetry.
- Strong understanding of SLI, SLO, SLA, error budgets, availability, latency, reliability, and performance metrics.
- Experience designing and operating highly available and fault-tolerant systems.
- Strong experience in incident management, troubleshooting, root-cause analysis, and production support.
- Experience building and managing CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI, or similar tools.
- Knowledge of disaster recovery, business continuity, capacity planning, security, IAM, and cloud cost optimization.
- Experience with automation, chaos engineering, performance engineering, and reducing operational toil is preferred.
- Demonstrated ability to lead technical initiatives, mentor engineers, and work effectively across engineering teams.
- Solid analytical, problem-solving, communication, and stakeholder-management skills.
- Bachelors degree in Computer Science, Information Technology, Computer Engineering, or a related discipline; M.Tech/MS/MCA is an advantage.
- Relevant certifications such as AWS/Azure/GCP, CKA/CKAD, Terraform, or ITIL are an advantage.
📌 Site Reliability Engineer Lead (India)
🏢 Good Co India
📍 India