01 Sep
|
Good Co India
|
India
01 Sep
Good Co India
India
Role & Responsibilities
- Lead and manage the Site Reliability Engineering (SRE) team responsible for the reliability, availability, scalability, and performance of production systems.
- Define and execute the SRE strategy, reliability roadmap, operational standards, and engineering best practices.
- Establish and continuously improve SLIs, SLOs, SLAs, error budgets, and reliability metrics across critical services.
- Design and implement highly available, scalable, resilient, and fault-tolerant cloud and distributed systems.
- Drive automation of infrastructure, deployments, monitoring, alerting, operational processes, and repetitive engineering tasks.
- Lead incident management, production support, root-cause analysis, post-incident reviews, and corrective action plans.
- Partner with Software Engineering, Platform Engineering, DevOps, Security, Product, and Architecture teams to improve system reliability.
- Design and maintain observability solutions covering monitoring, logging, tracing, metrics, dashboards, and alerting.
- Drive implementation and optimization of Kubernetes, Docker, CI/CD, Infrastructure as Code, and cloud-native platforms.
- Establish practices for capacity planning, performance optimization, disaster recovery, business continuity, and resilience testing.
- Identify reliability risks, technical debt, operational bottlenecks, and single points of failure and drive their remediation.
- Develop and enforce standards for production readiness, release management, deployment strategies, and operational excellence.
- Drive DevOps, automation, and self-service initiatives to improve developer productivity and reduce operational overhead.
- Monitor infrastructure and application performance and proactively identify opportunities to improve availability, latency,
scalability, and efficiency.
- Drive cloud infrastructure cost optimization and FinOps initiatives while maintaining reliability and performance.
- Evaluate and adopt new tools and technologies for monitoring, automation, infrastructure management, and reliability engineering.
- Hire, mentor, coach, and develop SRE, DevOps, and Platform Engineers.
- Conduct performance reviews, career planning, and technical mentoring to build a high-performing engineering team.
- Communicate reliability metrics, operational risks, incidents, and improvement initiatives to engineering leadership and senior stakeholders.
Preferred Candidate Profile
- 5 to 10 years of experience in SRE, DevOps, Cloud Engineering, Platform Engineering, Infrastructure Engineering, or Software Engineering, with demonstrated technical leadership experience.
- Proven experience leading or managing SRE, DevOps, Platform, or Infrastructure teams.
- Strong hands-on expertise in Site Reliability Engineering, Cloud Infrastructure, Distributed Systems, System Design, and Production Operations.
- Strong experience with at least one major cloud platform: AWS, Azure, or GCP.
- Hands-on expertise in Kubernetes, Docker, Terraform, Infrastructure as Code, CI/CD, Jenkins, GitLab CI, or GitHub Actions.
- Robust programming/scripting skills in Python, Go, Bash, or Shell scripting.
- Strong understanding of Linux, networking, TCP/IP,
DNS, load balancing, databases, APIs, microservices, and distributed systems.
- Proven experience implementing monitoring, logging, metrics, tracing, alerting, and observability using tools such as Prometheus, Grafana, ELK, Datadog, Splunk, New Relic, or OpenTelemetry.
- Strong understanding of SLI, SLO, SLA, error budgets, availability, latency, reliability, capacity planning, and performance optimization.
- Experience managing production incidents, on-call operations, root-cause analysis, postmortems, and disaster recovery.
- Strong knowledge of high availability, fault tolerance, resilience engineering, backup and recovery, and business continuity.
- Experience with cloud security, IAM, vulnerability management, DevSecOps, and infrastructure security.
- Strong understanding of automation, deployment strategies, release management, and operational excellence.
- Experience driving cloud cost optimization, capacity management, and FinOps initiatives is preferred.
- Proven ability to hire, mentor, coach, and develop SRE/DevOps engineers and build high-performing teams.
- Strong technical leadership, stakeholder management, communication, problem-solving, and decision-making skills.
- Ability to balance reliability, engineering velocity, security, scalability, performance, and cost.
- Bachelor's degree in Computer Science, Information Technology, Software Engineering, Electronics, or a related technical discipline.
- B.E./B.Tech/MCA/M.Tech or equivalent qualification preferred.
- Certifications such as AWS Solutions Architect/DevOps Engineer, Azure DevOps Engineer, GCP Professional Cloud DevOps Engineer, CKA/CKAD, or Terraform Associate are an added advantage.
📌 SRE Lead (India)
🏢 Good Co India
📍 India