03 Oct
|
Good Co India
|
India
03 Oct
Good Co India
India
Role & responsibilities
- Define and own the Site Reliability Engineering (SRE) architecture and reliability strategy for large-scale production systems.
- Design highly available, scalable, resilient, and fault-tolerant cloud and distributed-system architectures.
- Establish reliability standards around SLIs, SLOs, SLAs, error budgets, availability, latency, and performance.
- Lead the design and implementation of Kubernetes and cloud-native platforms across AWS, Azure, or GCP.
- Architect and improve observability platforms covering monitoring, logging, distributed tracing, alerting, and service health.
- Design robust incident management, disaster recovery, business continuity, backup, and failover strategies.
- Identify and eliminate reliability bottlenecks through capacity planning, performance engineering, automation, and proactive remediation.
- Drive adoption of Infrastructure as Code, CI/CD, GitOps, and automated operational processes.
- Establish and optimize production monitoring, alerting, on-call, and incident-response practices.
- Lead complex production incidents, root-cause analysis, post-incident reviews, and long-term corrective actions.
- Evaluate system resilience through load testing, stress testing, failure testing, and chaos engineering.
- Partner with Software Engineering, Cloud, Security, Data, and Platform teams to build reliable and scalable services.
- Define technical standards for cloud architecture, Kubernetes, networking, security, observability, and reliability engineering.
- Review architecture and infrastructure designs and provide guidance on scalability, resilience, security, and operational readiness.
- Identify opportunities for cloud cost optimization and infrastructure efficiency without compromising reliability.
- Mentor SRE, DevOps, Platform, and Infrastructure engineers and promote reliability engineering best practices.
- Track reliability KPIs and communicate availability, incidents, risks, capacity, and improvement initiatives to technical leadership.
Preferred candidate profile
- 5 to 10 years of experience in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, Platform Engineering, or distributed systems, with significant architecture experience.
- Strong expertise in AWS, Azure, or GCP and cloud-native architecture.
- Deep hands-on experience with Kubernetes, Docker, Terraform, Helm, Ansible, and Infrastructure as Code.
- Strong understanding of distributed systems, microservices, networking, Linux, databases, caching, messaging, and system design.
- Strong knowledge of SRE principles, SLIs, SLOs, SLAs, error budgets, availability, reliability, and operational excellence.
- Extensive experience with Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, Datadog, Dynatrace,
or equivalent observability technologies.
- Strong experience with CI/CD, GitOps, Jenkins, GitHub Actions, GitLab CI, ArgoCD, or equivalent tools.
- Proficiency in Python, Go, Bash, or other programming/scripting languages for automation and tooling.
- Experience designing high-availability, fault-tolerant, multi-region, and disaster-recovery architectures.
- Robust knowledge of incident management, root-cause analysis, capacity planning, performance optimization, and production troubleshooting.
- Experience with chaos engineering, load testing, resilience testing, and performance engineering is highly desirable.
- Strong understanding of cloud security, IAM, secrets management, vulnerability management, and compliance.
- Experience with service mesh, Istio, API gateways, distributed tracing, and event-driven architectures is advantageous.
- Knowledge of FinOps and cloud cost optimization is desirable.
- Demonstrated ability to influence architecture and engineering decisions across multiple teams and organizational boundaries.
- Strong technical leadership, communication, problem-solving, documentation, and stakeholder-management skills.
- Bachelor's degree in Computer Science, Information Technology, Computer Engineering, Software Engineering, or a related technical field; Master's degree is a plus.
- Certifications such as AWS Solutions Architect, Google Cloud Professional Cloud Architect, Azure Solutions Architect, CKA/CKS, or equivalent are advantageous.
📌 Site Reliability Architect (India)
🏢 Good Co India
📍 India