Site Reliability Architect (India)

Site Reliability Architect (India)

03 Oct
|
Good Co India
|
India

03 Oct

Good Co India

India

Role & responsibilities

- Define and own the Site Reliability Engineering (SRE) architecture and reliability strategy for large-scale production systems.
- Design highly available, scalable, resilient, and fault-tolerant cloud and distributed-system architectures.
- Establish reliability standards around SLIs, SLOs, SLAs, error budgets, availability, latency, and performance.
- Lead the design and implementation of Kubernetes and cloud-native platforms across AWS, Azure, or GCP.
- Architect and improve observability platforms covering monitoring, logging, distributed tracing, alerting, and service health.
- Design robust incident management, disaster recovery, business continuity, backup, and failover strategies.
- Identify and eliminate reliability bottlenecks through capacity planning, performance engineering, automation, and proactive remediation.
- Drive adoption of Infrastructure as Code, CI/CD, GitOps, and automated operational processes.
- Establish and optimize production monitoring, alerting, on-call, and incident-response practices.
- Lead complex production incidents, root-cause analysis, post-incident reviews, and long-term corrective actions.
- Evaluate system resilience through load testing, stress testing, failure testing, and chaos engineering.
- Partner with Software Engineering, Cloud, Security, Data, and Platform teams to build reliable and scalable services.
- Define technical standards for cloud architecture, Kubernetes, networking, security, observability, and reliability engineering.




- Review architecture and infrastructure designs and provide guidance on scalability, resilience, security, and operational readiness.
- Identify opportunities for cloud cost optimization and infrastructure efficiency without compromising reliability.
- Mentor SRE, DevOps, Platform, and Infrastructure engineers and promote reliability engineering best practices.
- Track reliability KPIs and communicate availability, incidents, risks, capacity, and improvement initiatives to technical leadership.

Preferred candidate profile

- 5 to 10 years of experience in SRE, DevOps, Cloud Engineering, Infrastructure Engineering, Platform Engineering, or distributed systems, with significant architecture experience.
- Strong expertise in AWS, Azure, or GCP and cloud-native architecture.
- Deep hands-on experience with Kubernetes, Docker, Terraform, Helm, Ansible, and Infrastructure as Code.
- Strong understanding of distributed systems, microservices, networking, Linux, databases, caching, messaging, and system design.
- Strong knowledge of SRE principles, SLIs, SLOs, SLAs, error budgets, availability, reliability, and operational excellence.
- Extensive experience with Prometheus, Grafana, OpenTelemetry, ELK/OpenSearch, Datadog, Dynatrace,



or equivalent observability technologies.
- Strong experience with CI/CD, GitOps, Jenkins, GitHub Actions, GitLab CI, ArgoCD, or equivalent tools.
- Proficiency in Python, Go, Bash, or other programming/scripting languages for automation and tooling.
- Experience designing high-availability, fault-tolerant, multi-region, and disaster-recovery architectures.
- Robust knowledge of incident management, root-cause analysis, capacity planning, performance optimization, and production troubleshooting.
- Experience with chaos engineering, load testing, resilience testing, and performance engineering is highly desirable.
- Strong understanding of cloud security, IAM, secrets management, vulnerability management, and compliance.
- Experience with service mesh, Istio, API gateways, distributed tracing, and event-driven architectures is advantageous.
- Knowledge of FinOps and cloud cost optimization is desirable.
- Demonstrated ability to influence architecture and engineering decisions across multiple teams and organizational boundaries.
- Strong technical leadership, communication, problem-solving, documentation, and stakeholder-management skills.
- Bachelor's degree in Computer Science, Information Technology, Computer Engineering, Software Engineering, or a related technical field; Master's degree is a plus.
- Certifications such as AWS Solutions Architect, Google Cloud Professional Cloud Architect, Azure Solutions Architect, CKA/CKS, or equivalent are advantageous.

📌 Site Reliability Architect (India)
🏢 Good Co India
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability architect (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: site reliability architect (india) / india