10 Sep
|
SKILLVENTORY
|
Guwahati
10 Sep
SKILLVENTORY
Guwahati
Key Responsibilities
Reliability &
Operations:
Own end-to-end production reliability, scalability, and performance. Establish SLIs, SLOs, SLAs, and Error Budgets while managing MTTR/MTTA through 24x7 on-call rotations and post-incident analysis.
Cloud &
Infrastructure:
Design, deploy, and manage GCP resources (GKE, VPC, IAM, Load Balancers, BigQuery, Pub/Sub) using Terraform (IaC).
Kubernetes &
Containers:
Deploy and troubleshoot containerized workloads on GKE using Helm, YAML, and Canary/Blue-Green deployment strategies.
Automation &
CI/CD:
Build automation pipelines using Jenkins, GitHub, Python, and Shell scripting to reduce manual toil.
Observability &
Monitoring:
Implement full-stack monitoring and alerting strategies using Dynatrace, Grafana, and cloud logging/metrics tools.
Troubleshooting: Perform root-cause analysis across distributed systems, Linux/networking layers, and Java or Golang applications.
Key Requirements
Experience:
11+ years in SRE, DevOps, or Cloud Engineering managing high-availability, 24x7 production settings.
Cloud &
Tools:
Solid hands-on experience with GCP, GKE, and Infrastructure as Code using Terraform.
CI/CD &
Scripting:
Proficiency in Jenkins pipeline creation, GitHub, Python, and Shell scripting.
Observability: Experience using Dynatrace, Grafana, or Prometheus for log, metric, and trace analysis.
Systems &
Networking:
Strong Linux system administration skills, TCP/IP networking fundamentals, and experience debugging Java/Golang microservices.
SRE Concepts: Solid understanding of SLIs/SLOs, Error Budgets, and contemporary incident management practices.
Pluses: Experience in banking/financial domains, high-scale distributed systems, and cloud security/compliance
📌 Lead Site Reliability Engineer Sre Guwahati
🏢 SKILLVENTORY
📍 Guwahati