09 Sep
|
SKILLVENTORY
|
Guwahati
09 Sep
SKILLVENTORY
Guwahati
Key Responsibilities
- Reliability &
- Operations:
Own end-to-end production reliability, scalability, and performance. Establish SLIs, SLOs, SLAs, and Error Budgets while managing MTTR/MTTA through 24x7 on-call rotations and post-incident analysis.
- Cloud &
- Infrastructure:
Design, deploy, and manage GCP resources (GKE, VPC, IAM, Load Balancers, BigQuery, Pub/Sub) using Terraform (IaC).
- Kubernetes &
- Containers:
Deploy and troubleshoot containerized workloads on GKE using Helm, YAML, and Canary/Blue-Green deployment strategies.
- Automation &
- CI/CD:
Build automation pipelines using Jenkins, GitHub, Python, and Shell scripting to reduce manual toil.
- Observability &
- Monitoring:
Implement full-stack monitoring and alerting strategies using Dynatrace, Grafana, and cloud logging/metrics tools.
- Troubleshooting: Perform root-cause analysis across distributed systems, Linux/networking layers, and Java or Golang applications.
Key Requirements
- Experience: 11+ years in SRE,
DevOps, or Cloud Engineering managing high-availability, 24x7 production environments.
- Cloud &
- Tools:
Strong hands-on experience with GCP, GKE, and Infrastructure as Code using Terraform.
- CI/CD &
- Scripting:
Proficiency in Jenkins pipeline creation, GitHub, Python, and Shell scripting.
- Observability: Experience using Dynatrace, Grafana, or Prometheus for log, metric, and trace analysis.
- Systems &
- Networking:
Strong Linux system administration skills, TCP/IP networking fundamentals, and experience debugging Java/Golang microservices.
- SRE Concepts: Solid understanding of SLIs/SLOs, Error Budgets, and contemporary incident management practices.
- Pluses: Experience in banking/financial domains, high-scale distributed systems, and cloud security/compliance
📌 Lead Site Reliability Engineer (SRE) (Guwahati)
🏢 SKILLVENTORY
📍 Guwahati