03 Sep
|
TechBlocks
|
Hyderabad
03 Sep
TechBlocks
Hyderabad
Job Summary
Position Title: SR SRE
Experience: 6 - 8 years
Job Location: Hyderabad/Ahmedabad
Work Mode: Hybrid Mode
Time Zone/Shift: Starts at 12:30 PM IST
Requirements
68 years of experience in SRE or Infrastructure Engineering within cloud-native environments, with strong expertise in GCP (GKE, Load Balancing, VPN, IAM), Kubernetes, and Docker. Hands-on experience with observability tools such as Prometheus, Grafana, ELK, and Datadog, along with Terraform and Helm for infrastructure automation. Solid background in incident management, RCA, on-call support, SLIs/SLOs, error budgets, and reducing MTTR.
Experience in designing highly available, scalable, and resilient platforms, with a focus on automation, performance optimization, capacity planning, and operational excellence.
Responsibilities
- Define and measure Service Level Indicators (SLIs), Service Level Objectives (SLOs), and manage error budgets across services.
- Lead incident management for critical production issues drive root cause analysis (RCA) and postmortems.
- Create and maintain runbooks and standard operating procedures for high availability services.
- Design and implement observability frameworks using ELK, Prometheus, and Grafana; drive telemetry adoption.
- Coordinate cross-functional war-room sessions during major incidents and maintain response logs.
- Develop and improve automated system recovery, alert suppression, and escalation logic.
- Use GCP tools like GKE, Cloud Monitoring, and Cloud Armor to improve performance and security posture.
- Collaborate with DevOps and Infrastructure teams to build highly available and scalable systems.
- Analyze performance metrics and conduct regular reliability reviews with engineering leads.
- Participate in capacity planning, failover testing, and resilience architecture reviews.
Who you'll be working with
- SRE, DevOs, Infrastructure, and Engineering Teams
- Engineering Leads and Technical Architects
- Cloud and Platform Engineering Teams
- Cross-functional teams during critical production incidents
- Security and Operations teams
- Product and Technology stakeholders
- Teams focused on observability, automation, and platform resilience
Required Skills Qualifications
- 68 years of experience in SRE or Infrastructure Engineering.
- Strong hands-on experience with GCP, GKE, Kubernetes, Docker, Terraform, and Helm.
- Proficiency in Prometheus, Grafana, ELK, and Datadog for observability and monitoring.
- Strong understanding of SLIs, SLOs, error budgets, incident management, RCA, and on-call operations.
- Experience in designing and maintaining highly available, scalable, and resilient cloud-native systems.
- Strong problem-solving, troubleshooting, and communication skills.
- Experience with PagerDuty/OpsGenie and automated incident response.
Preferred / Nice-to-Have Skills
- Familiarity with programming languages
- Define product strategy by connecting the dots
- Make product decisions without ambiguity in collaboration with the team
- Manage important stakeholders across the organization while leading initiatives
- Decision-making and accountability/ownership is critical
- One point of executive customer engagement for a large portfolio
- Contribute effectively outside your comfort zone
📌 Senior SRE Engineer (Hyderabad)
🏢 TechBlocks
📍 Hyderabad