02 Oct
|
Acesoft Labs
|
Hyderabad
02 Oct
Acesoft Labs
Hyderabad
Role & responsibilities
- Establish and lead the implementation of organizational reliability strategies, aligning SLAs, SLOs, and Error Budgets with business goals and customer expectations.
- Develop and institutionalize incident response frameworks, including escalation policies, on- scheduling, service ownership mapping, and RCA process governance.
- Lead technical reviews for infrastructure reliability design, high-availability architectures, and resiliency patterns across distributed cloud services. Champion observability and monitoring culture by standardizing tooling, alert definitions, dashboard templates, and telemetry data schemas across all product teams.
- Drive continuous improvement through operational maturity assessments, toil elimination initiatives, and SRE OKRs aligned with product objectives. Collaborate with cloud engineering and platform teams to introduce self-healing systems, capacity-aware autoscaling, and latency-optimized service mesh patterns.
- Act as the principal escalation point for reliability-related concerns and ensure incident retrospectives lead to measurable improvements in uptime and MTTR.
- Own runbook standardization, capacity planning, failure mode analysis,
and production readiness reviews for recent feature launches. Mentor and develop a high-performing SRE team, fostering a proactive ownership culture, encouraging cross-functional knowledge sharing, and establishing technical career pathways.
- Collaborate with leadership, delivery, and customer stakeholders to define reliability goals, track performance, and demonstrate ROI on SRE investments.
Preferred candidate profile
10+ years total experience, with 3+ years in a leadership role in SRE or Cloud Operations.
Technical Knowledge and Skills:
Mandatory:
- Deep understanding of Kubernetes, GKE, Prometheus, Terraform
- Cloud: Advanced GCP administration
- CI/CD: Jenkins, Argo CD, GitHub Actions
- Incident Management: Full lifecycle, tools like OpsGenie
Nice to Have :
- Knowledge of service mesh and observability stacks
- Strong scripting skills (Python, Bash)
- Big Query /Dataflow exposure for telemetry
-
- #### Immediate joiners can send CV to
- ******Follow/connect me on linkedln: https://www.linkedin.com/in/aman-b056091b1/
📌 Site Reliability Engineering (SRE) Manager (Hyderabad)
🏢 Acesoft Labs
📍 Hyderabad