05 Oct
|
Nodeascend Technologies
|
Faridabad
05 Oct
Nodeascend Technologies
Faridabad
Own uptime, latency and cost across every platform we run. Set the SLOs, build the tooling, and lead the response when something breaks.
View full description Hide description
What you will do
Define SLOs and error budgets per service, and hold delivery teams to them.
Run the incident process end to end: paging, comms, mitigation, and a blameless write-up that changes something.
Build the observability stack - metrics, traces, logs, alerts that fire on symptoms rather than causes.
Drive infrastructure cost down without trading away headroom.
Mentor engineers on operability so reliability is designed in, not bolted on.
What we are looking for
7 years in SRE, DevOps or platform engineering, including on-call ownership of a production system.
Fluent in Kubernetes and Terraform - you have built the cluster,
not just deployed to one.
Solid scripting in Python, Go or Bash, and comfort reading application code in a language you do not write.
Hands-on with Prometheus, Grafana and a tracing stack; opinionated about what deserves an alert.
Has led an incident review that a team actually acted on.
Stack Kubernetes Terraform Prometheus Grafana AWS GCP Go
Disclaimer: This job posting has been aggregated from external source. Role details, content, and availability are subject to change. Applicants are advised to confirm the latest information directly on the company website before applying.
📌 Site Reliability Manager Faridabad
🏢 Nodeascend Technologies
📍 Faridabad