Senior Platform & Site Reliability Engineer (India)

Senior Platform & Site Reliability Engineer (India)

22 Sep
|
Hyrhub
|
India

22 Sep

Hyrhub

India

What You Will Own

Platform Architecture

- Full architectural ownership of the non-AWS toolchain: CI/CD, observability, event streaming, automation, secrets, and deployment infrastructure
- Define, build, and enforce platform standards across portfolio products
- Terraform IaC for all infrastructure — nothing provisioned manually, everything versioned and reviewed
- Self-service developer platform so product teams ship without waiting on platform

Event Streaming & Pipeline Infrastructure

- Own the event streaming architecture, operational standards, and health monitoring across all products using real-time pipelines
- Design and maintain batch processing infrastructure alongside live event flows
- Ensure pipeline reliability, throughput, and cost are actively managed at scale

CI/CD & Deployment

- Build and maintain CI/CD pipelines (GitHub Actions) across all portfolio products
- Automate triage and retry logic for known failure classes — flaky tests, dependency timeouts, OOM kills — so engineers are only paged for genuinely novel failures
- Deployment standards: release management, rollback mechanisms, canary and blue-green patterns where justified

Observability & Reliability

- Own the full observability stack: Grafana, Prometheus, and Loki across all products
- SLOs and error budgets defined per product; reliability tracked consistently
- Build alerting that correlates signals and surfaces diagnostic context alongside notifications — so on-call engineers arrive at an incident with hypotheses,



not a blank screen
- Incident response: on-call design, escalation playbooks, post-mortem facilitation
- Automated remediation scoped to a defined set of safe, idempotent actions — container restarts, ECS task scaling, known rollback patterns. Novel or ambiguous failures escalate to a human with full context attached

Acquisition Onboarding

- Platform audit and gap analysis for every new acquisition — assessing CI/CD maturity, IaC coverage, observability gaps, and security posture
- Migration plan and execution for each portfolio company joining the platform.
- Target: full platform integration within a defined window per acquisition

What We're Looking For

Experience & Background

- 8–12 years in platform engineering, DevOps, or SRE — with clear evidence of increasing ownership over time
- Robust Terraform depth across multi-environment, multi-account setups
- CI/CD ownership across a multi-product environment with GitHub Actions
- Experience with event streaming infrastructure at production scale — design, operations, reliability, and cost management
- Hands-on Grafana, Prometheus, and Loki in production
- AWS operational depth: ECS, EKS, RDS, IAM, VPC, CloudWatch, Cost Explorer
- SRE fundamentals: SLOs, error budgets, on-call design, post-mortem culture
- Acquisition or greenfield platform integration experience strongly preferred

Skills:- Terraform, grafana, prometheus and Amazon Web Services (AWS)

📌 Senior Platform & Site Reliability Engineer (India)
🏢 Hyrhub
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior platform & site reliability engineer (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: senior platform & site reliability engineer (india) / india