What You Will Own
Platform Architecture
- Full architectural ownership of the non-AWS toolchain: CI/CD, observability, event streaming, automation, secrets, and deployment infrastructure
- Define, build, and enforce platform standards across portfolio products
- Terraform IaC for all infrastructure — nothing provisioned manually, everything versioned and reviewed
- Self-service developer platform so product teams ship without waiting on platform
Event Streaming & Pipeline Infrastructure
- Own the event streaming architecture, operational standards, and health monitoring across all products using real-time pipelines
- Design and maintain batch processing infrastructure alongside live event flows
- Ensure pipeline reliability, throughput, and cost are actively managed at scale
CI/CD & Deployment
- Build and maintain CI/CD pipelines (GitHub Actions) across all portfolio products
- Automate triage and retry logic for known failure classes — flaky tests, dependency timeouts, OOM kills — so engineers are only paged for genuinely novel failures
- Deployment standards: release management, rollback mechanisms, canary and blue-green patterns where justified
Observability & Reliability
- Own the full observability stack: Grafana, Prometheus, and Loki across all products
- SLOs and error budgets defined per product; reliability tracked consistently
- Build alerting that correlates signals and surfaces diagnostic context alongside notifications — so on-call engineers arrive at an incident with hypotheses,
not a blank screen
- Incident response: on-call design, escalation playbooks, post-mortem facilitation
- Automated remediation scoped to a defined set of safe, idempotent actions — container restarts, ECS task scaling, known rollback patterns. Novel or ambiguous failures escalate to a human with full context attached
Acquisition Onboarding
- Platform audit and gap analysis for every new acquisition — assessing CI/CD maturity, IaC coverage, observability gaps, and security posture
- Migration plan and execution for each portfolio company joining the platform.
- Target: full platform integration within a defined window per acquisition
What We're Looking For
Experience & Background
- 8–12 years in platform engineering, DevOps, or SRE — with clear evidence of increasing ownership over time
- Robust Terraform depth across multi-environment, multi-account setups
- CI/CD ownership across a multi-product environment with GitHub Actions
- Experience with event streaming infrastructure at production scale — design, operations, reliability, and cost management
- Hands-on Grafana, Prometheus, and Loki in production
- AWS operational depth: ECS, EKS, RDS, IAM, VPC, CloudWatch, Cost Explorer
- SRE fundamentals: SLOs, error budgets, on-call design, post-mortem culture
- Acquisition or greenfield platform integration experience strongly preferred
Skills:- Terraform, grafana, prometheus and Amazon Web Services (AWS)
📌 Senior Platform & Site Reliability Engineer (India)
🏢 Hyrhub
📍 India