Senior Site Reliability Engineer (India)

Senior Site Reliability Engineer (India)

25 Aug
|
Mindsprint
|
India

25 Aug

Mindsprint

India

Role Summary :

We are hiring a senior Site Reliability Engineer to own the availability, scalability, performance and operability of our production platform. The estate is AWS-first with a growing Azure footprint, and is fully provisioned as code using Terraform and CloudFormation. This is a hands-on engineering role covering all pillars of SRE infrastructure, automation, observability, SLO management, incident response, resilience and security with mentoring responsibility for mid-level engineers and participation in a shared on-call rotation.

Key Responsibilities :

- Cloud Infrastructure (AWS primary, Azure secondary): Architect, build and operate scalable AWS infrastructure (VPC and connectivity, IAM, compute, storage, managed databases, serverless, multi-account governance) and own production Amazon EKS clusters end to end provisioning, upgrades, autoscaling, networking and security. Build and support the secondary Azure footprint and drive cloud cost optimisation.
- Infrastructure as Code: Design and maintain reusable Terraform modules and CloudFormation stacks across multi-account and multi-subscription environments, with remote state, versioning, drift detection and policy-as-code guardrails. Eliminate manual console changes.
- Monitoring &
- Observability: Own the metrics, logs and tracing stack; build symptom- and SLO-burn-based alerting; instrument services with development teams; and maintain golden dashboards and runbooks that any on-call engineer can use.
- SLIs, SLOs &
- Error Budgets: Define SLIs for critical user journeys, agree SLOs with stakeholders, operate error-budget policy to balance feature velocity against reliability, and report reliability to leadership through regular service reviews.
- Automation &
- Toil Reduction: Measure and engineer away repetitive operational work; build tooling in Python, Bash or Go for provisioning, remediation, patching and diagnostics; and implement self-healing and auto-remediation for known failure modes.
- Incident Response &
- On-Call:



Participate in a compensated on-call rotation, act as Incident Commander for high-severity events, run blameless postmortems with corrective actions tracked to closure, and drive down MTTD and MTTR.
- CI/CD &
- Release Engineering: Build and harden application and infrastructure pipelines; enable blue-green, canary and progressive delivery with automated health gates and rapid rollback; and improve DORA metrics including change failure rate.
- Backup, DR &
- Business Continuity: Own the backup estate across Azure Backup (Recovery Services vaults, retention, immutability, cross-region restore) and AWS Backup; define RPO/RTO per service; and prove recoverability through scheduled restore tests and DR failover drills.
- Capacity Planning &
- Performance: Forecast demand and plan capacity across compute, storage, network and database tiers; run load and stress testing to validate headroom and autoscaling; tune performance and practise chaos engineering to validate resilience assumptions.
- Security, Compliance &
- Governance: Enforce least-privilege IAM and centralised secrets management, maintain patch and vulnerability hygiene across hosts and images, and support audit requirements (SOC 2 / ISO 27001 / PCI-DSS) with CIS-aligned hardened baselines.

Must-Have Skills:

- AWS (Primary): EC2, VPC & network design, IAM, S3, RDS/Aurora, ELB/ALB, Route 53, Lambda, CloudWatch, CloudTrail, AWS Backup, Organizations / multi-account governance.
- Kubernetes / EKS: Production EKS ownership cluster provisioning & upgrades, Karpenter / Cluster Autoscaler, HPA, Helm, ingress, IRSA, RBAC, network policies,



deep troubleshooting (CNI, CoreDNS, CSI).
- Azure (Secondary): Azure Backup &
- Recovery Services vaults (mandatory), Azure Site Recovery, VMs, VNets, Entra ID, Storage, AKS, Azure Monitor / Log Analytics.
- IaC: Terraform at expert level (modules, remote state, workspaces, drift detection, CI-driven plan/apply) and strong AWS CloudFormation (nested stacks, StackSets, change sets).
- Observability: Prometheus, Grafana, CloudWatch, Azure Monitor, ELK/OpenSearch, OpenTelemetry; plus one APM (Datadog / New Relic / Dynatrace). SLO and error-budget dashboards.
- SRE Practice: Hands-on SLI/SLO definition, error-budget policy, incident command, blameless postmortems, toil reduction, capacity planning, chaos/DR testing.
- CI/CD: GitHub Actions, GitLab CI, Jenkins or Azure DevOps
- ArgoCD or Flux; blue-green and canary deployment with automated rollback.

Qualifications:

- Experience: 812 years in IT infrastructure, cloud or DevOps, including a minimum of 5 years in a dedicated SRE / DevOps / Cloud Infrastructure engineering role with production ownership.
- Education: Bachelor's degree in Computer Science, Engineering or equivalent practical experience.
- Certifications (preferred): AWS Solutions Architect / DevOps Engineer Professional
- CKA or CKS
- Azure AZ-104 or AZ-305
- HashiCorp Terraform Associate.
- Soft skills: Strong ownership, calm and structured under production pressure, data-driven on reliability trade-offs, and able to explain risk clearly to non-technical stakeholders.

Good to Have:

- Go for operational tooling or Kubernetes operators
- AWS CDK or Pulumi; service mesh (Istio / Linkerd / App Mesh) in production.
- Chaos engineering tooling (AWS FIS, Chaos Mesh, Gremlin); database reliability engineering (Aurora, PostgreSQL, DynamoDB)
- Kafka / MSK / Kinesis at scale.
- FinOps and cloud cost ownership; data-centre-to-cloud or AWS-to-Azure migration experience; regulated-industry exposure.

📌 Senior Site Reliability Engineer (India)
🏢 Mindsprint
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: senior site reliability engineer (india) / india