Site Reliability Engineer (SRE) / DevOps Engineer (India)

Site Reliability Engineer (SRE) / DevOps Engineer (India)

02 Aug
|
QTrino Labs
|
India

02 Aug

QTrino Labs

India

We're QTrino Labs, a research-driven cybersecurity company focused on post-quantum readiness. We're building a next-generation platform that helps enterprises manage cryptographic controls and prepare for a quantum-safe future. As a Site Reliability Engineer / DevOps Engineer, you'll own the reliability, delivery, and cost-efficiency of a multi-tenant SaaS platform running on Kubernetes. You'll keep a Go macroservice backend and Next.js frontend online, fast, and observable; owning the path from merged PR to production and the infrastructure underneath it. Equal parts builder (IaC, pipelines, automation) and operator (on-call, incident response, capacity, cost).

Skillset Required:

Cloud Infrastructure (AWS + Kubernetes)

- 4+ yrs running production workloads on AWS and Kubernetes(managed control plane EKS, GKE, or AKS).
- Hands-on with Infrastructure-as-Code - AWS CDK (TypeScript) and/or Terraform/OpenTofu. You write it, review diffs, and reason about blast radius before applying.
- Core AWS services:VPC/networking, IAM (incl. OIDC federation + IRSA), RDS/Aurora, S3, ACM, Route53, KMS, SSM, ALB/WAF.
- Helmfor packaging and releasing K8s workloads; comfortable with ingress (Traefik),cert-manager, and pod-level controls (HPA, readiness/liveness probes, resource requests/limits).

Release Pipelines / CI-CD

- Built and maintained GitHub Actions(or equivalent) pipelines: build, test, scan, publish, deploy.
- OIDC-based cloud authfrom CI (no long-lived keys), container build + push toECR, multi-service/matrix builds.
- Supply-chain security in the pipeline:image scanning (Trivy), SBOM generation (CycloneDX), image signing (Cosign), dependency scanning (govulncheck/Dependabot).
- Release discipline: semver/tagging, environment promotion (dev -> staging -> prod), rollback strategy (helm rollback, re-deploy prior tag).

Observability & Monitoring

- Operate anOpenTelemetry-based stack: OTEL Collector+Grafana (Loki / Tempo / Mimir / Prometheus).




- Build dashboards and alerts on golden signals (latency, traffic, errors, saturation); instrument and debug distributed traces.
- Define and track SLO/SLI/error budgets; turn telemetry into actionable alerting (not noise).

Keeping Systems Online (Reliability)

- Production on-call experience: incident response, triage, mitigation, blameless postmortems.
- Designed for HA/failure: rolling deployments, autoscaling, graceful shutdown/draining, health checks.
- Disaster recovery: backup/restore procedures, RTO/RPO thinking, DB snapshot restore drills.

Patching & Maintenance

- Own patch lifecycle: base image updates, dependency bumps, K8s/node version upgrades, CVE remediation.
- Automate it (Dependabot/Renovate + scan gates) rather than chase it manually.

Cost Monitoring (FinOps)

- Monitor and optimize cloud spend: AWS Cost Explorer / Budgets / anomaly detection, resource tagging, rightsizing.
- Tune cost-sensitive services - Fargate sizing, Aurora Serverless v2 scaling (ACUs), S3 lifecycle, data transfer.
- Set up cost attribution per environment/tenant and surface trends to engineering leadership.

Robust Nice-to-Have

- Go familiarity (read/debug services, write tooling) - backend is a Go monorepo.
- PostgreSQL operations: connection pooling (RDS Proxy / PgBouncer), schema migrations as K8s Jobs (gooseor similar), read-replica scaling, Aurora ops.
- NATS JetStream(or Kafka/RabbitMQ) - operating a durable messaging/eventing layer.
- Multi-tenant SaaS operational experience: tenant isolation, per-tenant resource/cost attribution.
- Cloud-agnostic / on-prempackaging: K3s/K3d, CloudNativePG, SeaweedFS, OpenTofu running the same stack off AWS.
- GitOpsdirection: ArgoCD/Flux, progressive delivery (canary/blue-green) - where the platform is heading.
- AuthN/Z infra: Cognito(or other OIDC IdP),OpenFGA/Zanzibar-style authz.
- Edge/security: WAF, AWS Shield, rate limiting, secrets management, mTLS/service mesh.
- Writing and maintaining runbooks and operational docs.

Pay: From ₹1,000,000.00 per year

Work Location: In person

📌 Site Reliability Engineer (SRE) / DevOps Engineer (India)
🏢 QTrino Labs
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (sre) / devops engineer (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: site reliability engineer (sre) / devops engineer (india) / india