28 Aug
|
i2b Technologies
|
Hyderabad
28 Aug
i2b Technologies
Hyderabad
Senior DevOps / Cloud Infrastructure Engineer (AWS + Cloudflare)
Experience: 3–5 Years Location: [Hyderabad / Remote-India — edit as needed] Employment type: Full time
About the Role
We're hiring a hands-on Senior DevOps / Cloud Infrastructure Engineer who owns networking, cost, security, and observability end-to-end across AWS and Cloudflare. This is not a role where you consume dashboards or Terraform someone else wrote — you'll design the VPC, defend the NAT-vs-endpoint cost tradeoffs, author the Terraform modules, build the Grafana dashboards from scratch, and carry the pager when it all breaks at 3am.
Must-Have Experience: 3–5 Years
Networking — Cloudflare + AWS
- VPC design; VPC endpoints (PrivateLink) vs NAT routing decisions
- NAT gateways with per-component cost analysis (data-processing/egress, per-AZ tradeoffs, endpoint offload to cut cost)
- Service discovery (Cloud Map / K8s DNS / Consul)
- Multi-region (Transit Gateway, cross-region connectivity, latency routing, failover)
- Load balancing: ALB / NLB (NLB for low-latency sockets)
- Cloudflare: CDN, DNS, WAF, DDoS, caching (R2 / Tunnel / Workers as relevant)
- Log & Media Archival
Compute, IaC & Cost
- EC2, ECS/EKS, Fargate, Auto Scaling; Docker, Kubernetes, Helm
- Terraform (strong): reusable modules, remote state, multi-env/multi-region
- Scripting: Python, Bash
- Component-level FinOps: cost breakdown across NAT, VPC endpoints, cross-region transfer, Cloudflare vs AWS egress, RDS, LBs; Cost Explorer/Budgets, right-sizing, Savings Plans/RIs,
tagging
Security & Reliability
- IAM least-privilege, KMS, secrets management, TLS/cert management
- Cloudflare WAF/DDoS + AWS security groups; CloudTrail; data-residency (India DPDP; GDPR)
- Multi-region HA/DR (RTO/RPO, backups, restore drills); capacity & load-test planning (closes the NFR gap)
Observability — Hands-On Creation, Not Just Usage
- OTel instrumentation across services
- Metric creation: custom app + infra metrics, Prometheus exporters/PromQL
- Dashboard & alert creation: Grafana dashboards, Prometheus/Alertmanager alert rules
- Logs: Loki or ELK (either); distributed tracing
Incident Management & On-Call Ownership
- PagerDuty (or Opsgenie): alert ownership, routing & escalation policies
- Post-incident RCA / postmortems; SLOs and error budgets
- Tuning alerts to reduce noise / false pages
What Success Looks Like in This Role
- You can explain, with real numbers, why a given subnet uses a NAT gateway vs a VPC endpoint — and what it costs either way
- Your Terraform modules are reused across environments and regions, not copy-pasted
- Your Grafana dashboards and Prometheus alert rules were built by you, not inherited from a vendor template
- When you get paged, you know why within minutes — and your postmortems change something
- You can walk a security or compliance review through IAM boundaries, KMS usage, and data-residency posture without prep
Skills:- Amazon Web Services (AWS) and cloudflare
📌 Cloud Infrastructure Engineer (Hyderabad)
🏢 i2b Technologies
📍 Hyderabad