10 Aug
|
High Impact Talent
|
India
10 Aug
High Impact Talent
India
Role Summary :
You will own Mynts infrastructure and delivery lifecycle end-to-end: architect and provision our AWS footprint (EC2, ECS, Lambda, RDS, S3, ECR, Route 53), containerise and orchestrate services with Docker and Kubernetes, and run the data-and-messaging backbone (PostgreSQL, Redis, Kafka).
Youll build the CI/CD pipelines that ship our Node/React/Python stack safely, tame the runtime and web tier (NVM, Pyenv, nginx/Apache, Caddy/Traefik, Certbot), and keep every workplace Dev, QA, Staging, and Production fast, observable, and cheap to run.
Above all, youll bring seasoned, opinionated judgement from multiple cloud ecosystems to recommend the most cost-effective, seamlessly-scalable path for every decision.
The AI-Augmented Edge :
At Pyvot, AI is your primary workforce. Use Claude Code, Cursor, and Gemini to draft IaC, generate pipeline configs, review infra diffs, and reason about failure modes targeting a 3x5x efficiency gain.
Practise Cross-LLM Validation: one model proposes the architecture, another stress-tests it for cost, blast-radius, and scaling limits. Your value is measured by uptime, unit economics, and how gracefully the platform scales not tickets closed.
Core Responsibilities :
Cloud Architecture & Infrastructure-as-Code : Be the Architect :
- Multi-Service AWS Footprint : Design, provision, and harden EC2, ECS, Lambda, RDS (PostgreSQL), S3, ECR, and Route 53 as reproducible infrastructure not hand-crafted snowflakes.
- Infrastructure-as-Code : Own the estate in Terraform / CloudFormation (or Pulumi / CDK) with modular, reviewed, version-controlled definitions and least-privilege IAM baked in.
- Cloud-Agnostic Judgement : Bring real experience from more than one provider (AWS plus GCP / Azure / DigitalOcean / etc.) to pick the right primitive for the job and avoid needless lock-in.
- Cost-Effective by Design : Right-size compute, exploit spot / reserved / savings plans, and continuously trim the bill without sacrificing reliability treat the cloud invoice as a product metric.
Containers, Orchestration & the Data Backbone :
- Docker & Kubernetes : Containerise every service and run it on Kubernetes (EKS or self-managed) with sane resource limits, autoscaling, health checks, and zero-downtime rollouts.
- Data & Messaging Layer : Operate PostgreSQL (replication, backups, PITR, tuning), Redis (cache / queues), and Kafka (event streaming) as first-class, monitored, recoverable services.
- Runtime & Web Tier : Manage Node (NVM) and Python (Pyenv) runtimes and front the platform with nginx / Apache and Caddy / Traefik,
with automated TLS via Certbot / ACME.
- Seamless Scaling : Build horizontal-scaling paths (load balancing, autoscaling groups, HPA) so a 100x traffic jump is a config change, not a fire-drill.
CI/CD, Delivery & Automation :
- Pipelines That Ship Safely : Co-own GitHub Actions (build, test, scan, deploy) with reproducible builds, artifact versioning in ECR, and blue-green / canary rollouts.
- Git-Based, Push-Button Deploys : Standardise deployment across every service and box no manual scp, no editing prod by hand with fast, auditable rollbacks.
- Secrets & Config : Manage secrets and environment config safely (SSM / Secrets Manager / Vault) from a single source of truth, with no credentials in git.
- Environment Parity : Keep Dev, QA, Staging, and Production consistent and disposable so what passes staging behaves in production.
Reliability, Observability & Site Reliability Engineering :
- Monitoring & Alerting : Stand up metrics, logs, and traces (CloudWatch, Prometheus / Grafana, ELK, PagerDuty) with actionable alarms not alert fatigue.
- Availability & Recovery : Define and defend SLOs, RTO / RPO targets, backups, and disaster-recovery / failover drills you have actually rehearsed.
- Incident Response : Be first responder for infrastructure incidents diagnose, mitigate, and run blameless post-mortems that harden the system.
- Boot & Persistence Hygiene : Ensure services survive reboots, scaling events, and deploys (pm2 / systemd / docker restart policies) with no silent drift.
Advisory, Standards & Team Leadership : Be the Trusted Voice :
- Architecture Recommendations : Proactively bring cost / scaling / reliability trade-off proposals to the CTO youre hired for judgement, not just execution.
- Runbooks & Standards : Maintain infrastructure runbooks, deployment standards, and on-call playbooks the whole team can follow.
- Mentoring : Level up engineers on deployment, containers, and operational best-practice; make the platform something anyone can safely ship to.
- Security-Aware Ops : Partner with the CyberSecurity Lead on hardening, IAM, network segmentation, and audit-friendly, deny-delete logging.
Roles Requirements : Intersection of Cloud, Delivery & Reliability
Were looking for a full-spectrum infrastructure engineer equal parts cloud architect, platform / DevOps engineer, and reliability practitioner: someone at ease with a Terraform diff, a Kubernetes manifest, a Postgres failover, and a 2 AM scaling alert and seasoned enough to tell us the smartest, cheapest way to run all of it.
Must-Have Skills & Experience :
- DevOps / Cloud Engineering Experience : 6 - 10 years running production infrastructure for a SaaS / cloud-native product end-to-end.
- AWS Depth : Hands-on with EC2, ECS, Lambda, RDS, S3, ECR, Route 53, IAM, and VPC networking in real production.
- Multi-Cloud Exposure : Genuine experience across more than one cloud / ecosystem (AWS plus GCP / Azure / DO / etc.) able to compare and choose, not just operate one.
- Containers & Orchestration : Strong Docker and Kubernetes (EKS or self-managed) autoscaling, rollouts, and resource management.
- IaC & CI/CD : Terraform / CloudFormation and pipeline ownership (GitHub Actions or equivalent) with automated, rollback-safe deploys.
- Data & Runtime Ops : PostgreSQL operations (backup / replication / tuning), Redis, and the web / runtime tier (nginx / Apache, Caddy / Traefik, Certbot, NVM, Pyenv).
- Cost & Scale Judgement : A track record of cutting cloud spend and scaling systems smoothly with the numbers to prove it.
Should-Have Skills & Experience :
- Event Streaming : Production experience with Kafka (or equivalent) for high-throughput, ordered event pipelines.
- Observability & SRE : CloudWatch / Prometheus / Grafana / ELK / PagerDuty, SLOs, RTO / RPO, and rehearsed DR / failover.
- Networking & TLS : Solid grasp of DNS, load balancing, reverse proxies, and certificate automation.
- Scripting & Automation : Comfortable in Bash plus Python / Node to automate anything that repeats.
- Agile & Cross-Functional Collaboration : Comfortable in Agile / Scrum delivery and working closely with engineering, product, and leadership.
Good-To-Have Skills & Experience :
- Education & Certifications : Degree from a premier institute (IITs / NITs / BITS / IIITs) and/or AWS Solutions Architect / DevOps / CKA / CKAD a strong plus.
- Security-Adjacent Ops : Familiarity with IAM hardening, secrets management, and compliance friendly logging (works alongside our Security Lead).
- AI-Augmented Tooling : Experience using AI / LLM tools (Claude Code, Cursor, Gemini) for IaC drafting, config review, and incident reasoning.
- Startup / Scale-Up Experience : Prior ownership of infrastructure through rapid growth on a lean budget.
📌 DevOps & Cloud Infrastructure Engineer (India)
🏢 High Impact Talent
📍 India