DevOps Engineer (India)

DevOps Engineer (India)

10 Aug
|
APAS
|
India

10 Aug

APAS

India

Role

As our Senior DevOps Engineer, you will own our cloud infrastructure and lead the AWS migration in-house - end to end, without leaning on an external vendor. You are strong across cloud, containers, CI/CD, security, data, and AI infrastructure (with deep specialization in at least one of them).

You have more production experience than a mid-level engineer and can walk the talk: you’ve actually built and run these systems, not just talked about them. You have genuinely “been there, done that” with production databases and real infrastructure. You can lead and mentor.

What this role is NOT

- Not an ML researcher or data scientist. You will not train, fine-tune, or invent models, and you don’t need ML math, research, or a PhD.
- Not a narrow database-only or single-tool specialist. We need range.
- Not an application developer. You partner with the product/app team; you don’t build features.
- Not a pure on-prem sysadmin. This is cloud-native and AWS-first.
- Not an “AI thought leader” role. We need an operator who is AI-literate, not a keynote speaker.

You are an anti-candidate if (please don’t apply)…

- Your AWS experience is a certification and a tutorial, not production systems you owned and paid for.
- You’ve never owned infrastructure end to end - you only followed runbooks someone else wrote.
- You’re a one-trick specialist who “just does databases” (or just one tool) and can’t operate across the stack.
- You talk a great game but can’t walk it - you can’t point to systems you actually built and ran.
- You think “DevOps” means one CI tool, or that clicking around a cloud console is infrastructure.
- You’ve never written Infrastructure as Code, or you deploy to production by hand.
- You cannot explain, simply, what a container is, what an LLM token / context window is, or the difference between managed inference (Bedrock) and self-hosted GPU (RunPod).
- You dismiss AI as hype - or you’re all AI hype with no operational depth. This role needs both feet on the ground: real ops and real AI literacy.
- You treat security as someone else’s job (hardcoded secrets, open ports, “we’ll fix it later”).
- You go quiet when things break instead of owning the incident.
- You need a large team, heavy process, or perfect specs before you can move.

What you will own

- The AWS migration and the production AWS setting after it - led by you,



in-house.
- CI/CD pipelines: automated, repeatable build → test → deploy for every service.
- Infrastructure as code for everything - no click-ops in production.
- Containerized services on ECS/Fargate (and beyond as we grow).
- Production databases: PostgreSQL/RDS operations, backups, migrations, tuning, and the graph/ vector stores we run.
- Observability: centralized logs, metrics, traces, and alerting so we see problems before customers do.
- Security baseline: secrets management, least-privilege IAM, TLS everywhere, firewalling, hardening.
- Reliability: staging → production discipline, rollbacks, backups/DR, and incident ownership.
- Deploying and operating our AI workloads (LLM services, model routing, inference) reliably and cost-effectively.

Must-have (non-negotiable)

- 5+ years experience owning production infrastructure in DevOps / SRE / Platform / Cloud roles, at a senior level - you can lead a migration and set standards, not just execute tickets.
- Deep, hands-on AWS - not a certificate, real systems and a real bill: EC2, ECS/Fargate, VPC & networking, CloudFront, S3, RDS, IAM (least-privilege), service quotas, and cost optimization. You can own an AWS migration end to end.
- Real production database experience - Postgres/RDS operations (and comfort with graph/vector stores), not just connecting to a DB someone else runs.
- Docker and container orchestration in production.
- CI/CD you designed and maintained (e.g. GitHub Actions) - automated and repeatable.
- Infrastructure as Code (Terraform / CloudFormation / CDK) - you never configure production by hand.
- Observability - you can debug a live production incident from logs, metrics, and traces.
- DevSecOps - you treat security as a design property: no hardcoded secrets, least privilege, TLS, closed ports by default.
- Breadth across the stack with depth in a specialization - you’re an all-rounder who goes deep somewhere, not a narrow specialist who only does one thing.




- Scripting/automation in Bash and Python (or Go/Node).
- You own outcomes and own incidents, including the unglamorous 2am ones.

The AI foundation we require (this is a hard requirement, and it’s specific) We are an AI company. You do not need to be a data scientist - but you must have a solid, operational foundation in AI, because you will deploy, scale, secure, and cost-control AI workloads. You must be able to, in plain English:

- Explain how a modern LLM application works - prompts, tokens, context windows, embeddings and vector search, and retrieval-augmented generation (RAG) at a working level.
- Understand the infrastructure side of AI: managed inference (AWS Bedrock), model routing / gateways (OpenRouter, LiteLLM), self-hosted GPU (e.g. RunPod) vs. managed inference, and GPU vs. CPU compute with their cost and latency trade-offs.
- Deploy and operate AI-powered services in production, handle AI-provider API keys/secrets safely, and control cost (token / compute spend) and latency.
- Reason about the reliability of AI endpoints and design around their failure modes.

Nice to have

- You’ve led a startup’s migration to AWS before, in-house.
- Knowledge graphs and agentic systems. You've built or operated a production knowledge graph (GraphDB/Neo4j/RDF) and an agentic (multi-step, tool-using) AI system, and you know how to architect both into AWS — which pieces run on Bedrock vs. self-hosted, and how they connect securely and cost-effectively in production.
- Experience with vector/graph databases (GraphDB, RDF/SPARQL, Neo4j) and semantic search.
- FinOps / cost optimization at scale.
- Multi-tenant SaaS, row-level security, and a compliance-minded posture (utility/government clients).
- Familiarity with Cloudflare (Pages/Workers/KV) and DigitalOcean.

Our current stack AWS (EC2, Fargate, CloudFront, RDS, Bedrock, RunPod for GPU) · Cloudflare (Pages, Workers, KV) · DigitalOcean · Docker · GitHub Actions · Terraform · PostgreSQL / Supabase · GraphDB (semantic graph store) · Anthropic Claude with model routing (OpenRouter / LiteLLM).

Location

India, Remote

Compensation

- $1,500-$2,500 base salary plus performance bonuses
- Final salary will be based on a number of factors including level, relevant prior experience, skills, and expertise. This range is only inclusive of base salary, not benefits.

📌 DevOps Engineer (India)
🏢 APAS
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: devops engineer (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: devops engineer (india) / india