AI INFRASTRUCTURE ENGINEER (India)

AI INFRASTRUCTURE ENGINEER (India)

30 Sep
|
International Institute of Information Technology Banglore (A4I)
|
India

30 Sep

International Institute of Information Technology Banglore (A4I)

India

AI INFRASTRUCTURE ENGINEER

CLOUD INFRASTRUCTURE & AGENTIC SYSTEMS · A4I Lab (IIIT-B × Microsoft)

Organization: A4I Lab (IIIT-B × Microsoft)

Location: IIIT Bangalore

Type: Contract (through March 2027, extendable)

Experience: 5 to 8 years

About A4I Lab

AI Innovation and Inclusion Initiative (A4I) is a partnership between Microsoft and IIIT-B, collaborating with non-profit partners to harness AI for solving real-world challenges in education, healthcare, accessibility, and agriculture. A4I evolves innovations into open-source Digital Public Goods (DPGs) that are deployable at scale, fostering a strong AI innovation ecosystem for social impact.

The Opportunity

Imagine infrastructure that doesn't just run AI — it enables it. At A4I Lab, we're not spinning up VMs and hoping for the best. We're engineering the cloud backbone that powers autonomous multi-agent systems serving millions of users in India's most underserved communities: students learning in vernacular languages, frontline health workers operating without connectivity, farmers acting on real-time data.

Generic DevOps practices don't work here. The constraints are too real — cost ceilings, data sovereignty, multilingual AI services, and uptime requirements that leave no room for reactive firefighting. What we need is an infrastructure engineer who has already run AI-specific workloads in production, not just Kubernetes and Terraform in general.

Two things worth knowing before you apply. First: general DevOps/SRE experience (Kubernetes, Terraform, CI/CD, observability) is necessary but not sufficient on its own — we specifically need hands-on production experience with the AI/ML layer on top of that (model serving, vector stores, or MLOps tooling). Second: you will effectively be a one-person infrastructure function. That means full platform ownership, mentoring other engineers on infra practices, and writing clear incident communications for leadership under pressure — not just running commands.

What You Will Be Doing

Building and Owning Cloud Infrastructure

- Design, provision, and manage multi-application cloud environments on Azure (primary), with working knowledge of AWS and GCP — across compute, networking, storage, and managed services.
- Own environment parity between staging and production for a portfolio of live agentic AI applications, ensuring consistent, reproducible deployments.
- Drive Infrastructure as Code (IaC) practices using Terraform and Ansible, eliminating manual provisioning and enabling GitOps workflows.

Enabling Reliable AI Workloads in Production

- Deploy and manage containerised AI services using Docker and Kubernetes (AKS), including LLM inference endpoints, vector store services, and RAG pipeline backends.
- Maintain CI/CD pipelines (GitHub Actions or Azure DevOps) for AI and web application workloads — with staged rollouts, automated testing gates, and rollback capabilities.
- Implement and evolve MLOps tooling: model versioning, inference monitoring, and deployment pipelines that support safe, auditable releases into high-stakes social contexts.

Platform Resilience, Security, and Incident Response

- Design and implement backup and restoration strategies across heterogeneous data stores — PostgreSQL, MongoDB Atlas, CosmosDB, and vector databases — with tested recovery procedures.
- Architect and maintain disaster recovery (DR) plans covering multi-region failover, RTO/RPO targets, and regular DR drills for all production systems.
- Own incidents end-to-end, including writing clear, structured post-mortem communications for non-technical leadership under time pressure.
- Manage Kubernetes cluster resilience: node auto-scaling, pod disruption budgets, health checks, and stateful workload recovery.

Observability and Security





- Build and maintain the end-to-end observability stack — Azure Monitor, Grafana, Prometheus, and ELK/OpenSearch — with dashboards, alerting, and incident runbooks for all A4I applications.
- Bring a proactive security instinct: spotting hardcoded secrets, privileged containers, and overly permissive mounts on sight, not just applying IAM policy from a checklist.
- Manage Cloudflare for CDN, DNS routing, and edge security across all A4I-hosted applications.

Cost Governance and Platform Strategy

- Own cloud cost visibility and optimisation — right-sizing, reserved capacity planning, and rationalisation across multi-application Azure subscriptions.
- Evaluate and recommend cloud provider and tooling decisions, balancing cost, open-source compliance, and DPG deployment requirements.
- Automate operational workflows using Python and Bash, moving the team from reactive support to proactive, self-healing infrastructure.

What You Will Not Be Doing

- Managing tickets and waiting for engineers to raise infrastructure requests — this is a proactive ownership role.
- Working on a single-application environment — you will own a portfolio of distinct AI products simultaneously.
- Operating with a large dedicated infra team behind you — you are the infrastructure function, working closely with (but not handing off to) AI architects, engineers, and partner teams.
- Maintaining legacy on-premise infrastructure — the entire stack is cloud-native.

Must-Have Skills

Each bullet represents a skill area we will probe directly during screening.

- Cloud platform ownership: Hands-on production experience with Azure, including App Services, Function Apps, AKS, Container Registry, CosmosDB, and Entra ID. Familiarity with native VMs is required; AWS and GCP experience is optional.
- Infrastructure as Code: Proficiency with Terraform and/or Ansible for managing multi-environment cloud infrastructure; you write IaC, you don't patch portals.
- Kubernetes and containers: Real production experience with Docker and Kubernetes (AKS or equivalent) — including stateful workload management, cluster scaling, and service mesh basics.
- AI/ML infrastructure — hard requirement: Hands-on production experience with at least one of: MLflow, Kubeflow, BentoML, NVIDIA Triton, Seldon, vLLM, or Ray Serve for model serving; or hands-on production deployment of vector databases (pgvector, Qdrant, Pinecone) and RAG pipeline infrastructure. General DevOps experience alone does not meet this bar.
- AI-native engineering practice: You use AI coding agents (Cursor, GitHub Copilot, Windsurf, or equivalent) as a core part of your daily workflow — including agent/composer modes and awareness of context-grounding approaches such as MCP. We're looking for genuine AI-first practice, not occasional autocomplete or copy-pasting into a chat window without reviewing the output.
- Backup, restoration, and DR: Demonstrated experience designing and testing backup strategies and disaster recovery plans across cloud-native databases and stateful services.
- CI/CD pipelines: Proven track record building and maintaining CI/CD pipelines for containerised AI/ML and web workloads using GitHub Actions, Azure DevOps, or equivalent.
- Observability: Hands-on experience with Azure Monitor, Grafana, Prometheus, and ELK/OpenSearch for production monitoring, alerting, and incident response.




- Security and compliance: Working knowledge of IAM, Entra ID, secrets management (Key Vault or equivalent), encryption practices, and PI data handling requirements — with a proactive instinct for spotting misconfigurations, not just policy knowledge.
- Python and Bash: Production-quality scripting for automation, infrastructure tooling, and operational workflows — not just glue scripts.
- Incident communication: Ability to write clear, structured incident and post-mortem communications under time pressure — evaluated directly in our process.

Qualifications

- Bachelor's or Master's in Computer Science, Engineering, or equivalent demonstrated expertise.
- 4 to 6 years in cloud infrastructure, platform engineering, or DevOps roles, with at least 2 years supporting live AI or ML workloads in production.
- No AI/ML infra experience yet? If you have strong general DevOps/SRE experience (Kubernetes, Terraform, multi-cloud) but haven't yet worked hands-on with MLOps tooling or model serving, we'd still like to hear from you for a separate, more junior AI-infra track — but please don't apply to this senior req expecting to learn the AI layer on the job, as this specific role is scoped to already have it.
- Demonstrated experience managing infrastructure for more than one live application simultaneously.
- Verifiable track record of cloud cost governance and right-sizing in real production environments — references or case studies will be requested.

Preferred Qualifications

- Hands-on experience with Azure-specific AI services: Azure OpenAI, Cognitive Services, Bhashini integration, or Sarvam AI endpoints in production.
- Familiarity with vector databases (pgvector, Qdrant, or equivalent) and the infrastructure considerations for embedding pipelines at scale.
- Prior work in open-source DPG deployments or multi-tenant cloud environments serving low-resource or low-connectivity end users.
- Experience with Cloudflare Workers, DNS management, and edge security configuration.
- Contributions to open-source infrastructure tooling, IaC libraries, or internal platform frameworks.
- Exposure to multilingual AI service infrastructure — Indic language models, ASR/TTS endpoints, or translation API pipelines.

A4I Lab is committed to building AI systems that serve everyone. We actively encourage applications from candidates of all backgrounds, disciplines, and geographies.

Pay: ₹1,500,000.00 - ₹2,800,000.00 per year

Benefits:

- Health insurance
- Provident Fund

Application Question(s):

- •Have you deployed or operated any of the following in production: MLflow, Kubeflow, BentoML, NVIDIA Triton, Seldon, vLLM, or Ray Serve?

◦Yes
◦No

- •Do you use AI coding agents (Cursor, Copilot, Windsurf) as a core part of your daily workflow, including agent modes or context-grounding techniques like MCP?

◦Yes
◦No

- •Have you deployed or operated a vector database (pgvector, Qdrant, Pinecone, or equivalent) in production?

◦Yes
◦No

- •Are you comfortable being the sole owner of infrastructure — including writing incident communications for leadership — rather than working within a larger dedicated infra team?

◦Yes
◦No

- •Have you managed cloud infrastructure for more than one live application simultaneously?•Have you managed cloud infrastructure for more than one live application simultaneously?

YES
NO

- •Have you used Azure hands-on in a production workplace?

◦Yes
◦No

- •Have you used AWS or GCP hands-on in a production environment?

◦Yes
◦No

- •Have you designed a disaster recovery plan for a live cloud system?

◦Yes
◦No

- •Have you used Kubernetes to manage your production environment?

◦Yes
◦No

Work Location: In person

📌 AI INFRASTRUCTURE ENGINEER (India)
🏢 International Institute of Information Technology Banglore (A4I)
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: ai infrastructure engineer (india) / india