01 Oct
|
International Institute of Information Technology Banglore (A4I)
|
Electronic City
01 Oct
International Institute of Information Technology Banglore (A4I)
Electronic City
AI INFRASTRUCTURE ENGINEER
CLOUD INFRASTRUCTURE & AGENTIC SYSTEMS · A4I Lab (IIIT-B × Microsoft)
Organization: A4I Lab (IIIT-B × Microsoft)
Location: IIIT Bangalore
Type: Contract (through March 2027, extendable)
Experience: 5 to 8 years
About A4I Lab
AI Innovation and Inclusion Initiative (A4I) is a partnership between Microsoft and IIIT-B, collaborating with non-profit partners to harness AI for solving real-world challenges in education, healthcare, accessibility, and agriculture. A4I evolves innovations into open-source Digital Public Goods (DPGs) that are deployable at scale, fostering a strong AI innovation ecosystem for social impact.
The Chance
Imagine infrastructure that doesn't just run AI — it enables it. At A4I Lab, we're not spinning up VMs and hoping for the best. We're engineering the cloud backbone that powers autonomous multi-agent systems serving millions of users in India's most underserved communities: students learning in vernacular languages, frontline health workers operating without connectivity, farmers acting on real-time data.
Generic DevOps practices don't work here. The constraints are too real — cost ceilings, data sovereignty, multilingual AI services, and uptime requirements that leave no room for reactive firefighting. What we need is an infrastructure engineer who has already run AI-specific workloads in production, not just Kubernetes and Terraform in general.
Two things worth knowing before you apply.
First: general DevOps/SRE experience (Kubernetes, Terraform, CI/CD, observability) is necessary but not sufficient on its own — we specifically need hands-on production experience with the AI/ML layer on top of that (model serving, vector stores, or MLOps tooling).
Second: you will effectively be a one-person infrastructure function. That means full platform ownership, mentoring other engineers on infra practices, and writing clear incident communications for leadership under pressure — not just running commands.
What You Will Be Doing
Building and Owning Cloud Infrastructure
- Design, provision, and manage multi-application cloud environments on Azure (primary), with working knowledge of AWS and GCP — across compute, networking, storage, and managed services.
- Own environment parity between staging and production for a portfolio of live agentic AI applications, ensuring consistent, reproducible deployments.
- Drive Infrastructure as Code (IaC) practices using Terraform and Ansible, eliminating manual provisioning and enabling GitOps workflows.
Enabling Reliable AI Workloads in Production
- Deploy and manage containerised AI services using Docker and Kubernetes (AKS), including LLM inference endpoints, vector store services, and RAG pipeline backends.
- Maintain CI/CD pipelines (GitHub Actions or Azure DevOps) for AI and web application workloads — with staged rollouts, automated testing gates, and rollback capabilities.
- Implement and evolve MLOps tooling: model versioning, inference monitoring, and deployment pipelines that support safe, auditable releases into high-stakes social contexts.
Platform Resilience, Security, and Incident Response
- Design and implement backup and restoration strategies across heterogeneous data stores — PostgreSQL, MongoDB Atlas, CosmosDB, and vector databases — with tested recovery procedures.
- Architect and maintain disaster recovery (DR) plans covering multi-region failover, RTO/RPO targets, and regular DR drills for all production systems.
- Own incidents end-to-end, including writing clear, structured post-mortem communications for non-technical leadership under time pressure.
- Manage Kubernetes cluster resilience: node auto-scaling, pod disruption budgets, health checks, and stateful workload recovery.
Observability and Security
- Build and maintain the end-to-end observability stack — Azure Monitor, Grafana, Prometheus, and ELK/OpenSearch — with dashboards, alerting, and incident runbooks for all A4I applications.
- Bring a proactive security instinct: spotting hardcoded secrets, privileged containers, and overly permissive mounts on sight, not just applying IAM policy from a checklist.
- Manage Cloudflare for CDN, DNS routing, and edge security across all A4I-hosted applications.
Cost Governance and Platform Strategy
- Own cloud cost visibility and optimisation — right-sizing, reserved capacity planning, and rationalisation across multi-application Azure subscriptions.
- Evaluate and recommend cloud provider and tooling decisions, balancing cost, open-source compliance, and DPG deployment requirements.
- Automate operational workflows using Python and Bash, moving the team from reactive support to proactive, self-healing infrastructure.
What You Will Not Be Doing
- Managing tickets and waiting for engineers to raise infrastructure requests — this is a proactive ownership role.
- Working on a single-application environment — you will own a portfolio of distinct AI products simultaneously.
- Operating with a large dedicated infra team behind you — you are the infrastructure function, working closely with (but not handing off to) AI architects, engineers, and partner teams.
- Maintaining legacy on-premise infrastructure — the entire stack is cloud-native.
Must-Have Skills Each bullet represents a skill area we will probe directly during screening.
- Cloud platform ownership: Hands-on production experience with Azure, including App Services, Function Apps, AKS, Container Registry, CosmosDB, and Entra ID. Familiarity with native VMs is required; AWS and GCP experience is optional.
- Infrastructure as Code: Proficiency with Terraform and/or Ansible for managing multi-environment cloud infrastructure; you write IaC, you don't patch portals.
- Kubernetes and containers: Real production experience with Docker and Kubernetes (AKS or equivalent) — including stateful workload management, cluster scaling, and service mesh basics.
- AI/ML infrastructure — hard requirement: Hands-on production experience with at least one of: MLflow, Kubeflow, BentoML, NVIDIA Triton, Seldon, vLLM, or Ray Serve for model serving; or hands-on production deployment of vector databases (pgvector, Qdrant, Pinecone) and RAG pipeline infrastructure. General DevOps experience alone does not meet this bar.
- AI-native engineering practice: You use AI coding agents (Cursor, GitHub Copilot, Windsurf, or equivalent) as a core part of your daily workflow — including agent/composer modes and awareness of context-grounding approaches such as MCP. We're looking for genuine AI-first practice, not occasional autocomplete or copy-pasting into a chat window without reviewing the output.
- Backup, restoration, and DR: Demonstrated experience designing and testing backup strategies and disaster recovery plans across cloud-native databases and stateful services.
- CI/CD pipelines: Proven track record building and maintaining CI/CD pipelines for containerised AI/ML and web workloads using GitHub Actions, Azure DevOps, or equivalent.
- Observability: Hands-on experience with Azure Monitor, Grafana, Prometheus, and ELK/OpenSearch for production monitoring, alerting, and incident response.
- Security and compliance: Working knowledge of IAM, Entra ID, secrets management (Key Vault or equivalent), encryption practices, and PI data handling requirements — with a proactive instinct for spotting misconfigurations, not just policy knowledge.
- Python and Bash: Production-quality scripting for automation, infrastructure tooling, and operational workflows — not just glue scripts.
- Incident communication: Ability to write clear, structured incident and post-mortem communications under time pressure — evaluated directly in our process.
Qualifications
- Bachelor's or Master's in Computer Science, Engineering, or equivalent demonstrated expertise.
- 4 to 6 years in cloud infrastructure, platform engineering, or DevOps roles, with at least 2 years supporting live AI or ML workloads in production.
- No AI/ML infra experience yet? If you have strong general DevOps/SRE experience (Kubernetes, Terraform, multi-cloud) but haven't yet worked hands-on with MLOps tooling or model serving, we'd still like to hear from you for a separate, more junior AI-infra track — but please don't apply to this senior req expecting to learn the AI layer on the job, as this specific role is scoped to already have it.
- Demonstrated experience managing infrastructure for more than one live application simultaneously.
- Verifiable track record of cloud cost governance and right-sizing in real production environments — references or case studies will be requested.
Preferred Qualifications
- Hands-on experience with Azure-specific AI services: Azure OpenAI, Cognitive Services, Bhashini integration, or Sarvam AI endpoints in production.
- Familiarity with vector databases (pgvector, Qdrant, or equivalent) and the infrastructure considerations for embedding pipelines at scale.
- Prior work in open-source DPG deployments or multi-tenant cloud environments serving low-resource or low-connectivity end users.
- Experience with Cloudflare Workers, DNS management, and edge security configuration.
- Contributions to open-source infrastructure tooling, IaC libraries, or internal platform frameworks.
- Exposure to multilingual AI service infrastructure — Indic language models, ASR/TTS endpoints, or translation API pipelines.
A4I Lab is committed to building AI systems that serve everyone. We actively encourage applications from candidates of all backgrounds, disciplines, and geographies. Pay: ₹1,500,000.00 - ₹2,800,000.00 per year
Benefits
- Health insurance
- Provident Fund
Application Question(s):
- •Have you deployed or operated any of the following in production: MLflow, Kubeflow, BentoML, NVIDIA Triton, Seldon, vLLM, or Ray Serve?
◦Yes ◦No
- •Do you use AI coding agents (Cursor, Copilot, Windsurf) as a core part of your daily workflow, including agent modes or context-grounding techniques like MCP?
◦Yes ◦No
- •Have you deployed or operated a vector database (pgvector, Qdrant, Pinecone, or equivalent) in production?
◦Yes ◦No
- •Are you comfortable being the sole owner of infrastructure — including writing incident communications for leadership — rather than working within a larger dedicated infra team?
◦Yes ◦No
- •Have you managed cloud infrastructure for more than one live application simultaneously?•Have you managed cloud infrastructure for more than one live application simultaneously?
YES NO
- •Have you used Azure hands-on in a production environment?
◦Yes ◦No
- •Have you used AWS or GCP hands-on in a production environment?
◦Yes ◦No
- •Have you designed a disaster recovery plan for a live cloud system?
◦Yes ◦No
- •Have you used Kubernetes to manage your production environment?
◦Yes ◦No
Work Location: In person
📌 AI INFRASTRUCTURE ENGINEER (Electronic City)
🏢 International Institute of Information Technology Banglore (A4I)
📍 Electronic City