02 Oct
|
International Institute of Information Technology Banglore (A4I)
|
Electronic City
02 Oct
International Institute of Information Technology Banglore (A4I)
Electronic City
AI INFRASTRUCTURE ENGINEER
CLOUD INFRASTRUCTURE & AGENTIC SYSTEMS · A4I Lab (IIIT-B × Microsoft)
Organization: A4I Lab (IIIT-B × Microsoft)
Location: IIIT Bangalore
Type: Contract (through March 2027, extendable)
Experience: 5 to 8 years
About A4I Lab
AI Innovation and Inclusion Initiative (A4I) is a partnership between Microsoft and IIIT-B, collaborating with non-profit partners to harness AI for solving real-world challenges in education, healthcare, accessibility, and agriculture. A4I evolves innovations into open-source Digital Public Goods (DPGs) that are deployable at scale, fostering a strong AI innovation ecosystem for social impact.
The Opportunity
Imagine infrastructure that doesn't just run AI — it enables it. At A4I Lab, we're not spinning up VMs and hoping for the best. We're engineering the cloud backbone that powers autonomous multi-agent systems serving millions of users in India's most underserved communities: students learning in vernacular languages, frontline health workers operating without connectivity, farmers acting on real-time data.
Generic DevOps practices don't work here. The constraints are too real — cost ceilings, data sovereignty, multilingual AI services, and uptime requirements that leave no room for reactive firefighting. What we need is an infrastructure engineer who has already run AI-specific workloads in production, not just Kubernetes and Terraform in general.
Two things worth knowing before you apply.
First: general DevOps/SRE experience (Kubernetes, Terraform, CI/CD, observability) is necessary but not sufficient on its own — we specifically need hands-on production experience with the AI/ML layer on top of that (model serving, vector stores, or MLOps tooling).
Second: you will effectively be a one-person infrastructure function. That means full platform ownership, mentoring other engineers on infra practices, and writing clear incident communications for leadership under pressure — not just running commands.
What You Will Be Doing
Building and Owning Cloud Infrastructure
- Design, provision, and manage multi-application cloud environments on Azure (primary), with working knowledge of AWS and GCP — across compute, networking, storage, and managed services.
- Own environment parity between staging and production for a portfolio of live agentic AI applications, ensuring consistent, reproducible deployments.
- Drive Infrastructure as Code (IaC) practices using Terraform and Ansible, eliminating manual provisioning and enabling GitOps workflows.
Enabling Reliable AI Workloads in Production
- Deploy and manage containerised AI services using Docker and Kubernetes (AKS), including LLM inference endpoints, vector store services, and RAG pipeline backends.
- Maintain CI/CD pipelines (GitHub Actions or Azure DevOps) for AI and web application workloads — with staged rollouts, automated testing gates, and rollback capabilities.
- Implement and evolve MLOps tooling: model versioning, inference monitoring, and deployment pipelines that support safe, auditable releases into high-stakes social contexts.
Platform Resilience, Security, and Incident Response
- Design and implement backup and restoration strategies across heterogeneous data stores — PostgreSQL, MongoDB Atlas, CosmosDB, and vector databases — with tested recovery procedures.
- Architect and maintain disaster recovery (DR) plans covering multi-region failover, RTO/RPO targets, and regular DR drills for all production systems.
- Own incidents end-to-end, including writing clear, structured post-mortem communications for non-technical leadership under time pressure.
- Manage Kubernetes cluster resilience: node auto-scaling, pod disruption budgets, health checks, and stateful workload recovery.
Observability and Security
- Build and maintain the end-to-end observability stack — Azure Monitor, Grafana, Prometheus, and ELK/OpenSearch — with dashboards, alerting, and incident runbooks for all A4I applications.
- Bring a proactive security instinct: spotting hardcoded secrets, privileged containers, and overly permissive mounts on sight, not just applying IAM policy from a checklist.
- Manage Cloudflare for CDN, DNS routing, and edge security across all A4I-hosted applications.
Cost Governance and Platform Strategy
- Own cloud cost visibility and optimisation — right-sizing, reserved capacity planning, and rationalisation across multi-application Azure subscriptions.
- Evaluate and recommend cloud provider and tooling decisions, balancing cost, open-source compliance, and DPG deployment requirements.
- Automate operational workflows using Python and Bash, moving the team from reactive support to proactive, self-healing infrastructure.
What You Will Not Be Doing
- Managing tickets and waiting for engineers to raise infrastructure requests — this is a proactive ownership role.
- Working on a single-application environment — you will own a portfolio of distinct AI products simultaneously.
- Operating with a large dedicated infra team behind you — you are the infrastructure function, working closely with (but not handing off to) AI architects, engineers, and partner teams.
- Maintaining legacy on-premise infrastructure — the entire stack is cloud-native.
Must-Have Skills
Each bullet represents a skill area we will probe directly during screening.
- Cloud platform ownership: Hands-on production experience with Azure, including App Services, Function Apps, AKS, Container Registry, CosmosDB, and Entra ID. Familiarity with native VMs is required; AWS and GCP experience is optional.
- Infrastructure as Code: Proficiency with Terraform and/or Ansible for managing multi-environment cloud infrastructure; you write IaC, you don't patch portals.
- Kubernetes and containers: Real production experience with Docker and Kubernetes (AKS or equivalent) — including stateful workload management, cluster scaling, and service mesh basics.
- AI/ML infrastructure — hard requirement: Hands-on production experience with at least one of: MLflow, Kubeflow, BentoML, NVIDIA Triton, Seldon, vLLM, or Ray Serve for model serving; or hands-on production deployment of vector databases (pgvector, Qdrant, Pinecone) and RAG pipeline infrastructure. General DevOps experience alone does not meet this bar.
- AI-native engineering practice: You use AI coding agents (Cursor, GitHub Copilot, Windsurf, or equivalent) as a core part of your daily workflow — including agent/composer modes and awareness of context-grounding approaches such as MCP. We're looking for genuine AI-first practice, not occasional autocomplete or copy-pasting into a chat window without reviewing the output.
- Backup, restoration, and DR: Demonstrated experience designing and testing backup strategies and disaster recovery plans across cloud-native databases and stateful services.
- CI/CD pipelines: Proven track record building and maintaining CI/CD pipelines for containerised AI/ML and web workloads using GitHub Actions, Azure DevOps, or equivalent.
- Observability: Hands-on experience with Azure Monitor, Grafana, Prometheus, and ELK/OpenSearch for production monitoring, alerting, and incident response.
- Security and compliance: Working knowledge of IAM, Entra ID, secrets management (Key Vault or equivalent), encryption practices, and PI data handling requirements — with a proactive instinct for spotting misconfigurations, not just policy knowledge.
- Python and Bash: Production-quality scripting for automation, infrastructure tooling, and operational workflows — not just glue scripts.
- Incident communication: Ability to write explicit, structured incident and post-mortem communications under time pressure — evaluated directly in our process.
Qualifications
- Bachelor's or Master's in Computer Science, Engineering, or equivalent demonstrated expertise.
- 4 to 6 years in cloud infrastructure, platform engineering, or DevOps roles, with at least 2 years supporting live AI or ML workloads in production.
- No AI/ML infra experience yet? If you have strong general DevOps/SRE experience (Kubernetes, Terraform, multi-cloud) but haven't yet worked hands-on with MLOps tooling or model serving, we'd still like to hear from you for a separate, more junior AI-infra track — but please don't apply to this senior req expecting to learn the AI layer on the job, as this specific role is scoped to already have it.
- Demonstrated experience managing infrastructure for more than one live application simultaneously.
- Verifiable track record of cloud cost governance and right-sizing in real production environments — references or case studies will be requested.
Preferred Qualifications
- Hands-on experience with Azure-specific AI services: Azure OpenAI, Cognitive Services, Bhashini integration, or Sarvam AI endpoints in production.
- Familiarity with vector databases (pgvector, Qdrant, or equivalent) and the infrastructure considerations for embedding pipelines at scale.
- Prior work in open-source DPG deployments or multi-tenant cloud environments serving low-resource or low-connectivity end users.
- Experience with Cloudflare Workers, DNS management, and edge security configuration.
- Contributions to open-source infrastructure tooling, IaC libraries, or internal platform frameworks.
- Exposure to multilingual AI service infrastructure — Indic language models, ASR/TTS endpoints, or translation API pipelines.
Pay: ₹1,800,000.00 - ₹2,400,000.00 per year
Application Question(s)
- Three of our 20 Kubernetes nodes keep evicting pods, and node memory looks fine on all of them. Someone has a PR ready to raise the memory limits. What do you check first and in what order? What would make you approve or reject that PR?
- After we moved a service to a new namespace, about 1 in 6 requests from a client pod takes roughly 5 seconds longer. The rest are normal. The team suspects the network policies added during the move. How would you test that, and what would you look at in the delay itself?
- A Terraform state lock has been held for 3 hours. The owner is out sick, so a teammate force-unlocks and applies a small change. The next plan wants to create 14 resources that already exist. Write your first 10 minutes of actions and what each one tells you. Which actions would you not allow anyone to take?
- A CI job that has been green for months now fails about 1 in 8 runs on the same commit. Retrying usually passes, and the team wants to add automatic retries. What would you compare across passing and failing runs? When is a retry acceptable, and what would you add so it doesn't hide a real problem?
- A pull request from a fork triggers a CI workflow that has access to deploy credentials and runs code from the PR. A reviewer notices after the run finished. The logs look clean. Was anything exposed? What do you do today versus this week? What would you tell a stakeholder who asks "are we safe?"
Work Location: In person
📌 AI INFRASTRUCTURE ENGINEER (Electronic City)
🏢 International Institute of Information Technology Banglore (A4I)
📍 Electronic City