Agentic AI Engineer Technical Lead (India)

Agentic AI Engineer Technical Lead (India)

07 Aug
|
Headway It Solutions
|
India

07 Aug

Headway It Solutions

India

Job Descriptions Agentic AI Platform

Role 1: Agentic AI Engineer — Technical Lead

Location: [City / Hybrid] | Function: Technology & Data | Level: Senior / Lead | Type: Permanent

About the Role

We are building a next-generation Agentic AI development platform that enables autonomous, multi-step AI workflows across the enterprise. As the Agentic AI Engineer Technical Lead, you will be the primary architect and hands-on engineer responsible for designing, building, and evolving the core agentic layer of this platform.

You will lead the end-to-end technical delivery of agentic systems — from LLM reasoning and tool orchestration to memory management, safety gating, and real-time event-driven execution. You will work at the intersection of LLM engineering, distributed systems, and enterprise AI, translating complex business problems into production-grade agentic solutions.

Key Responsibilities

- Agentic Architecture & Design: Lead the design and implementation of multi-step, multi-agent workflows using LangGraph, including branching logic, conditional routing, loops, and state management across agent nodes.
- LLM Engineering: Configure, version, and optimise prompts and model parameters via AI Foundry (GPT-4/5x); implement chain-of-thought, ReAct, and tool-use reasoning patterns.
- Tool & API Integration: Build and maintain tool-call layers — RAG retrieval via Azure AI Search, API integrations via Azure Functions (real and mock), ML model calls via Databricks Model Serving, and file I/O via Blob Storage.
- Event-Driven Orchestration: Design and implement Kafka-based event triggers for agent activation; define message schemas, queue management, and dispatcher logic via Azure Functions.
- Memory & State Management: Implement short-term (LangGraph in-context), checkpoint (Postgres / Cosmos DB), and long-term (AI Search / Cosmos DB) memory strategies for resumable, cross-session agent workflows.
- Safety & Output Governance: Integrate Azure AI Content Safety for PII detection, toxicity filtering, and groundedness checks; design LLM output review stores for governance and audit.
- Technical Leadership: Define engineering standards, conduct code reviews, mentor engineers, and drive architectural decisions across the agentic platform team.
- Collaboration: Partner closely with data scientists, MLOps engineers, product owners, and business stakeholders to translate use cases into agentic solutions.
- Platform Evolution: Contribute to the platform roadmap — identifying capability gaps, evaluating new frameworks, and driving continuous improvement of the agentic stack.

Required Skills & Experience

- LLM & Agent Frameworks: 3+ years hands-on experience with LLM-based systems; deep expertise in LangGraph and/or LangChain for agentic workflow orchestration.
- Python Engineering: Strong Python development skills; experience building production-grade, modular, testable AI systems.
- Azure AI Stack: Practical experience with Azure AI Foundry, Azure AI Search (vector indexing, RAG), Azure Functions, and Azure Blob Storage.
- Data & ML Platforms: Experience with Databricks (Unity Catalog, SQL endpoints, Model Serving REST APIs) for data querying and ML model integration.
- Event-Driven Systems: Familiarity with Kafka or equivalent message streaming platforms for event-driven agent triggers.
- State & Memory Stores: Experience with Postgres, Cosmos DB, or equivalent for workflow checkpointing and persistent memory.
- Prompt Engineering: Strong understanding of prompt design patterns — chain-of-thought, ReAct, few-shot, tool-use — and prompt lifecycle management.
- Safety & Responsible AI: Awareness of LLM safety challenges (hallucination, PII leakage, prompt injection); experience with content safety tooling.
- CI/CD & DevOps: Proficiency with GitHub-based CI/CD pipelines; experience deploying AI/ML workloads in cloud environments.
- Security:



Understanding of secrets management via Azure Key Vault and Managed Identity authentication patterns.

Preferred / Nice-to-Have

- Experience with multi-agent architectures (supervisor agents, specialist sub-agents, agent-to-agent communication).
- Familiarity with MLflow for experiment tracking and model evaluation.
- Experience with Application Insights / Azure Monitor for AI system observability.
- Background in insurance, financial services, or other regulated industries.
- Contributions to open-source LLM or agent framework projects.
- Experience with vector databases beyond Azure AI Search (e.g., Pinecone, Weaviate, pgvector).

What You Will Be Working On

From day one, you will be working on real agentic use cases being piloted across the business. The platform stack is live — LangGraph orchestration, GPT-4/5x via AI Foundry, RAG via Azure AI Search, Databricks ML model calls, Azure Content Safety, and full observability via Application Insights are all operational. You will be building and refining agents that reason, retrieve, act, and respond — with Kafka event triggers and live system API integrations as the next frontier.

What We Offer

- A greenfield agentic AI platform with a contemporary, fully cloud-native stack.
- Direct impact on enterprise AI strategy and early-stage product development.
- Collaborative, cross-functional team with data scientists, MLOps engineers, and domain experts.
- Continuous learning culture — access to AI research, tooling, and experimentation time.
- Competitive compensation, adaptable working, and a clear career pathway in AI engineering leadership.

Role 2: Agentic AI & MLOps Engineer — Lead

Location: [City / Hybrid] | Function: Technology & Data | Level: Senior / Lead | Type: Permanent

About the Role

As the Agentic AI & MLOps Engineer Lead, you will own the operational backbone of the Agentic AI platform — ensuring that AI agents, LLM pipelines, and ML models are reliably deployed, continuously evaluated, and observable in production. You will bridge the gap between agentic AI engineering and machine learning operations, bringing engineering rigour to a fast-moving AI development environment.

This is a dual-domain role: you will contribute to agentic workflow development while simultaneously owning the MLOps infrastructure — model serving, evaluation pipelines, experiment tracking, CI/CD automation, observability, and governance. You will be a key enabler of the platform's path to production.

Key Responsibilities

- MLOps Platform Ownership: Design, build, and maintain the MLOps infrastructure underpinning the Agentic AI platform — model serving, experiment tracking, evaluation pipelines, and deployment automation using Databricks MLflow.
- Model Serving & Lifecycle Management: Manage ML model endpoints on Databricks Model Serving (REST); oversee model versioning, promotion, and deprecation workflows across dev, staging, and production environments.
- LLM Evaluation Pipelines: Build and operate automated evaluation pipelines for LLM outputs — accuracy, relevance, groundedness, and safety metrics — integrated with MLflow experiment tracking.
- CI/CD for AI Workloads: Own and evolve the GitHub-based CI/CD pipeline for agent and model code — automated testing, quality gates, environment promotion, and deployment to production.
- Observability & Monitoring: Implement and maintain end-to-end observability for agentic workflows using Application Insights and Azure Monitor — execution traces, token usage, latency, error rates, and alerting.
- Data Pipeline Automation:



Automate data ingestion pipelines into Databricks Unity Catalog and Azure AI Search (RAG document ingestion); ensure data quality, freshness, and governed access.
- Secrets & Identity Management: Maintain secure credential management via Azure Key Vault and Managed Identity; enforce zero-secrets-in-code policies across all platform components.
- Agentic Engineering Contribution: Contribute to agentic workflow development in LangGraph — particularly around tool-call reliability, error handling, retry logic, and workflow resilience.
- Governance & Audit: Maintain the LLM output review store (Cosmos DB / Blob Storage) for governance sign-off; support audit and compliance requirements for AI outputs.
- Platform Roadmap: Collaborate with the Technical Lead and product stakeholders to prioritise and deliver the platform's path-to-production roadmap.

Required Skills & Experience

- MLOps & Model Lifecycle: 3+ years of MLOps experience; deep expertise in Databricks MLflow for experiment tracking, model registry, and evaluation.
- Databricks Platform: Strong hands-on experience with Databricks — Unity Catalog, SQL endpoints, Model Serving (REST), and Delta Lake.
- CI/CD for AI/ML: Proven experience building and maintaining GitHub Actions or equivalent CI/CD pipelines for ML and AI workloads — including automated testing, staging, and production promotion.
- Cloud Observability: Proficiency with Application Insights and Azure Monitor for distributed tracing, metrics, alerting, and dashboarding of AI systems.
- Python Engineering: Strong Python skills; experience writing production-quality ML pipeline code, evaluation scripts, and automation tooling.
- Azure Cloud Services: Hands-on experience with Azure Functions, Blob Storage, Cosmos DB, Key Vault, and Managed Identity.
- LLM Evaluation: Understanding of LLM evaluation methodologies — groundedness, relevance, faithfulness, toxicity — and experience implementing automated eval pipelines.
- Data Engineering: Experience building and maintaining data ingestion pipelines; familiarity with event-driven data flows and batch processing.
- Security & Governance: Strong understanding of secrets management, RBAC, and data access governance in cloud AI environments.
- Agentic AI Familiarity: Working knowledge of LangGraph or LangChain for agentic workflow orchestration; ability to contribute to agent development alongside the engineering team.

Preferred / Nice-to-Have

- Experience with Kafka or event streaming platforms for pipeline triggering and data flow.
- Familiarity with Azure AI Content Safety and responsible AI tooling.
- Experience with vector search and RAG pipeline management in Azure AI Search.
- Background in regulated industries (insurance, financial services, healthcare).
- Knowledge of Kubernetes or containerised ML workload deployment.
- Experience with data quality frameworks (Great Expectations, dbt, or equivalent).
- Familiarity with cost management and FinOps practices for cloud AI workloads.

What You Will Be Working On

The platform's observability, model serving, evaluation, CI/CD, and secrets management layers are all live and operational. Your immediate focus will be on hardening these capabilities for production scale — automating evaluation pipelines, building robust CI/CD gates for agent code, wiring live data ingestion pipelines, and preparing the platform for Kafka-based event-driven operation. You will be the engineering force that takes the platform from 'demo-ready' to 'production-grade'.

What We Offer

- End-to-end ownership of the MLOps and operational backbone of a cutting-edge Agentic AI platform.
- A modern, fully cloud-native stack — Databricks, Azure, LangGraph, MLflow, GitHub Actions.
- Direct collaboration with AI engineers, data scientists, and business stakeholders.
- High-impact role on the critical path to production for enterprise AI use cases.
- Competitive compensation, flexible working, and a clear career pathway in AI/MLOps leadership.

📌 Agentic AI Engineer Technical Lead (India)
🏢 Headway It Solutions
📍 India

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: agentic ai engineer technical lead (india) / india

Subscribe to this job alert:

Get the latest job offers by email for: agentic ai engineer technical lead (india) / india