22 Aug
|
Concept Dash
|
India
22 Aug
Concept Dash
India
AI Platform Engineer – Agents, Retrieval & AI Infrastructure
Position
Designation:
AI Platform Engineer
Department:
Technology / AI Engineering
Experience:
5+ years in software engineering, including 2+ years in production LLM/AI engineering
Employment Type:
Full-Time
Work Mode:
Remote
Role Overview
We are looking for an experienced AI Platform Engineer to build and operate production-grade AI systems across LLM agents, retrieval/RAG, data pipelines, backend services, cloud infrastructure, and AI model serving
. The role requires strong depth in agent and retrieval engineering, with practical expertise in data engineering, databases, cloud, and GPU infrastructure.
Key Responsibilities
AI Agents & LLM Engineering
Design and deploy production LLM-powered agents with tool calling, structured outputs, JSON-schema validation, streaming, prompt caching, and human-in-the-loop workflows.
Work with agent frameworks such as LangGraph, LangChain, CrewAI, AutoGen/AG2, Pydantic AI, Google ADK, or Microsoft/OpenAI Agent SDKs and select the appropriate architecture for each use case.
Build reliable agent workflows with checkpointing, pause/resume, memory, context management, guardrails, and controlled tool execution.
Retrieval, RAG & Knowledge Systems
Build enterprise-grade RAG and retrieval systems using vector search, pgvector, metadata filtering, contextual chunking, embeddings, cross-encoder reranking, and graph-based retrieval where appropriate.
Develop document pipelines covering PDF/DOCX extraction, chunking, deduplication, change detection, incremental indexing, and permission-aware retrieval
.
Work with LlamaIndex, Haystack, Qdrant, Weaviate, Milvus
, or equivalent technologies and measure retrieval using Recall@K, nDCG, and labelled evaluation datasets.
AI Evaluation & Observability
Create automated AI evaluation frameworks using real-world test cases, expected outputs, failure categories, and regression testing.
Implement production observability for agent traces, tool calls, token usage, latency, model performance, errors, and cost using tools such as LangSmith, Langfuse, or OpenTelemetry.
Support online evaluation and LLM-as-a-judge approaches using tools such as RAGAS, DeepEval, promptfoo, or Braintrust
.
Backend, Database & Data Engineering
Develop production services using Python and/or TypeScript/Node.js
, with working knowledge of REST APIs and Next.js.
Solid PostgreSQL experience is required, including EXPLAIN ANALYZE, indexing, advanced SQL, window functions, materialized views,
and production migrations
. Working knowledge of MongoDB and complex nested business documents is expected.
Build reliable data pipelines using Airflow, Dagster, Prefect, dbt, Great Expectations
, or equivalent technologies, including monitoring, data-quality checks, retries, and incremental processing.
MCP & AI Interoperability
Build or integrate Model Context Protocol (MCP)
servers and clients and understand tools, resources, schemas, authentication, elicitation, and sampling.
Work with emerging agent-to-agent (A2A)
interoperability standards and design secure model-facing tools and interfaces.
Cloud, DevOps & Infrastructure
Deploy and operate AI workloads using AWS, Docker, Kubernetes, Terraform/CDK, IAM, secrets management, and CI/CD
.
Own production infrastructure including monitoring, logging, alerting, backups, recovery, database migrations, rollback procedures, and incident resolution.
Build durable queues and event-driven workflows using SQS, Kafka, NATS, Redis Streams
, or equivalent technologies for long-running AI and data workloads.
AI Models & GPU Infrastructure
Deploy and operate open-weight models such as Llama, Qwen, or Mistral using vLLM, SGLang
, or equivalent inference platforms.
Manage model quantization, LoRA/QLoRA fine-tuning, embeddings, rerankers, GPU instances, CUDA, batching, KV-cache, autoscaling, model storage, and inference cost optimization.
Experience with NVIDIA GPU Operator, Kubernetes GPU scheduling, MIG/time-slicing, DCGM, Triton, NIM, or TensorRT-LLM is highly desirable.
Security & AI Governance
Implement permission-aware retrieval and enforce user authorization at the application and data layers rather than through prompts.
Ensure AI systems cannot independently perform sensitive actions such as sending emails, deleting records, modifying financial data, or writing to enterprise systems without appropriate controls and human approval.
Apply secure secrets management, IAM, access controls, guardrails, and safe handling of untrusted web or document content.
Required Technical Skills
Programming:
Python, TypeScript/Node.js, REST APIs, Next.js, Pydantic/Zod or equivalent.
AI/LLM:
LLM APIs,
prompt engineering, tool calling, structured outputs, streaming, prompt caching, agent orchestration, human-in-the-loop, agent memory, guardrails.
Agent Frameworks:
LangGraph, LangChain, CrewAI, AutoGen/AG2, Pydantic AI, Google ADK, Microsoft/OpenAI Agent SDKs or equivalent.
RAG & Retrieval:
Vector search, pgvector, embeddings, reranking, chunking, metadata filtering, LlamaIndex/Haystack, retrieval evaluation, document indexing.
Databases:
Advanced PostgreSQL, query optimization, indexing, migrations, window functions, materialized views; working knowledge of MongoDB.
Data Engineering:
Airflow/Dagster/Prefect, dbt/Great Expectations, data quality, incremental pipelines, CDC and reliable job orchestration.
Cloud & DevOps:
AWS, Docker, Kubernetes, Terraform/CDK, IAM, secrets management, CI/CD, monitoring and alerting.
AI Infrastructure:
vLLM/SGLang, open-weight models, GPU deployment, CUDA, quantization, LoRA/QLoRA, NVIDIA GPU Operator, GPU scheduling and optimization.
AI Integration:
MCP, A2A, Microsoft Graph, SharePoint, browser automation, secure external-data ingestion.
Required Experience & Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Artificial Intelligence, Software Engineering, or a related field .
- 5+ years of professional software engineering experience.
- 2+ years of hands-on experience building and deploying production LLM/AI applications.
- Proven experience taking AI solutions from development through production and supporting real users.
- Strong experience in at least two of the three areas: AI Agents, Retrieval/Data Engineering, and AI Infrastructure .
- Experience troubleshooting production AI systems using logs, traces, metrics, and evaluation results.
- Strong understanding of system reliability, scalability, security, performance, and cost optimization.
Nice to Have
- AEC, architecture, engineering, construction, or infrastructure domain knowledge
- Knowledge of RFPs, tenders, proposals, procurement, or fee/manpower modelling
- Microsoft Graph and SharePoint APIs
- Debezium or PostgreSQL logical replication
- LiteLLM, OpenRouter, or AI gateway platforms
- Browserbase, Firecrawl, Playwright, E2B, or Daytona
- Vercel AI SDK, CopilotKit, or AG-UI
- RAGAS, DeepEval, promptfoo, or Braintrust
- LLM-as-a-judge and human-labelled evaluation methodologies
- A2A and MCP-based integrations
- Cloud cost optimization and AI infrastructure cost management
📌 AI Platform Engineer – Agents, Retrieval & AI Infrastructure (India)
🏢 Concept Dash
📍 India