AI Engineer (Mumbai)

AI Engineer (Mumbai)

13 Aug
|
Unico Connect Private
|
Mumbai

13 Aug

Unico Connect Private

Mumbai

AI Engineer

LLMs, Agents & AI Services

Mumbai (On-site) | Full-time | 2-4 years

About the Role:

Unico Connect is an AI-first technology partner that builds custom mobile, web, and AI products for clients across multiple geographies.

AI is core to how we design, deliver, and scale software for our customers.

We are hiring an AI Engineer for a dedicated client engagement building a complex production AI platform, working on the AI capabilities and agentic features at the core of the product.

The mandatory requirement for this role is at least one AI feature personally shipped to production for real users, with operational ownership.

The role suits someone who thinks quickly on solutioning, can take an ambiguous problem to a working prototype in days, and has the discipline to carry it through to production with predictable economics.

You will work alongside the Senior AI Engineer and the wider pod, with ownership of parts of the AI surface area of the product.

Responsibilities:

Solutioning and POCs

Translate ambiguous customer problems into working POCs at speed.

Pick the right model, framework, and architecture, and demonstrate value early before scaling investment.

LLM Application Development

Build AI features and services using LLM APIs from OpenAI, Anthropic, Google, and self-hosted open-weight models (Llama, Qwen, Mistral).

Choose the right model per use case based on cost, latency, capability, and context-window trade-offs.

Agentic System Design

Design and implement agentic workflows using LangGraph, CrewAI, AutoGen, LlamaIndex Agents, or custom orchestration.

Cover tool use, planning, memory, and multi-step reasoning appropriate to the problem.

API and Service Development

Build production AI services and APIs using Python and FastAPI.

Handle streaming responses, async processing, structured outputs, retries,



and graceful degradation when models or tools fail.

Retrieval and Tool Integration

Implement RAG pipelines with vector databases (Pinecone, Weaviate, Qdrant, pgvector, Chroma), embeddings, chunking strategies, hybrid search, and reranking.

Integrate external tools, internal APIs, and document sources through tool-calling and MCP-style patterns.

Cost Analysis and Unit Economics

Model the per-request and per-user cost of every AI feature before it ships.

Track token usage, prompt caching, batching, and model-routing strategies.

Drive measurable improvements in unit economics.

Production Hardening

Add observability and tracing (LangSmith, Langfuse, OpenTelemetry), guardrails, content safety checks, prompt injection defences, and fallback behaviour.

Prompt Engineering and Evaluation

Design, test, and iterate prompts with measured outcomes.

Build evaluation harnesses for accuracy, hallucination, latency, and cost.

Run benchmarks across models and prompt variants before locking in a design.

Requirements:

AI Feature Shipped to Production (Mandatory)

Must have personally built and shipped at least one AI feature that runs in production for real users, with operational ownership.

POCs, internal demos, and one-off scripts do not qualify.

2 to 4 Years of Professional Software or AI Engineering Experience

With at least one production AI feature owned end to end.

Strong Python Proficiency and API Development with FastAPI

Comfort with type hints, async, packaging, testing, streaming responses,



and authentication.

Production-grade Python, not notebook-only code.

Hands-on Depth Across the LLM and Agent Stack

Working experience with at least two of OpenAI, Anthropic Claude, Google Gemini, or self-hosted open-weight models (vLLM, Ollama, Together, Replicate).

Working familiarity with at least one agent framework (LangGraph, CrewAI, AutoGen, LlamaIndex Agents) or hand-rolled equivalent.

Working knowledge of RAG, embeddings, and vector databases (Pinecone, Weaviate, Qdrant, pgvector, Chroma).

Solutioning Speed and POC Velocity

Demonstrated ability to move from a fuzzy problem to a working prototype in days.

Strong instinct for what to build first, what to defer, and what to throw away.

Cost Discipline for Production AI

Ability to calculate, monitor, and optimise the cost of LLM APIs, tokens, embeddings, vector store usage, and infrastructure.

Treats unit economics as a first-class concern.

AWS Familiarity

Working knowledge of EC2, S3, IAM, and at least one of Bedrock, SageMaker, or equivalent.

Comfortable in a Quick-Moving Environment

Self-directed, comfortable with ambiguity, takes ownership without being asked, and ships under shifting priorities.

Strong Written and Spoken English Communication

Able to explain trade-offs to non-AI engineers, designers, product managers, and clients in plain language.

Nice to Have

- fine-tuning or LoRA, QLoRA, PEFT exposure
- MCP server authoring
- eval framework experience (LangSmith, Promptfoo, Ragas, DeepEval)
- open-source AI contributions
- multi-modal models (vision, audio)

Skills:- Python, Large Language Models (LLM), Generative AI, LangGraph, FastAPI, Retrieval Augmented Generation (RAG), AI Agents, OpenAI, Anthropic Claude, Google Gemini, Vector database and Prompt engineering

📌 AI Engineer (Mumbai)
🏢 Unico Connect Private
📍 Mumbai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: ai engineer (mumbai) / mumbai

Subscribe to this job alert:

Get the latest job offers by email for: ai engineer (mumbai) / mumbai