AI Engineer (LLM & Agent Systems) (Mumbai)

AI Engineer (LLM & Agent Systems) (Mumbai)

20 Sep
|
SPEEGILE CONSULTING
|
Mumbai

20 Sep

SPEEGILE CONSULTING

Mumbai

About the Role

We are looking for an experienced AI Engineer to design, build, and deploy LLM-driven agentic systems end-to-end — from data and model fine-tuning to retrieval, evaluation, guardrails, and production deployment. You should be equally comfortable reasoning about how a transformer works internally and shipping a reliable, observable agent system in production. This role owns the AI/ML depth of our stack, with support from backend and frontend engineers for integration.

Experience: 3+ Years

Qualification: B.E. / B.Tech / M.Sc. / M.Tech / MCA — Computer Science / IT or related field

Location: Preference Mumbai / Relocation to Mumbai / Remote

Employment Type: Full-Time

Notice Period: Max 30 Days

Must-Have Skills

• 3+ years of experience building Python-based ML/AI systems in production

• Strong, practical understanding of how transformer, encoder-decoder, and deep learning models work internally — not just API-level usage

• Hands-on experience fine-tuning LLMs (LoRA / QLoRA / full fine-tuning) and working with large-scale datasets end-to-end

• Foundational Knowledge of LLMs and Can Lay the foundation for building a LLM model from scratch.

• Solid grasp of the full LLM landscape — RAG, vectored and vector-less retrieval, evaluation, observability, and guardrails

• Experience with voice/speech models (ASR/TTS & Translation) in Fine Tuning the models can even lay the foundation to build a voice model from scratch

• Knowledge graphs, Redis, Advanced RAG , Self corrective RAG

• Experience with cloud AI deployments (AWS / GCP / Azure)

• Hands-on experience with LLM agent frameworks and vector databases (ChromaDB, Weaviate, pgvector)

• Strong knowledge of PyTorch or TensorFlow and Scikit-Learn

• Production experience with FastAPI, Docker, and MLOps

Good-to-Have Skills

• Advanced prompt optimization and agent evaluation techniques

• Experience building AI observability tools from scratch





• Exposure to multilingual or low-resource language models

• Expert-level usage of agentic coding IDEs (Cursor, Windsurf, Claude Code)

Soft Skills

• Solid problem-solving and system design mindset

• Ability to work across AI, backend, and front-end teams

• Comfortable mentoring freshers/junior engineers on AI fundamentals

• Clear communication and documentation skills

• Passion for building production-grade AI systems

What You'll Work On

LLM Agents & Prompt Engineering

• Design and implement LLM agents using LangGraph, PydanticAI, and Google ADK

• Build tool-augmented reasoning pipelines: RAG, Chain-of-Thought, ReAct, and planner–executor architectures

• Develop robust, tested prompt strategies to improve reliability and reduce hallucination

Model Fundamentals & Fine-Tuning

• Deep working knowledge of transformer architecture, Mamba architecture— encoder-only, decoder-only, and encoder–decoder models — and how attention, embeddings, and positional encoding actually work under the hood

• Fine-tune LLMs using LoRA, QLoRA, and full fine-tuning depending on the use case and compute budget. If the fine tuning doesn't get the desired results, a LLM model will be built from scratch

• Prepare, clean, and process large-scale datasets for pre-training/fine-tuning — deduplication, tokenization, sampling, and quality filtering at scale

• Understanding of core deep learning model families (CNNs, RNNs/LSTMs, transformers) and when to use each





• Work with voice/speech models — ASR (speech-to-text) and TTS (text-to-speech) — and understand how they integrate into conversational AI pipelines

Retrieval, RAG & Knowledge Systems

• Implement both vectored (embedding-based) and vector-less (keyword/graph/hybrid) retrieval strategies, choosing the right approach per use case

• Integrate vector databases (FAISS, Pinecone, pgvector, ChromaDB, Weaviate) and knowledge graphs (Neo4j)

• Design chunking, embedding, and re-ranking strategies for high-precision retrieval

• Implement Prompt Engineering, Context Engineering, Loop Engineering & Efficient Low token retrieval

Evaluation, Observability & Guardrails

• Build offline and online evaluation harnesses for agent and LLM outputs (accuracy, groundedness, latency, cost)

• Implement guardrails for safety, PII redaction, and scope control (e.g. NeMo Guardrails, Guardrails AI, or custom rule/LLM-based filters)

• Build tooling for trace analysis, state debugging, and hallucination detection

• Set up observability dashboards for LLM/agent systems (e.g. LangSmith, Arize, custom logging pipelines)

• Benchmark agent orchestration frameworks for performance, cost, and reliability

Backend & MCP Integration

• Build scalable APIs using FastAPI (sync & async execution)

• Implement Model Context Protocol (MCP) for secure tool and data access

• Manage agent state, context routing, and plugin-based workflows

MLOps & Deployment

• Deploy and monitor models in cloud environments (AWS / GCP / Azure)

• Work with model serving frameworks (e.g. vLLM, TGI) and apply quantization for efficient inference

• Implement logging, observability dashboards, and automated recovery workflows

Front-End Collaboration

• Build or collaborate on UI using React, TypeScript, or Next.js

• Create seamless UI–API bridges for agent interactions and basic dashboards

📌 AI Engineer (LLM & Agent Systems) (Mumbai)
🏢 SPEEGILE CONSULTING
📍 Mumbai

Reply to this offer

Impress this employer describing Your skills and abilities, fill out the form below and leave Your personal touch in the presentation letter.

Subscribe to this job alert:

Get the latest job offers by email for: ai engineer (llm & agent systems) (mumbai) / mumbai

Subscribe to this job alert:

Get the latest job offers by email for: ai engineer (llm & agent systems) (mumbai) / mumbai