20 Sep
|
SPEEGILE CONSULTING
|
Mumbai
20 Sep
SPEEGILE CONSULTING
Mumbai
About the Role
We are looking for an experienced AI Engineer to design, build, and deploy LLM-driven agentic systems end-to-end — from data and model fine-tuning to retrieval, evaluation, guardrails, and production deployment. You should be equally comfortable reasoning about how a transformer works internally and shipping a reliable, observable agent system in production. This role owns the AI/ML depth of our stack, with support from backend and frontend engineers for integration.
Experience: 3+ Years
Qualification: B.E. / B.Tech / M.Sc. / M.Tech / MCA — Computer Science / IT or related field
Location: Preference Mumbai / Relocation to Mumbai / Remote
Employment Type: Full-Time
Notice Period: Max 30 Days
Must-Have Skills
- 3+ years of experience building Python-based ML/AI systems in production
- Strong, practical understanding of how transformer, encoder-decoder, and deep learning models work internally — not just API-level usage
- Hands-on experience fine-tuning LLMs (LoRA / QLoRA / full fine-tuning) and working with large-scale datasets end-to-end
- Foundational Knowledge of LLMs and Can Lay the foundation for building a LLM model from scratch.
- Solid grasp of the full LLM landscape — RAG, vectored and vector-less retrieval, evaluation, observability, and guardrails
- Experience with voice/speech models (ASR/TTS & Translation) in Fine Tuning the models can even lay the foundation to build a voice model from scratch
- Knowledge graphs, Redis, Advanced RAG , Self corrective RAG
- Experience with cloud AI deployments (AWS / GCP / Azure)
- Hands-on experience with LLM agent frameworks and vector databases (ChromaDB, Weaviate, pgvector)
- Strong knowledge of PyTorch or TensorFlow and Scikit-Learn
- Production experience with FastAPI, Docker, and MLOps
Good-to-Have Skills
- Advanced prompt optimization and agent evaluation techniques
- Experience building AI observability tools from scratch
- Exposure to multilingual or low-resource language models
- Expert-level usage of agentic coding IDEs (Cursor, Windsurf, Claude Code)
Soft Skills
- Solid problem-solving and system design mindset
- Ability to work across AI, backend, and front-end teams
- Comfortable mentoring freshers/junior engineers on AI fundamentals
- Clear communication and documentation skills
- Passion for building production-grade AI systems
What You'll Work On
LLM Agents & Prompt Engineering
- Design and implement LLM agents using LangGraph, PydanticAI, and Google ADK
- Build tool-augmented reasoning pipelines: RAG, Chain-of-Thought, ReAct, and planner–executor architectures
- Develop robust, tested prompt strategies to improve reliability and reduce hallucination
Model Fundamentals & Fine-Tuning
- Deep working knowledge of transformer architecture, Mamba architecture— encoder-only, decoder-only, and encoder–decoder models — and how attention, embeddings, and positional encoding actually work under the hood
- Fine-tune LLMs using LoRA, QLoRA, and full fine-tuning depending on the use case and compute budget. If the fine tuning doesn’t get the desired results, a LLM model will be built from scratch
- Prepare, clean, and process large-scale datasets for pre-training/fine-tuning — deduplication, tokenization, sampling, and quality filtering at scale
- Understanding of core deep learning model families (CNNs, RNNs/LSTMs, transformers) and when to use each
- Work with voice/speech models — ASR (speech-to-text) and TTS (text-to-speech) — and understand how they integrate into conversational AI pipelines
Retrieval, RAG & Knowledge Systems
- Implement both vectored (embedding-based) and vector-less (keyword/graph/hybrid) retrieval strategies, choosing the right approach per use case
- Integrate vector databases (FAISS, Pinecone, pgvector, ChromaDB, Weaviate) and knowledge graphs (Neo4j)
- Design chunking, embedding, and re-ranking strategies for high-precision retrieval
- Implement Prompt Engineering, Context Engineering, Loop Engineering & Efficient Low token retrieval
Evaluation, Observability & Guardrails
- Build offline and online evaluation harnesses for agent and LLM outputs (accuracy, groundedness, latency, cost)
- Implement guardrails for safety, PII redaction, and scope control (e.g. NeMo Guardrails, Guardrails AI, or custom rule/LLM-based filters)
- Build tooling for trace analysis, state debugging, and hallucination detection
- Set up observability dashboards for LLM/agent systems (e.g. LangSmith, Arize, custom logging pipelines)
- Benchmark agent orchestration frameworks for performance, cost, and reliability
Backend & MCP Integration
- Build scalable APIs using FastAPI (sync & async execution)
- Implement Model Context Protocol (MCP) for secure tool and data access
- Manage agent state, context routing, and plugin-based workflows
MLOps & Deployment
- Deploy and monitor models in cloud environments (AWS / GCP / Azure)
- Work with model serving frameworks (e.g. vLLM, TGI) and apply quantization for efficient inference
- Implement logging, observability dashboards, and automated recovery workflows
Front-End Collaboration
- Build or collaborate on UI using React, TypeScript, or Next.js
- Create seamless UI–API bridges for agent interactions and basic dashboards
📌 AI Engineer (LLM & Agent Systems) (Mumbai)
🏢 SPEEGILE CONSULTING
📍 Mumbai