Architect, fine-tune, and deploy large language models (LLMs) and advanced retrieval-augmented generation (RAG) search pipelines for enterprise platforms.
Role Overview
As an LLM Engineer at Fospe, you will be at the forefront of designing, fine-tuning, and operationalizing cutting-edge Large Language Models and Retrieval-Augmented Generation (RAG) pipelines. You will collaborate directly with our AI research team to embed deterministic cognitive capabilities and high-throughput semantic reasoning into our vertical enterprise SaaS products.
Key Responsibilities
- ✓Architect, fine-tune, and quantize open-source foundation models (Llama 3, Mistral, Qwen, DeepSeek) for domain-specific enterprise workloads.
- ✓Design high-accuracy Retrieval-Augmented Generation (RAG) architectures with hybrid dense/sparse search, reranking, and semantic chunking.
- ✓Build low-latency streaming inference pipelines utilizing vLLM, TensorRT-LLM, and Triton Inference Server.
- ✓Implement guardrail validation, safety filters, prompt routing, and automated hallucination benchmarking.
- ✓Collaborate with backend engineers to expose scalable gRPC and REST interfaces for seamless frontend consumption.
Mandatory Requirements
Eligibility & Qualifications
- ✓Bachelor’s or Master’s degree in Computer Science, Artificial Intelligence, Data Science, or related quantitative field.
- ✓3+ years of qualified experience in machine learning and deep learning, with 1+ years dedicated to LLM engineering.
- ✓Deep hands-on expertise with PyTorch, Hugging Face Transformers, vLLM, and LangChain/LlamaIndex.
- ✓Proven experience with vector databases such as Pinecone, Qdrant, Milvus, or pgvector.
- ✓Strong proficiency in Python, asynchronous programming, Docker containerization, and Linux/CUDA environments.
- ✓Demonstrated understanding of model optimization techniques including LoRA, QLoRA, AWQ, and GPTQ quantization.
Preferred / Advantageous Qualifications
- Experience with multi-agent orchestration frameworks (LangGraph, CrewAI, AutoGen).
- Contributions to open-source AI libraries or published research in NLP/LLMs.
- Experience deploying models on Kubernetes and cloud infrastructure (AWS/GCP/Azure GPU instances).
Primary Technologies & Environment
PythonPyTorchHugging FacevLLMLangChainLlamaIndexQdrantDockerKubernetesCUDAFastAPI
Our Hiring Process
01
Application Review
Our technical talent leads review your GitHub portfolio, code repositories, and background.
02
Technical Screen
45-minute architectural discussion on NLP, transformer architectures, and RAG pipelines.
03
Deep-Dive Live Coding
Practical problem-solving session focusing on vector search, prompt routing, and latency optimization.
04
Executive Offer
Meet the AI leadership team, align on vision, and finalize your offer.
Summary Information
Department:Software
Location:Bangalore, India
Workplace:Hybrid
Employment:Full-time
Date Posted:2026-07-15
📌 LLM Engineer (Large Language Models) (India)
🏢 Fospe
📍 India