14 Aug
|
Nice Software Solutions
|
Nagpur
14 Aug
Nice Software Solutions
Nagpur
Role: LLM Engineer
Experience: 4+ Years
Location: Pune
Role Overview
We are seeking an experienced LLM Engineer to build, deploy, and scale production-grade Generative AI systems spanning agentic workflows, Retrieval-Augmented Generation (RAG), evaluation, and backend infrastructure. The role focuses on reliably serving LLM and agentic applications at scale while optimizing quality, cost, and latency.
Key Responsibilities
Agentic Systems
- Build agent workflows using frameworks such as LangGraph, CrewAI, AutoGen, or OpenAI Agents SDK.
- Implement function calling and tool use to enable agents to reliably interact with APIs, databases, and internal services.
- Use MCP (Model Context Protocol) or similar standards to connect agents with external tools and data sources.
- Implement guardrails and fallback mechanisms for predictable and safe multi-step agent behavior.
Retrieval-Augmented Generation (RAG)
- Build and optimize RAG pipelines, including chunking strategies, embedding model selection, and re-ranking.
- Deploy, scale, and evaluate vector databases such as Pinecone, Qdrant, FAISS, pgvector, ChromaDB, Milvus, or Azure AI Search.
- Diagnose hallucinations and poor responses by identifying retrieval-related issues and improving the underlying RAG pipeline.
Evaluation & Quality
- Build and maintain evaluation datasets, golden datasets, and regression suites covering retrieval accuracy, answer quality, and agent task completion.
- Use evaluation frameworks such as RAGAS, DeepEval, TruLens, LangSmith, or LangFuse.
- Track metrics including faithfulness, context precision/recall, answer relevance, and groundedness.
- Monitor quality metrics across models and prompt versions to identify regressions and support model/provider decisions.
- Combine automated evaluation with human-in-the-loop reviews for nuanced correctness, tone, and safety scenarios.
1. Backend Engineering & Deployment
- Build and maintain APIs and microservices using FastAPI or Flask to expose agent and RAG capabilities.
- Integrate LLM providers such as OpenAI, Anthropic Claude, Gemini, and open-source models.
- Handle retries, rate limits, and provider fallback mechanisms.
- Deploy and scale LLMs using vLLM, TGI, TensorRT-LLM, SGLang, or LMDeploy.
- Build and maintain inference infrastructure including GPU provisioning, autoscaling, load balancing, and multi-model routing.
- Optimize latency, throughput, and cost using quantization, batching, caching, asynchronous processing, and model routing.
- Containerize and orchestrate services using Docker and Kubernetes/ECS with CI/CD pipelines.
- Manage AWS/Azure cloud and GPU infrastructure for training and inference workloads.
Observability & MLOps
- Implement observability for LLM applications covering latency, cost, token usage, quality, and error rates.
- Version prompts, models, and pipelines to support rollback and reproducibility.
- Set up monitoring, logging, and alerting for production systems, including uptime, error rates, and cost tracking.
Applied ML Knowledge
- Demonstrate strong understanding of embeddings, fine-tuning,
and prompting to make sound architecture decisions.
- Collaborate with model evaluation teams to select appropriate models and approaches for specific use cases.
Required Skills
- Solid backend/systems engineering experience with Python, FastAPI/Flask, and REST APIs.
- Hands-on experience with agent frameworks such as LangGraph, CrewAI, AutoGen, or OpenAI Agents SDK.
- Hands-on experience with MCP (Model Context Protocol).
- Experience with LLM serving frameworks such as vLLM, TGI, TensorRT-LLM, SGLang, or LMDeploy.
- Strong experience building and tuning RAG pipelines.
- Experience with vector databases such as FAISS, Pinecone, Qdrant, ChromaDB, Milvus, or pgvector.
- Experience with evaluation frameworks such as RAGAS, DeepEval, TruLens, LangSmith, or LangFuse.
- Strong knowledge of Docker, Kubernetes/ECS, CI/CD, and cloud infrastructure including AWS/Azure GPU, EC2, and S3.
- Experience with model optimization techniques such as quantization, distillation, and batching.
- Familiarity with observability tools such as Prometheus, Grafana, and LangFuse.
- Strong understanding of LLM, RAG, and agentic AI fundamentals, including embeddings, fine-tuning, and prompting.
Valuable to Have
- Experience with streaming architectures such as Kafka for real-time AI pipelines.
- Cost optimization or FinOps experience for GPU workloads.
- Experience deploying multi-agent or RAG systems at enterprise scale.
- Experience integrating multiple LLM providers including OpenAI, Anthropic Claude, Gemini, and open-source models.
Education
- B.Tech/M.Tech in Computer Science, Information Technology, or a related field.
📌 LLM Engineer (Nagpur)
🏢 Nice Software Solutions
📍 Nagpur