17 Aug
|
Iprogrammer
|
Pune
Key Responsibilities
Design, develop, and maintain GenAI-powered applications such as chatbots, customer support assistants, recommendation assistants, and document/query intelligence systems.
Build and optimize RAG pipelines including document ingestion, chunking, embeddings, vector search, retrieval logic, reranking, and response generation.
Work on LLM orchestration using frameworks such as LangChain, LangGraph, LlamaIndex, or similar tools.
Develop backend AI services using Python, FastAPI, REST APIs , and microservice-based architecture.
Implement intent classification, entity extraction, query parsing, and structured JSON output generation for business workflows.
Work with open-source and cloud-based LLMs, including local model setup, model serving, and inference optimization.
Implement guardrails to prevent hallucination, unsafe responses, competitor comparisons, irrelevant answers, prompt injection, and data leakage.
Optimize AI systems for latency, accuracy, reliability, token usage, and scalability .
Integrate AI workflows with backend systems, databases, APIs, CRM/ERP systems, payment flows, and business applications.
Create evaluation datasets, regression test cases, accuracy reports, and failure analysis for AI responses.
Collaborate with product managers, backend developers, DevOps teams, and business stakeholders to convert business requirements into scalable AI solutions.
Required Skills
Robust hands-on experience with Python and backend development.
Practical experience in building LLM-based applications using OpenAI, Claude, Gemini, Llama, Mistral, or other open-source models.
Good understanding of RAG architecture , embeddings, vector databases, semantic search, hybrid search, and retrieval optimization.
Experience with frameworks such as LangChain, LangGraph, LlamaIndex , or similar orchestration tools.
Knowledge of vector databases such as FAISS, ChromaDB, Pinecone, Weaviate, Qdrant, Milvus, or pgvector.
Experience with prompt engineering , system prompts, structured output generation, function calling/tool calling, and JSON schema-based responses.
Understanding of LLM guardrails , safety filters, fallback handling, confidence scoring, and hallucination control.
Experience with FastAPI / Flask / Django , REST APIs, and backend integration.
Understanding of Docker, Git, Linux basics , and deployment workflows.
Ability to debug production AI issues related to latency, incorrect responses, token limits, retrieval failure, context mismatch, and model output inconsistency.
Good to Have Skills
Experience with vLLM, Ollama, Hugging Face Transformers, TensorRT-LLM , or other model serving frameworks.
Knowledge of model quantization, inference optimization, batching, GPU utilization , and token streaming.
Experience with AWS, Azure, or GCP for AI/ML deployment.
Exposure to telecom, fintech, customer support, billing, recharge, payments, or high-scale consumer applications.
Experience in multilingual AI systems , especially Hinglish or Indian language handling.
Understanding of ASR/transcription-based input normalization will be a plus.
Experience with monitoring tools, logging, prompt/version management, and AI evaluation frameworks.
Knowledge of MCP, agentic workflows, tool-based reasoning, or multi-step AI orchestration will be an advantage.
Candidate Profile
Experience: 2 to 5 years
Education: BE/BTech/MTech/MCA in Computer Science, AI/ML, Data Science, IT, or equivalent practical experience.
The candidate should have built at least one real-world AI/GenAI application involving LLMs, backend integration, RAG, or production deployment.
The candidate should be able to explain implementation details clearly, not just theoretical AI concepts.
📌 AI Engineer (Pune)
🏢 Iprogrammer
📍 Pune