28 Aug
|
iProgrammer Solutions
|
Pune
28 Aug
iProgrammer Solutions
Pune
Key Responsibilities:
- Design, develop, and maintain GenAI-powered applications such as chatbots, customer support assistants, recommendation assistants, and document/query intelligence systems.
- Build and optimize RAG pipelines including document ingestion, chunking, embeddings, vector search, retrieval logic, reranking, and response generation.
- Work on LLM orchestration using frameworks such as LangChain, LangGraph, LlamaIndex, or similar tools.
- Develop backend AI services using Python, FastAPI, REST APIs, and microservice-based architecture.
- Implement intent classification, entity extraction, query parsing, and structured JSON output generation for business workflows.
- Work with open-source and cloud-based LLMs, including local model setup, model serving, and inference optimization.
- Implement guardrails to prevent hallucination, unsafe responses, competitor comparisons, irrelevant answers, prompt injection, and data leakage.
- Optimize AI systems for latency, accuracy, reliability, token usage, and scalability.
- Integrate AI workflows with backend systems, databases, APIs, CRM/ERP systems, payment flows, and business applications.
- Create evaluation datasets, regression test cases, accuracy reports, and failure analysis for AI responses.
- Collaborate with product managers, backend developers, DevOps teams, and business stakeholders to convert business requirements into scalable AI solutions.
Required Skills:
- Robust hands-on experience with Python and backend development.
- Practical experience in building LLM-based applications using OpenAI, Claude, Gemini, Llama, Mistral, or other open-source models.
- Good understanding of RAG architecture, embeddings, vector databases, semantic search, hybrid search, and retrieval optimization.
- Experience with frameworks such as LangChain, LangGraph, LlamaIndex, or similar orchestration tools.
- Knowledge of vector databases such as FAISS, ChromaDB, Pinecone, Weaviate, Qdrant, Milvus, or pgvector.
- Experience with prompt engineering, system prompts, structured output generation, function calling/tool calling, and JSON schema-based responses.
- Understanding of LLM guardrails, safety filters, fallback handling, confidence scoring, and hallucination control.
- Experience with FastAPI / Flask / Django, REST APIs, and backend integration.
- Understanding of Docker, Git, Linux basics, and deployment workflows.
- Ability to debug production AI issues related to latency, incorrect responses, token limits, retrieval failure, context mismatch, and model output inconsistency.
Good to Have Skills:
- Experience with vLLM, Ollama, Hugging Face Transformers, TensorRT-LLM, or other model serving frameworks.
- Knowledge of model quantization, inference optimization, batching, GPU utilization, and token streaming.
- Experience with AWS, Azure, or GCP for AI/ML deployment.
- Exposure to telecom, fintech, customer support, billing, recharge, payments, or high-scale consumer applications.
- Experience in multilingual AI systems, especially Hinglish or Indian language handling.
- Understanding of ASR/transcription-based input normalization will be a plus.
- Experience with monitoring tools, logging, prompt/version management, and AI evaluation frameworks.
- Knowledge of MCP, agentic workflows, tool-based reasoning, or multi-step AI orchestration will be an advantage.
Candidate Profile
- Experience: 2 to 5 years
- Education: BE/BTech/MTech/MCA in Computer Science, AI/ML, Data Science, IT, or equivalent practical experience.
- The candidate should have built at least one real-world AI/GenAI application involving LLMs, backend integration, RAG, or production deployment.
- The candidate should be able to explain implementation details clearly, not just theoretical AI concepts.
📌 Artificial Intelligence Engineer (Pune)
🏢 iProgrammer Solutions
📍 Pune