Company Description Heu.ai is an AI platform company focused on helping enterprises enhance their business and solve complex problems through AI-driven solutions.
About the Role
We are looking for a hands-on AI / LLM Engineer to design, build, and deploy intelligent system architectures and agentic workflows. In this role, you will go beyond simple wrapper scripts to build production-grade Retrieval-Augmented Generation (RAG) pipelines, complex multi-agent orchestrations, and robust backend APIs.
Experience
1-3 Years
Key Roles and Responsibilities
- Agent & Workflow Engineering: Design and deploy stateful, multi-agent orchestrations and dynamic decision-making graphs using LangGraph.
- Advanced RAG Pipelines: Build, evaluate, and optimize hybrid retrieval strategies (vector search, keyword/BM25, metadata filtering, reranking) and chunking techniques for high-precision context retrieval.
- Prompt Engineering & Evaluation: Architect, test, and iterate on structured prompts, system instructions, and dynamic context windows. Establish evaluation frameworks (e.g., automated benchmarks, ground-truth testing) to monitor quality and hallucination rates.
- API & Backend Development: Develop scalable, asynchronous REST/gRPC APIs and microservices using Python frameworks (FastAPI, Flask) to integrate AI capabilities into production applications.
- Data & Vector Store Management: Manage and optimize vector databases (e.g., Pinecone,
Qdrant, Chroma, PGVector) for efficient embedding indexing and search performance.
- System Monitoring & Optimization: Monitor latency, token usage, rate limits, and cost across LLM providers. Implement caching strategies (e.g., semantic caching) and streaming outputs for low-latency end-user experience.
Qualifications & Skills
- Programming Proficiency: Strong proficiency in Python with an emphasis on async programming, clean architecture, and type safety.
- Framework Experience: Demonstrated hands-on experience with LangGraph, LangChain, or custom agentic execution systems.
- API Engineering: Solid background in building and maintaining production APIs via FastAPI or up-to-date Web frameworks, including authentication, rate-limiting, and error-handling patterns.
- RAG & Vector Search: Proven experience implementing end-to-end RAG workflows with vector stores (Pinecone, Qdrant, PGVector, etc.) and embedding models.
- Prompt Architecture: Mastery of structured outputs (JSON schema, Pydantic), chain-of-thought prompting, and defensive prompt engineering against edge cases and injection.
- Tooling & Integrations: Experience integrating third-party APIs, webhooks, and custom tools into LLM workflows.
- Experience with fine-tuning open-source models (e.g., Llama, Mistral) or local LLM deployments (Ollama, vLLM).
📌 AI Engineer (Kochi)
🏢 Heu.Ai
📍 Kochi