Scale existing RAG solution in Azure cloud to much more users and data, improve performance
Essential functions
- Refactor current solution to make it scalable and stable
- Prompt Engineering Libraries/Plugins: Develop libraries or plugins to improve response consistency and quality through pre-built prompts, templates, and best practices for effective prompt crafting.
- Performance tuning to handle large payloads, such as performance testing scripts, to deliver complete responses.
- Develop strategies to manage token limitations, enabling the tool to facilitate complex and informative interactions without restrictions.
- Extend support for a wide range of file formats, including XML, HTML, and HAR, in addition to standard Microsoft product documents (Word, PowerPoint, Excel, PDF).
- Leverage multi-vector retrieval techniques to enhance the tool's functionality and effectively handle large data sets.
Qualifications
- Robust expertise in Python and PySpark with over 5 years of hands-on experience in developing scalable data processing and analytics solutions.
- Azure cloud experience.
- Experience of developing and refactoring scalable and robust solutions.
- Experience with GenAI/LLM frameworks and technics, like guardrails, Langchain, etc.
Would be a plus
- Strong experience with GenAI components, like vector databases (e.g. Cosmos DB, FAISS etc)
- Practical experience of developing and deploying GenAI/LLMs/SLMs-based solutions.