Scale existing RAG solution in Azure cloud to much more users and data, improve performance
Essential functions
Refactor current solution to make it scalable and reliable
Prompt Engineering Libraries/Plugins: Develop libraries or plugins to improve response consistency and quality through pre-built prompts, templates, and best practices for effective prompt crafting.
Performance tuning to handle large payloads, such as performance testing scripts, to deliver complete responses.
Develop strategies to manage token limitations, enabling the tool to facilitate complex and informative interactions without restrictions.
Extend support for a wide range of file formats, including XML, HTML, and HAR, in addition to standard Microsoft product documents (Word, PowerPoint, Excel, PDF).
Leverage multi-vector retrieval techniques to enhance the tool's functionality and effectively handle large data sets.
Qualifications
Robust expertise in Python and PySpark with over 5 years of hands-on experience in developing scalable data processing and analytics solutions.
Azure cloud experience.
Experience of developing and refactoring scalable and robust solutions.
Experience with GenAI/LLM frameworks and technics, like guardrails, Langchain, etc.
Would be a plus
Solid experience with GenAI components, like vector databases (e.g. Cosmos DB, FAISS etc)
Practical experience of developing and deploying GenAI/LLMs/SLMs-based solutions.