Looking for an AI Engineer to help test, evaluate, and improve the Agentic System being developed for our Operator Assist product. Key requirementsinclude:
• Strong expertise in Python and a solid understanding of APIs and system integrations.
• Deep knowledge of Generative AI, LLMs, and Agentic Systems.
• Experience defining and measuring agent evaluation metrics, including tool-calling performance, accuracy, reliability, and overall system effectiveness.
• Ability to quickly understand complex business domains by reviewing and analyzing large volumes of documentation and data.
• Experience designing evaluation datasets, creating questions and answers, developing multi-turn conversational scenarios, benchmarking agent performance, and identifying areas for improvement.
• Robust analytical and problem-solving skills, with a focus on improving agent quality and user outcomes.