Looking for an AI Engineer to help test, evaluate, and improve the Agentic System being developed for our Operator Assist product. Key requirementsinclude:
Solid expertise in Python and a solid understanding of APIs and system integrations.
Deep knowledge of Generative AI, LLMs, and Agentic Systems.
Experience defining and measuring agent evaluation metrics, including tool-calling performance, accuracy, reliability, and overall system effectiveness.
Ability to quickly understand complex business domains by reviewing and analyzing large volumes of documentation and data.
Experience designing evaluation datasets, creating questions and answers, developing multi-turn conversational scenarios, benchmarking agent performance, and identifying areas for improvement.
Robust analytical and problem-solving skills, with a focus on improving agent quality and user outcomes.