Next upHack for Humanity: San Francisco (powered by Google Gemini)
1:01:39

GPT on a Leash: Evaluating LLM-based Apps & Mitigating Their Risks

The task of testing and evaluating AI systems is extremely challenging, especially when it involves text and unstructured data. In the case of LLM-based applications, these challenges are magnified by the fact that there isn't "one correct answer" and by a combination of various external constraints such as topics that shouldn't be discussed. Speaker: Philip is the co-founder and CEO of Deepchecks. Philip is an experienced Data Scientist and in the past, he led a top-tier ML research gr

Dmytro Spodarets
Dmytro Spodarets
Dec 15, 2023
Summary

Philip, co-founder and CEO of Deepchecks, on the challenges of testing and evaluating LLM-based applications, where there is no single correct answer, and how to mitigate their risks under external constraints.