SHYENA
Blog

Blog

Notes on testing conversational AI at scale: what we have learned about evaluation models, judge design, and the difference between a passing metric and a release-worthy agent.