Blog
Blog
Notes on testing conversational AI at scale: what we have learned about evaluation models, judge design, and the difference between a passing metric and a release-worthy agent.
S
Testing Strategy6 min read
Why Conversational AI Needs a Different Testing Model
The same test persona can take a different but equally valid path every run. That breaks the assumptions conventional QA is built on.
Read article
S
Quality Assurance5 min read
The Problem With Green Checkmarks on Broken Conversations
When a test runner reports success on a conversation that never reached its goal, your metrics are lying to you.
Read article
S
Evaluation Model
What 'LLM-as-Judge' Actually Means in Practice
Subjective quality is not a bug. It is a dimension that deterministic assertions were never designed to capture.
Coming soon