Pricing
Simple pricing for serious testing
Flexible seat and usage-based pricing that scales from your first pilot to a company-wide release gate. Every plan includes core evaluation and execution-integrity gating.
Starter and Growth prices are list price. Enterprise is scoped to your deployment.
Starter
Piloting evaluation on one agent with core quality metrics.
$750/mo
2,000 evaluated conversations included · $0.20 per conversation after
- One agent / one environment
- 2,000 evaluated conversations / mo
- Core LLM-as-judge metrics
- Deterministic assertion contracts
- Execution-integrity hard gate
- Email support
Most Popular
Growth
Multi-agent teams that need scheduled regression runs and richer quality gates.
$3,500/mo
Starting price · 15,000 conversations included · $0.15 per conversation after
- Up to 5 agents & environments
- 15,000 evaluated conversations / mo
- Full metric suite (semantic assurance, accessibility gates, orchestrator quality)
- Scheduled regression runs
- Slack & webhook alerts
- Priority support
Enterprise
Organization-wide deployment with custom metrics and enterprise controls.
Custom
Typically $75k–150k+/yr, scoped to your deployment and volume
- Unlimited agents & environments
- Negotiated evaluation volume
- SSO / SAML and role-scoped access
- Dedicated success engineer
- Custom metric development
- VPC / on-prem deployment option
- Contractual SLAs
Compare plans
What's included in each tier
All plans include the core evaluation engine. The differences are in scale, advanced quality gates, and enterprise controls.
CapabilityStarterGrowthEnterprise
Agentic test personas
LLM-as-judge metrics
Deterministic assertions
Execution-integrity gate
Semantic assurance
Accessibility scanning
Full audit trail
SSO / SAML
Dedicated support
Custom integrations
FAQ
Questions about pricing and plans
See it evaluate your own agent
Bring one real scenario. We'll run it against your live conversational AI agent and walk through every judged turn with you.