Quality checks
Quality checks let AI judges score a sample of an agent’s real answers, so you can see answer quality over time and get an alert when it drops. They’re off unless you pick a set for an agent, because each scored check costs credits.
Create a set
Section titled “Create a set”Open Quality checks → New set:
| Setting | Meaning |
|---|---|
| What to check | Ready-made groups: Answer quality (helpful, correct, on topic), Safety (no harmful content or stereotyping), Tool use (right tool, right details, job done). |
| Your own criteria | One per line, in your own words - e.g. Mentions the 14-day return window when refunds come up. An AI judge checks each sampled answer against each line. |
| How many answers to check | From 5% to every answer. |
| Alert me below | Optional, 1-100. You’re notified when an agent’s average over the last day drops below it. |
The form shows an estimate of the cost per 100 answers and at your recent traffic.
Attach to agents
Section titled “Attach to agents”Pick the set under Quality checks in the Agent Builder. Editing a set later updates every agent using it, with no redeploy.
Read the results
Section titled “Read the results”Scores arrive a few minutes after each answer. On the agent’s page, Answer quality shows:
- the average (0-100, higher is better) over 7 or 30 days,
- each check’s average and its change from the period before,
- the latest scores with the judge’s one-line reason.
The Dashboard lists your agents’ quality, lowest first.
Each scored check costs credits - a judge that returns no score costs nothing. When your workspace runs out of credits or reaches its monthly limit, checks pause instead of running, and resume after a top-up. See Credits.