Skip to content

Quality checks

Quality checks let AI judges score a sample of an agent’s real answers, so you can see answer quality over time and get an alert when it drops. They’re off unless you pick a set for an agent, because each scored check costs credits.

Open Quality checks → New set:

Setting Meaning
What to check Ready-made groups: Answer quality (helpful, correct, on topic), Safety (no harmful content or stereotyping), Tool use (right tool, right details, job done).
Your own criteria One per line, in your own words - e.g. Mentions the 14-day return window when refunds come up. An AI judge checks each sampled answer against each line.
How many answers to check From 5% to every answer.
Alert me below Optional, 1-100. You’re notified when an agent’s average over the last day drops below it.

The form shows an estimate of the cost per 100 answers and at your recent traffic.

Pick the set under Quality checks in the Agent Builder. Editing a set later updates every agent using it, with no redeploy.

Scores arrive a few minutes after each answer. On the agent’s page, Answer quality shows:

  • the average (0-100, higher is better) over 7 or 30 days,
  • each check’s average and its change from the period before,
  • the latest scores with the judge’s one-line reason.

The Dashboard lists your agents’ quality, lowest first.

Each scored check costs credits - a judge that returns no score costs nothing. When your workspace runs out of credits or reaches its monthly limit, checks pause instead of running, and resume after a top-up. See Credits.