> For the complete documentation index, see [llms.txt](https://docs.uptiq.ai/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.uptiq.ai/task-guides/evaluate-agent-quality.md).

# Evaluate agent quality

Every Evals task in Control Center: create evaluators, score agents' live traffic, read scores and runs, and review low-confidence results.

This guide is organized by task. Find what you want to do in the list below, then follow the steps. Each task links to the reference page it comes from.

An evaluator is a named scoring method backed by a judge model, defined at the account level. You enable evaluators per agent to score its live traffic, and low-confidence results can go to people for review.

| Task                                          | Go to                                                        |
| --------------------------------------------- | ------------------------------------------------------------ |
| See the account's evaluators and judge models | [Evaluators](#see-the-evaluators)                            |
| Create an evaluator                           | [Create](#create-an-evaluator)                               |
| Choose between one judge and a jury           | [Framework](#choose-between-a-single-judge-and-a-jury)       |
| Edit or remove an evaluator                   | [Edit](#edit-or-remove-an-evaluator)                         |
| Choose which evaluators score an agent        | [Assign](#assign-evaluators-to-an-agent)                     |
| Start scoring an agent's live traffic         | [Turn on](#turn-on-online-eval-for-an-agent)                 |
| Score only part of an agent's traffic         | [Sampling](#score-only-part-of-the-traffic)                  |
| Read the scores                               | [Scores](#read-the-scores)                                   |
| See why a result went to Human Review         | [Routing](#see-why-a-result-went-to-human-review)            |
| See the judge's reasoning and evidence        | [Reasoning](#see-the-judge-reasoning-and-evidence)           |
| Find out why a score is missing               | [Runs](#find-out-why-a-score-is-missing)                     |
| Read the conversation behind a score          | [Linked conversation](#read-the-conversation-behind-a-score) |
| Review a low-confidence result                | [Human Review](#review-a-low-confidence-result)              |
| Export evaluation results                     | [Export](#export-evaluation-results)                         |

***

## <i class="fa-scale-balanced">:scale-balanced:</i> Evaluators

### See the evaluators

Open **Evals** › **Evaluators**. The summary shows how many **Evaluators** exist and are active, how many **Judge models** are available and from which providers, and the **Online eval agents** scoring live traffic out of all agents, such as `9 / 71`.

Each evaluator row shows its name and **Active** status, its method and how many agents use it, its **Framework**, its **Judge model(s)**, its **Confidence Threshold**, and an **Actions** menu.

<figure><img src="https://1326225582-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F0qmgQjJ5aArDTj2ACFHG%2Fuploads%2Fgit-blob-09c733456ef7dc554abc33f0ba234e3a60430219%2Fqore_control-center_evals-evaluators_rounded_shadow.png?alt=media" alt="The Evaluators tab: summary cards for evaluators, judge models and online eval agents, above a table of evaluators with their framework, judge models and confidence threshold"><figcaption><p>Evaluators.</p></figcaption></figure>

### Create an evaluator

{% stepper %}
{% step %}

#### Open the dialog

Select **+ New evaluator**. You need the **Build Evals** capability; without it the button is disabled.
{% endstep %}

{% step %}

#### Name it and choose the framework

* **Evaluator name** — required, such as *Answer relevancy*.
* **Framework** — **Single judge** or **Jury**.
  {% endstep %}

{% step %}

#### Choose the method

**Evaluation method**: **Answer Relevancy**, **Faithfulness**, **Safety**, or **Task Completion**.
{% endstep %}

{% step %}

#### Choose the judge models

**Judge model** for a single judge, or **Jury models** (2 to 5, from any providers) for a jury.
{% endstep %}

{% step %}

#### Set the threshold and routing

* **Confidence threshold** — default `50%`. Results below it are flagged, regardless of rating.
* **Route low-confidence results to human review** — off by default. While it's off, low-confidence results are still scored but auto-accepted.
  {% endstep %}

{% step %}

#### Save

Select **Create**, or **Create & assign agents** to go straight to choosing the agents it scores. Both are enabled once you set a name, a method, and the judge model or jury models.
{% endstep %}
{% endstepper %}

<figure><img src="https://1326225582-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F0qmgQjJ5aArDTj2ACFHG%2Fuploads%2Fgit-blob-2a8588db801be86987209321641f750eb76fb84b%2Fqore_control-center_evals-new-evaluator-single-judge_rounded_shadow.png?alt=media" alt="The New evaluator dialog with Single judge selected: evaluator name, framework, evaluation method, judge model, a 50% confidence threshold, and the route-to-human-review toggle"><figcaption><p>New evaluator.</p></figcaption></figure>

**Faithfulness** rates how much of what the agent said is supported by its context: retrieved context, tool results, and what the user said. The score is supported claims divided by all factual claims. Reference: [Evaluators › Faithfulness](/control-center/evals/evaluators.md#faithfulness).

### Choose between a single judge and a jury

| Framework        | Judge models  | Scoring                                                                                                |
| ---------------- | ------------- | ------------------------------------------------------------------------------------------------------ |
| **Single judge** | One model     | That model scores each output directly.                                                                |
| **Jury**         | 2 to 5 models | Each judge scores independently with its own confidence; Qore combines them into one weighted verdict. |

Each evaluator's confidence threshold is independent of every other evaluator's. Reference: [Evaluators › Choose a framework](/control-center/evals/evaluators.md#choose-a-framework).

### Edit or remove an evaluator

Use the evaluator's **Actions** menu.

***

## <i class="fa-gauge-high">:gauge-high:</i> Score live traffic

### Assign evaluators to an agent

{% stepper %}
{% step %}

#### Open Agent enablement

On **Evaluators**, select the **Agent enablement** tab. Filter by project, or search for an agent. An agent with the same name in two projects has two rows.
{% endstep %}

{% step %}

#### Choose the evaluators

Select **Edit** in the agent's **Assigned evaluators** cell. **Assign evaluators** lists every evaluator in the account. Tick the ones that should score this agent, then select **Save**.
{% endstep %}
{% endstepper %}

<figure><img src="https://1326225582-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F0qmgQjJ5aArDTj2ACFHG%2Fuploads%2Fgit-blob-7c67dc82118a4380fd1b996d53406dc1e81a78fc%2Fqore_control-center_evals-assign-evaluators_rounded_shadow.png?alt=media" alt="The Assign evaluators dialog for an agent, with a checkbox for each evaluator in the account and Cancel and Save buttons"><figcaption><p>Assign evaluators.</p></figcaption></figure>

Assigning evaluators doesn't start scoring on its own.

### Turn on Online Eval for an agent

On **Agent enablement**, turn on the agent's **Online eval** toggle. While it's on, every enabled evaluator assigned to the agent scores the sampled share of its conversations.

An agent can show **Unsupported** instead of a toggle: its type doesn't support Online Eval yet. You can still pre-assign evaluators with **Edit**.

<figure><img src="https://1326225582-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F0qmgQjJ5aArDTj2ACFHG%2Fuploads%2Fgit-blob-ae1679bcbddda6aa2aa3079554bc5f703d2a6a44%2Fqore_control-center_evals-agent-enablement_rounded_shadow.png?alt=media" alt="The Agent enablement tab, listing agents with their project, assigned evaluators, a sampling percentage, and an Online eval toggle"><figcaption><p>Agent enablement.</p></figcaption></figure>

### Score only part of the traffic

On **Agent enablement**, set the agent's **Sampling**: the chance that each live conversation is evaluated. The default is `100%`; lower it on high-traffic agents to keep judge costs down. Online Eval scoring is billed as `OnlineEval` on **Credit Usage** › **Features**.

Reference: [Evaluators › Agent enablement](/control-center/evals/evaluators.md#agent-enablement).

***

## <i class="fa-square-poll-vertical">:square-poll-vertical:</i> Read results

### Read the scores

Open **Evals** › **Online Eval**, on the **Scores** tab. Filter by project, agent, minimum score (**Score ≥**), verdict, or routing, and search by execution, agent, or evaluator.

The summary counts what's in view: **Evaluations in view**, **Passed**, **Routed to Human Review**, and **Avg. score** with average confidence. Each row shows the **Execution**, **Agent**, **Evaluator**, **Score** (`pass`/`fail` or a percentage), **Conf.** (confidence over the threshold, such as `87%` over `≥ 95%`), **Verdict**, **Routing**, and **Scored** time.

<figure><img src="https://1326225582-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F0qmgQjJ5aArDTj2ACFHG%2Fuploads%2Fgit-blob-636b637a8ce5ca98ee165012e966b0da2f1fd95b%2Fqore_control-center_evals-online-eval-scores_rounded_shadow.png?alt=media" alt="The Online Eval Scores tab: project, agent, score, verdict and routing filters; summary cards for evaluations in view, passed, routed to Human Review and average score; and a table of scored executions"><figcaption><p>Online Eval, Scores tab.</p></figcaption></figure>

### See why a result went to Human Review

An execution routes to Human Review when its confidence is below the evaluator's threshold **and** the evaluator has **Route low-confidence results to human review** on. This can happen even when the verdict is *Passed*. A *Failed* execution is auto-accepted when confidence meets the threshold.

### See the judge reasoning and evidence

Expand a score row. **Judge reasoning** explains the score. **Evidence** lists up to three points, from the one that moved the score most, each citing something the agent did or failed to do. For a **Jury** evaluator, each judge shows its own.

### Find out why a score is missing

Open the **Runs** tab. Each sampled conversation creates one run per assigned evaluator. Filter by project, agent, or status, and search by execution. Statuses are **Pending**, **Running**, **Completed**, **Failed**, and **Skipped**. For **Failed**, hover over the status for the reason:

* `Judge model unavailable: <model>`
* `Evaluation method not found`
* `Evaluation timed out` — the run stayed Pending or Running for over two hours.
* `Evaluation failed`

**Skipped** means the evaluator was deleted before scoring, or the conversation had nothing to score.

### Read the conversation behind a score

Select **Open linked conversation**, labelled with the run's execution ID. It opens Runs & Traces filtered to that conversation; the filtered view survives a reload, so you can share its link.

Reference: [Online Eval](/control-center/evals/online-eval.md).

***

## <i class="fa-user-check">:user-check:</i> Review

### Review a low-confidence result

{% stepper %}
{% step %}

#### Open Human Review

Open **Evals** › **Human Review**, or select **Open board** on **Online Eval**.
{% endstep %}

{% step %}

#### Choose a queue

Switch between **My tasks**, **All tasks**, and **Unassigned**, and filter by team, type, status, or priority. **View unassigned** opens unassigned tasks.
{% endstep %}

{% step %}

#### Record a verdict

Open a task. Its evaluation summary shows the judge's reasoning and evidence. Record a verdict; the task moves from **Review Pending** to **Reviewed**.
{% endstep %}
{% endstepper %}

<figure><img src="https://1326225582-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F0qmgQjJ5aArDTj2ACFHG%2Fuploads%2Fgit-blob-c14b5c4b2324732bee5e6e1c303b72092e17b62d%2Fqore_control-center_evals-human-review_rounded_shadow.png?alt=media" alt="The Human Review screen: My tasks, All tasks and Unassigned queues, team, type, status and priority filters, summary cards, and a table of tasks with team, assignee, priority, status and verdict"><figcaption><p>Human Review.</p></figcaption></figure>

{% hint style="info" %}
Human Review assigns tasks through teams, not to individuals. A reviewer without a team has no assigned tasks. Set up teams in [Manage users and access › Create a team](/task-guides/manage-users-and-access.md#create-a-team).
{% endhint %}

Reference: [Human Review](/control-center/evals/human-review.md).

### Export evaluation results

Online Eval has no export button of its own. Open **Observe** › **Export** and choose the **Eval Results** dataset: scores, verdicts, reasoning, evidence, and human overrides. See [Monitor activity › Export](/task-guides/monitor-activity.md#export-activity-data).

***

## <i class="fa-link">:link:</i> Related pages

* [Evals](/control-center/evals.md), [Evaluators](/control-center/evals/evaluators.md), [Online Eval](/control-center/evals/online-eval.md), and [Human Review](/control-center/evals/human-review.md)
* Test an agent before release: [Create an agent › Build a test dataset](/task-guides/create-an-agent.md#build-a-test-dataset)

***

{% columns %}
{% column width="83.33333333333334%" %}

<p align="right"><em>Maintained by</em> <mark style="color:green;">Abhishek Paul</mark><br><code>AI-assisted, human-approved</code></p>
{% endcolumn %}

{% column width="16.666666666666664%" %}

<div align="left"><img src="https://1326225582-files.gitbook.io/~/files/v0/b/gitbook-x-prod.appspot.com/o/spaces%2F0qmgQjJ5aArDTj2ACFHG%2Fuploads%2Fgit-blob-d034645b4f8f8ee661985f5306b60ded3289653c%2Fmaintainer-abhishek-paul.png?alt=media" alt="" width="60"></div>
{% endcolumn %}
{% endcolumns %}


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.uptiq.ai/task-guides/evaluate-agent-quality.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
