Evaluators
Create scoring methods and configure confidence-based review routing.
An evaluator is a named scoring method backed by a judge model. Open Evaluation & Quality → Evaluators to view, edit, or create evaluators.
At the top, the page summarizes your setup:
Total and active evaluators
Available judge models from Anthropic, OpenAI, and Google
Agents with Online Eval enabled
Use Evaluators to manage scoring methods. Use Agent enablement to enable evaluators for each agent.

Evaluator fields
Evaluator
The evaluator name, method, and number of agents using it.
Framework
The scoring approach. Single judge uses one model for each output.
Judge model(s)
The model that performs scoring, such as gpt-5.1.
Conf. → HITL
The confidence threshold that routes results to Human Review, such as < 67%.
Status
Whether the evaluator is active.
Actions
A menu for editing or removing the evaluator.
Create an evaluator
Configure the evaluator
Complete the following fields:
Evaluator name — Required. Use a label such as Answer relevancy.
Framework — Single judge scores each output with one model. Jury vote is coming soon and unavailable.
Evaluation method — Select Answer Relevancy, Safety, or Task Completion.


Judge model — Required. Select the model that judges output.
Confidence threshold — Defaults to
50%. The judge returns a rating and confidence from 0–100%. Results below this threshold are flagged for human review, regardless of rating.Route low-confidence results to human review — Controls whether below-threshold results enter Human Review.
Last updated

