Creating an eval
A short wizard walks you through three steps:- Target: what to run on. Choose
trace(score the whole run) orspan(score individual steps). Optionally filter by agent, trace name, and for spans, span type and model. - Check: what to verify. Pick a preset (below).
- Score: how to score. For LLM judges, pick a judge model. For checks that take a parameter, set it (a substring, pattern, or max length). Set a sample rate (1% to 100%) to control how much matching traffic is scored.
Code checks
Deterministic, free, and run without any external calls:LLM judges
Judges send the input and output to a model that returns a 0.00 to 1.00 score or a pass/fail verdict with a reason. Presets cover relevance, helpfulness, coherence, conciseness, instruction following, completeness, toxicity and safety, tool selection, and RAG checks (faithfulness, context relevance, and correctness against a reference).LLM judges use your own provider keys. Add one (below) before creating a
judge eval. An eval with no usable key shows the status needs key and
doesn’t score until a key is added. The available judge models depend on the
deployment.
.png?fit=max&auto=format&n=1sml0kwCw1BNtiz-&q=85&s=a7c846a133aa3480dc1e03a2da5fec2b)
.png?fit=max&auto=format&n=1sml0kwCw1BNtiz-&q=85&s=6534218552b79980770f11c40aafd6ec)