Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 16 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -240,6 +240,22 @@ evaluators:
threshold: 0.7
```

Rubric-based metrics take their rubrics from the same config, as `{id, text}` mappings or plain strings, and score each rubric per invocation:

```yaml
evaluators:
- name: rubric_based_final_response_quality_v1
type: builtin
judge_model: gemini-2.5-flash
rubrics:
- id: corrected_value_wins
text: The response uses the operator's latest correction, never a superseded value.
- id: no_invented_cause
text: The response states no root cause the conversation did not establish.
```

Rubrics on a matched eval case (`rubrics` on the case, or on one of its invocations) are added to the metric's own at run time; see the [eval set format](docs/eval-set-format.md#rubrics).

Evaluators with a `requirements.txt` get automatic virtual environment management. You can also use `type: remote` for community evaluators from GitHub, or `type: openai_eval` to delegate grading to the [OpenAI Evals API](https://developers.openai.com/api/reference/resources/evals/methods/create) (requires `pip install "agentevals-cli[openai]"`).

See the [Custom Evaluators guide](docs/custom-evaluators.md) for the full protocol reference, SDK helpers, and how to contribute evaluators.
Expand Down
15 changes: 15 additions & 0 deletions docs/custom-evaluators.md
Original file line number Diff line number Diff line change
Expand Up @@ -89,6 +89,21 @@ agentevals run traces/my_trace.json \

Each evaluator entry in the `evaluators` list uses the following fields. The `type` field determines which other fields are valid.

### `type: builtin` (ADK metrics)

| Field | Required | Default | Description |
|---|---|---|---|
| `name` | yes | | The metric, for example `tool_trajectory_avg_score` or `rubric_based_final_response_quality_v1` |
| `type` | yes | | `builtin` |
| `threshold` | no | `0.5` | Score at or above this value means PASSED |
| `judge_model` | no | ADK default | Judge model for LLM-backed metrics |
| `trajectory_match_type` | no | `EXACT` | `EXACT`, `IN_ORDER` or `ANY_ORDER`, for `tool_trajectory_avg_score` |
| `rubrics` | for `rubric_based_*` | | A list of `{id, text}` mappings, or plain strings (given ids `rubric_0`, `rubric_1`, ...). Ids must be unique. An entry may also carry `type`, an ADK rubric type; unset, the metric's own type is used |
| `credential_ref` | no | | Name of a credential the judge key is resolved from (API runs) |
| `judge_base_url` | no | | Base URL of an OpenAI-compatible judge endpoint |

A `rubric_based_*` metric with no rubric from the config, the matched eval case or its invocations reports an error rather than scoring nothing. Its result carries `details.per_invocation[].rubric_scores` and `details.overall_rubric_scores`, one entry per rubric id with its score and the judge's rationale.

### `type: code` (local scripts)

| Field | Required | Default | Description |
Expand Down
16 changes: 16 additions & 0 deletions docs/eval-set-format.md
Original file line number Diff line number Diff line change
Expand Up @@ -176,6 +176,22 @@ An eval case can have multiple invocations to represent a conversation. Each inv
| `rubrics` | list[Rubric] | no | Scoring rubrics for this specific invocation |
| `creation_timestamp` | float | no | Unix timestamp |

### Rubrics

`rubrics` on an eval case apply to every invocation of that case; `rubrics` on an invocation apply to that invocation alone. Each rubric is an ADK `Rubric`: `rubric_id`, `rubric_content.text_property`, and an optional `type`. When a `rubric_based_*` metric runs, the runner adds the matched case's rubrics and each invocation's rubrics to the rubrics the metric's config declares, the way ADK's own eval service does, and a rubric without a `type` is given the metric's (`FINAL_RESPONSE_QUALITY` or `TOOL_USE_QUALITY`), since ADK applies invocation-level rubrics by type. Rubric ids must be distinct across the metric, the case and the invocation; a clash is reported as an error. With no rubric from any of the three, the metric reports an error.

The `rubric_based_*` metrics grade each invocation on its own, with that invocation's prompt, steps and final response, so a case-level rubric is judged once per turn. In a multi-turn case a rubric that describes the end state ("produces the completed draft") is therefore judged against the intermediate turns as well, where it fails, and lowers the case's mean. Write such a rubric on the final invocation instead, and keep case-level rubrics for properties every turn must have.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added this note after testing the draft on multi-turn traces. Without it, someone using case-level rubrics on a conversation can see a low mean that doesn't really reflect the quality of the final answer. If we choose option 2 or 3 from the attach_case_rubrics comment, I'll update this paragraph to match the behavior we settle on.


```json
{
"eval_id": "resume-correction",
"rubrics": [
{"rubric_id": "corrected_value_wins", "rubric_content": {"text_property": "The response uses the operator's latest correction."}}
],
"conversation": [ ... ]
}
```

### Content

Uses the Google GenAI `Content` format:
Expand Down
2 changes: 1 addition & 1 deletion src/agentevals/api/routes.py
Original file line number Diff line number Diff line change
Expand Up @@ -242,7 +242,7 @@ async def list_metrics():
requires_gcp=m.metric_name in METRICS_NEEDING_GCP,
requires_rubrics=m.metric_name in _METRICS_NEEDING_RUBRICS,
description=m.description or "No description available",
working=m.metric_name not in _METRICS_NEEDING_RUBRICS,
working=True,
)
)

Expand Down
121 changes: 117 additions & 4 deletions src/agentevals/builtin_metrics.py
Original file line number Diff line number Diff line change
Expand Up @@ -160,6 +160,19 @@ def _enrich_app_details(invocations: list[Invocation]) -> list[Invocation]:
return [inv.model_copy(update={"app_details": app_details}) for inv in invocations]


METRICS_NEEDING_RUBRICS = {
"rubric_based_final_response_quality_v1",
"rubric_based_tool_use_quality_v1",
}

# The ADK rubric type each rubric metric applies to invocation-level rubrics;
# a rubric written for the metric without a type is given this one.
_METRIC_RUBRIC_TYPES = {
"rubric_based_final_response_quality_v1": "FINAL_RESPONSE_QUALITY",
"rubric_based_tool_use_quality_v1": "TOOL_USE_QUALITY",
}


def rubric_strings_to_objects(rubric_texts: list[str]) -> list[Rubric]:
"""Convert plain-text rubric strings into ADK Rubric objects."""
return [
Expand All @@ -171,14 +184,88 @@ def rubric_strings_to_objects(rubric_texts: list[str]) -> list[Rubric]:
]


def to_rubric_objects(rubrics: list[str | Rubric] | None) -> list[Rubric]:
"""Plain strings become positional rubrics; ADK ``Rubric`` objects pass through."""
if not rubrics:
return []
texts = [r for r in rubrics if isinstance(r, str)]
objects = [r for r in rubrics if isinstance(r, Rubric)]
return rubric_strings_to_objects(texts) + objects


def attach_case_rubrics(
metric_name: str,
actual_invocations: list[Invocation],
expected_invocations: list[Invocation] | None,
case_rubrics: list[Rubric] | None,
) -> list[Invocation]:
"""Give the actual invocations the eval case's rubrics, as ADK's own eval
service does: case-level rubrics on every invocation, invocation-level

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried this default against the pinned 0.9.10 CLI using real multi-turn traces (three-turn conversations: two clarification turns followed by a final draft, rubric_based_final_response_quality_v1, five samples). One behavior stood out: case-level rubrics describing the end state ("produces the labelled draft", "carries the operator's answers") failed on every intermediate turn. The metric evaluates each invocation in isolation, using only that turn's prompt, steps, and response. In this sample, two-thirds of the case mean came from those intermediate-turn penalties, while the same rubrics mostly passed on the final invocation. I also saw an answer from two turns earlier treated as "not supplied in the prompt" when the final invocation was graded.

Keeping case-level rubrics on every invocation matches ADK's eval service, so I left the implementation that way for now. For a per-invocation metric, though, I'm not sure it's the best default for a conversation. A few options we could consider:

  1. Keep ADK parity (the current implementation) and document the tradeoff; I've added a note to docs/eval-set-format.md.
  2. Apply case-level rubrics only to the final invocation for rubric_based_* metrics, since those metrics score final response and tool use per turn and case-level properties often describe the overall outcome.
  3. Add per-rubric applies_to: final | every, defaulting to ADK's every, so the eval set can opt into the intended behavior without duplicating the rubric.

My preference is 2 or 3, but I'm happy to follow the maintainers' preference. The id-keyed per-rubric verdicts in details—the piece downstream consumers rely on—stay the same in all three approaches.

rubrics from the expected invocation at the same position. A rubric
without a type gets the metric's, since ADK filters invocation rubrics by
type. Duplicate ids within one invocation are an error."""
metric_type = _METRIC_RUBRIC_TYPES.get(metric_name)
attached = []
for index, actual in enumerate(actual_invocations):
extra: list[Rubric] = []
if case_rubrics:
extra.extend(case_rubrics)
if expected_invocations and index < len(expected_invocations) and expected_invocations[index].rubrics:
extra.extend(expected_invocations[index].rubrics or [])
if not extra:
attached.append(actual)
continue
merged: dict[str, Rubric] = {r.rubric_id: r for r in actual.rubrics or []}
for rubric in extra:
if rubric.rubric_id in merged:
raise ValueError(
f"Rubric id '{rubric.rubric_id}' is defined more than once for invocation {index} "
f"(eval case, invocation and metric rubrics must have distinct ids)."
)
merged[rubric.rubric_id] = (
rubric if rubric.type or not metric_type else rubric.model_copy(update={"type": metric_type})
)
attached.append(actual.model_copy(update={"rubrics": list(merged.values())}))
return attached


def extract_rubric_details(eval_result: EvaluationResult) -> dict[str, Any]:
"""Per-rubric scores, per invocation and overall, from a rubric metric."""

def _score(rubric_score: Any) -> dict[str, Any]:
return {
"rubric_id": rubric_score.rubric_id,
"score": rubric_score.score,
"rationale": rubric_score.rationale,
}

per_invocation = []
for per_inv_result in eval_result.per_invocation_results:
actual_inv = per_inv_result.actual_invocation
per_invocation.append(
{
"invocation_id": actual_inv.invocation_id if actual_inv else None,
"rubric_scores": [_score(r) for r in per_inv_result.rubric_scores or []],
}
)
return {
"per_invocation": per_invocation,
"overall_rubric_scores": [_score(r) for r in eval_result.overall_rubric_scores or []],
}


def build_eval_metric(
metric_name: str,
judge_model: str | None,
threshold: float | None,
rubrics: list[str] | None = None,
rubrics: list[str | Rubric] | None = None,
match_type: str | None = None,
) -> EvalMetric:
"""Construct an ADK ``EvalMetric`` with the appropriate criterion."""
"""Construct an ADK ``EvalMetric`` with the appropriate criterion.

``rubrics`` feed the ``rubric_based_*`` metrics' criterion: plain strings
get positional ids, ADK ``Rubric`` objects are used as given.
"""
effective_threshold = threshold if threshold is not None else 0.5

criterion: BaseCriterion | None = None
Expand Down Expand Up @@ -211,7 +298,7 @@ def build_eval_metric(
judge_opts = JudgeModelOptions()
if judge_model:
judge_opts.judge_model = judge_model
rubric_objects = rubric_strings_to_objects(rubrics) if rubrics else []
rubric_objects = to_rubric_objects(rubrics)
criterion = RubricsBasedCriterion(
threshold=effective_threshold,
judge_model_options=judge_opts,
Expand Down Expand Up @@ -370,9 +457,17 @@ async def evaluate_builtin_metric(
match_type: str | None = None,
credential_ref: str | None = None,
judge_base_url: str | None = None,
rubrics: list[str | Rubric] | None = None,
case_rubrics: list[Rubric] | None = None,
) -> dict[str, Any]:
"""Evaluate a single built-in ADK metric.

``rubrics`` are the metric's own (from the eval config); ``case_rubrics``
are the matched eval case's, applied to every invocation, and each
expected invocation's own rubrics apply to the actual invocation at the
same position. A rubric metric with no rubric from any of these is an
error, not an empty evaluation.

Returns a dict with keys: metric_name, score, eval_status,
per_invocation_scores, error, details.
"""
Expand All @@ -387,8 +482,24 @@ async def evaluate_builtin_metric(
),
)

if metric_name in METRICS_NEEDING_RUBRICS:
try:
actual_invocations = attach_case_rubrics(

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I noticed one related issue in the same measurement that had an even larger effect on scores. The rubric metrics aren't in _METRICS_NEEDING_INVOCATION_EVENTS, so they don't receive the app_details tool declarations synthesized by _enrich_app_details. ADK's rubric prompt only trusts output from "procedurally sound" tool calls and flags a call to a tool that doesn't exist. Without those declarations, the rater treated every tool call—including the clarification tool carrying the operator's answers—as invalid, and therefore treated the tool output as untrusted evidence. That led factual rubrics to fail with explanations such as "the values come only from the invalid tool call," even when the value also appeared in the user's prompt.

This feels adjacent to, rather than part of, the rubric input path covered by this PR. That said, it affects whether rubrics can pass on traces with tool calls. Would you prefer that I extend _enrich_app_details to METRICS_NEEDING_RUBRICS here, or open a focused follow-up issue? I'm happy with either direction.

metric_name, actual_invocations, expected_invocations, case_rubrics
)
except ValueError as exc:
return MetricResult(metric_name=metric_name, error=str(exc))
if not rubrics and not any(inv.rubrics for inv in actual_invocations):
return MetricResult(
metric_name=metric_name,
error=(
f"Metric '{metric_name}' requires rubrics: set 'rubrics' on the evaluator in the eval "
"config, or 'rubrics' on the matched eval case or its invocations."
),
)

try:
eval_metric = build_eval_metric(metric_name, judge_model, threshold, match_type=match_type)
eval_metric = build_eval_metric(metric_name, judge_model, threshold, rubrics=rubrics, match_type=match_type)
evaluator: Evaluator = get_evaluator(eval_metric)

if credential_ref:
Expand Down Expand Up @@ -425,6 +536,8 @@ async def evaluate_builtin_metric(
details = None
if metric_name == "tool_trajectory_avg_score":
details = extract_trajectory_details(eval_result)
elif metric_name in METRICS_NEEDING_RUBRICS:
details = extract_rubric_details(eval_result)

return MetricResult(
metric_name=metric_name,
Expand Down
74 changes: 73 additions & 1 deletion src/agentevals/config.py
Original file line number Diff line number Diff line change
Expand Up @@ -4,11 +4,14 @@

from collections import Counter
from pathlib import Path
from typing import Annotated, Any, Literal
from typing import TYPE_CHECKING, Annotated, Any, Literal

from pydantic import BaseModel, ConfigDict, Field, field_validator
from pydantic.alias_generators import to_camel

if TYPE_CHECKING:
from google.adk.evaluation.eval_rubrics import Rubric


def _normalize_trajectory_match_type(v: str | None) -> str | None:
valid = {"EXACT", "IN_ORDER", "ANY_ORDER"}
Expand All @@ -17,6 +20,53 @@ def _normalize_trajectory_match_type(v: str | None) -> str | None:
return v.upper() if v is not None else v


class RubricDef(BaseModel):
"""One rubric for a ``rubric_based_*`` metric, as written in an eval config.

A rubric is a testable statement about the response or the tool use, with
an id that names it in the per-rubric scores. In YAML an entry is either a
mapping ``{id, text}`` or a plain string, which gets a positional id
(``rubric_0``, ``rubric_1``, ...).
"""

model_config = ConfigDict(alias_generator=to_camel, populate_by_name=True, extra="forbid")

id: str = Field(min_length=1, description="Unique id of the rubric within the metric.")
text: str = Field(min_length=1, description="The testable statement the judge assesses.")
type: str | None = Field(
default=None,
description=(
"Optional ADK rubric type (for example FINAL_RESPONSE_QUALITY). Left unset, the metric's own type is used."
),
)

def to_adk(self, default_type: str | None = None) -> Rubric:
from google.adk.evaluation.eval_rubrics import Rubric, RubricContent

return Rubric(
rubric_id=self.id,
rubric_content=RubricContent(text_property=self.text),
type=self.type or default_type,
)


def _normalize_rubrics(value: Any) -> Any:
"""Accept plain strings beside ``{id, text}`` mappings; reject duplicate ids."""
if value is None:
return None
if not isinstance(value, list):
raise ValueError("'rubrics' must be a list of strings or {id, text} mappings")
normalized: list[Any] = []
for index, entry in enumerate(value):
if isinstance(entry, str):
if not entry.strip():
raise ValueError(f"Rubric {index} is empty")
normalized.append({"id": f"rubric_{index}", "text": entry})
else:
normalized.append(entry)
return normalized


class BuiltinMetricDef(BaseModel):
"""A built-in ADK metric, optionally with threshold/judge overrides."""

Expand All @@ -27,6 +77,13 @@ class BuiltinMetricDef(BaseModel):
threshold: float | None = Field(default=None, ge=0, le=1)
judge_model: str | None = None
trajectory_match_type: str | None = None
rubrics: list[RubricDef] | None = Field(
default=None,
description=(
"Rubrics for rubric_based_* metrics: a list of {id, text} mappings or plain strings. "
"Rubrics on the matched eval case or its invocations are added at run time."
),
)
credential_ref: str | None = Field(
default=None,
description="Logical name of a RunSpec.credential_refs entry whose resolved value is the judge API key.",
Expand All @@ -41,6 +98,21 @@ class BuiltinMetricDef(BaseModel):
def _validate_trajectory_match_type(cls, v: str | None) -> str | None:
return _normalize_trajectory_match_type(v)

@field_validator("rubrics", mode="before")
@classmethod
def _normalize_rubrics(cls, v: Any) -> Any:
return _normalize_rubrics(v)

@field_validator("rubrics")
@classmethod
def _unique_rubric_ids(cls, v: list[RubricDef] | None) -> list[RubricDef] | None:
if v is None:
return v
duplicates = sorted(rubric_id for rubric_id, count in Counter(r.id for r in v).items() if count > 1)
if duplicates:
raise ValueError("Rubric ids must be unique within a metric. Duplicate ids: " + ", ".join(duplicates))
return v


class BaseEvaluatorDef(BaseModel):
"""Shared fields for all executable evaluator definitions."""
Expand Down
6 changes: 6 additions & 0 deletions src/agentevals/custom_evaluators.py
Original file line number Diff line number Diff line change
Expand Up @@ -432,6 +432,7 @@ async def evaluate_custom_evaluator(
actual_invocations: list[Invocation],
expected_invocations: list[Invocation] | None,
performance_metrics: dict[str, Any] | None = None,
eval_case=None,
):
"""Evaluate a single custom evaluator and return a ``MetricResult``.

Expand All @@ -446,6 +447,9 @@ async def evaluate_custom_evaluator(
from .runner import MetricResult

if isinstance(evaluator_def, BuiltinMetricDef):
from .builtin_metrics import _METRIC_RUBRIC_TYPES

rubric_type = _METRIC_RUBRIC_TYPES.get(evaluator_def.name)
return await evaluate_builtin_metric(
metric_name=evaluator_def.name,
actual_invocations=actual_invocations,
Expand All @@ -455,6 +459,8 @@ async def evaluate_custom_evaluator(
match_type=evaluator_def.trajectory_match_type,
credential_ref=evaluator_def.credential_ref,
judge_base_url=evaluator_def.judge_base_url,
rubrics=[r.to_adk(rubric_type) for r in evaluator_def.rubrics] if evaluator_def.rubrics else None,
case_rubrics=list(eval_case.rubrics) if eval_case is not None and eval_case.rubrics else None,
)

if isinstance(evaluator_def, OpenAIEvalDef):
Expand Down
Loading
Loading