Skip to content

Give rubric_based_* metrics a public rubric path (#132) - #229

Draft
erauner12 wants to merge 2 commits into
agentevals-dev:mainfrom
erauner12:rubrics-public-path
Draft

erauner12 wants to merge 2 commits into
agentevals-dev:mainfrom
erauner12:rubrics-public-path

Conversation

@erauner12

Copy link
Copy Markdown
Contributor

Draft, opened to make the shape of #132 concrete; happy to reshape it to the design you have in mind.

What this does

  • Config. A builtin evaluator takes rubrics: a list of {id, text} mappings, or plain strings that get positional ids (rubric_0, ...). An entry may carry type (an ADK rubric type); unset, the metric's own type is used. Duplicate ids are refused at load time. The field survives run-level overrides.
  • Eval set. The runner now matches the eval case, not only its conversation (_find_eval_case; _find_expected_invocations stays as a wrapper). For a rubric_based_* metric, the matched case's rubrics are attached to every actual invocation and each expected invocation's rubrics to the actual invocation at the same position, the way ADK's LocalEvalService does it, with the metric's type given to untyped rubrics because ADK's evaluators filter invocation rubrics by type. A duplicate id across metric, case and invocation is reported as an error.
  • No silent empty runs. A rubric metric with no rubric from the config, the case or an invocation reports an error before any judge call: "requires rubrics: set 'rubrics' on the evaluator in the eval config, or 'rubrics' on the matched eval case or its invocations".
  • Per-rubric scores. The metric's details carries per_invocation[].rubric_scores and overall_rubric_scores, each entry rubric_id, score, rationale.
  • API. /api/metrics lists the rubric metrics as working; requires_rubrics is unchanged.
  • Docs. README example, the type: builtin reference in docs/custom-evaluators.md, and a Rubrics section in docs/eval-set-format.md.

Example:

evaluators:
  - name: rubric_based_final_response_quality_v1
    type: builtin
    judge_model: gemini-2.5-flash
    rubrics:
      - id: corrected_value_wins
        text: The response uses the operator's latest correction, never a superseded value.
      - The response states no root cause the conversation did not establish.

Not in this draft

  • A CLI flag for rubrics. The config file and the eval set cover the CLI; a --rubric option can follow if wanted.
  • MCP: the evaluate tool already accepts an eval config path, so it gets rubrics from that.

Tests

tests/test_rubrics.py: config parsing (mappings, strings, refusals), overrides preserving rubrics, YAML loading, strings and objects reaching the criterion, case and invocation attachment with type defaulting and duplicate detection, the error without rubrics before any judge call, the dispatch passing both sources, case matching, and the details shape. The unit suite passes (765 passed, 7 skipped).

Closes #132 if the shape is acceptable.

The rubric-based metrics have been reachable only through the internal
build_eval_metric(..., rubrics=...) call: the metric builder accepted
rubrics, the eval-set format documented them on cases and invocations,
and nothing on the config, CLI or API path passed any of them through,
so the metrics were listed as not working.

Rubrics can now be declared on a builtin evaluator in the eval config,
as {id, text} mappings or plain strings with positional ids, with
duplicate ids refused. At run time the runner matches the eval case,
not only its conversation, and the matched case's rubrics and each
expected invocation's rubrics are attached to the actual invocations
the way ADK's own eval service attaches them, with the metric's rubric
type given to rubrics that have none, since ADK applies invocation
rubrics by type. A rubric metric with no rubric from the config, the
case or an invocation reports an error before any judge call rather
than an empty evaluation. Its result carries per-rubric scores per
invocation and overall in details. The API lists the rubric metrics as
working.

@erauner12 erauner12 left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tested this draft on a few real multi-turn traces and left three notes on behavior that seems worth deciding before the design is finalized. The rubric input path itself works but the main open questions are how case-level rubrics should apply across invocations, and whether tool declarations should be handled here or in a separate follow-up.

My preference is final-invocation-only behavior or an explicit applies_to option, but I’m happy to follow the maintainers’ direction. The per-rubric result shape stays the same either way, so downstream consumers are not blocked on that decision.

case_rubrics: list[Rubric] | None,
) -> list[Invocation]:
"""Give the actual invocations the eval case's rubrics, as ADK's own eval
service does: case-level rubrics on every invocation, invocation-level

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tried this default against the pinned 0.9.10 CLI using real multi-turn traces (three-turn conversations: two clarification turns followed by a final draft, rubric_based_final_response_quality_v1, five samples). One behavior stood out: case-level rubrics describing the end state ("produces the labelled draft", "carries the operator's answers") failed on every intermediate turn. The metric evaluates each invocation in isolation, using only that turn's prompt, steps, and response. In this sample, two-thirds of the case mean came from those intermediate-turn penalties, while the same rubrics mostly passed on the final invocation. I also saw an answer from two turns earlier treated as "not supplied in the prompt" when the final invocation was graded.

Keeping case-level rubrics on every invocation matches ADK's eval service, so I left the implementation that way for now. For a per-invocation metric, though, I'm not sure it's the best default for a conversation. A few options we could consider:

  1. Keep ADK parity (the current implementation) and document the tradeoff; I've added a note to docs/eval-set-format.md.
  2. Apply case-level rubrics only to the final invocation for rubric_based_* metrics, since those metrics score final response and tool use per turn and case-level properties often describe the overall outcome.
  3. Add per-rubric applies_to: final | every, defaulting to ADK's every, so the eval set can opt into the intended behavior without duplicating the rubric.

My preference is 2 or 3, but I'm happy to follow the maintainers' preference. The id-keyed per-rubric verdicts in details—the piece downstream consumers rely on—stay the same in all three approaches.


if metric_name in METRICS_NEEDING_RUBRICS:
try:
actual_invocations = attach_case_rubrics(

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I noticed one related issue in the same measurement that had an even larger effect on scores. The rubric metrics aren't in _METRICS_NEEDING_INVOCATION_EVENTS, so they don't receive the app_details tool declarations synthesized by _enrich_app_details. ADK's rubric prompt only trusts output from "procedurally sound" tool calls and flags a call to a tool that doesn't exist. Without those declarations, the rater treated every tool call—including the clarification tool carrying the operator's answers—as invalid, and therefore treated the tool output as untrusted evidence. That led factual rubrics to fail with explanations such as "the values come only from the invalid tool call," even when the value also appeared in the user's prompt.

This feels adjacent to, rather than part of, the rubric input path covered by this PR. That said, it affects whether rubrics can pass on traces with tool calls. Would you prefer that I extend _enrich_app_details to METRICS_NEEDING_RUBRICS here, or open a focused follow-up issue? I'm happy with either direction.

Comment thread docs/eval-set-format.md

`rubrics` on an eval case apply to every invocation of that case; `rubrics` on an invocation apply to that invocation alone. Each rubric is an ADK `Rubric`: `rubric_id`, `rubric_content.text_property`, and an optional `type`. When a `rubric_based_*` metric runs, the runner adds the matched case's rubrics and each invocation's rubrics to the rubrics the metric's config declares, the way ADK's own eval service does, and a rubric without a `type` is given the metric's (`FINAL_RESPONSE_QUALITY` or `TOOL_USE_QUALITY`), since ADK applies invocation-level rubrics by type. Rubric ids must be distinct across the metric, the case and the invocation; a clash is reported as an error. With no rubric from any of the three, the metric reports an error.

The `rubric_based_*` metrics grade each invocation on its own, with that invocation's prompt, steps and final response, so a case-level rubric is judged once per turn. In a multi-turn case a rubric that describes the end state ("produces the completed draft") is therefore judged against the intermediate turns as well, where it fails, and lowers the case's mean. Write such a rubric on the final invocation instead, and keep case-level rubrics for properties every turn must have.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I added this note after testing the draft on multi-turn traces. Without it, someone using case-level rubrics on a conversation can see a low mean that doesn't really reflect the quality of the final answer. If we choose option 2 or 3 from the attach_case_rubrics comment, I'll update this paragraph to match the behavior we settle on.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Clarify supported rubric input path for rubric_based_* metrics

1 participant