Conversation
The rubric-based metrics have been reachable only through the internal
build_eval_metric(..., rubrics=...) call: the metric builder accepted
rubrics, the eval-set format documented them on cases and invocations,
and nothing on the config, CLI or API path passed any of them through,
so the metrics were listed as not working.
Rubrics can now be declared on a builtin evaluator in the eval config,
as {id, text} mappings or plain strings with positional ids, with
duplicate ids refused. At run time the runner matches the eval case,
not only its conversation, and the matched case's rubrics and each
expected invocation's rubrics are attached to the actual invocations
the way ADK's own eval service attaches them, with the metric's rubric
type given to rubrics that have none, since ADK applies invocation
rubrics by type. A rubric metric with no rubric from the config, the
case or an invocation reports an error before any judge call rather
than an empty evaluation. Its result carries per-rubric scores per
invocation and overall in details. The API lists the rubric metrics as
working.
erauner12
left a comment
There was a problem hiding this comment.
I tested this draft on a few real multi-turn traces and left three notes on behavior that seems worth deciding before the design is finalized. The rubric input path itself works but the main open questions are how case-level rubrics should apply across invocations, and whether tool declarations should be handled here or in a separate follow-up.
My preference is final-invocation-only behavior or an explicit applies_to option, but I’m happy to follow the maintainers’ direction. The per-rubric result shape stays the same either way, so downstream consumers are not blocked on that decision.
| case_rubrics: list[Rubric] | None, | ||
| ) -> list[Invocation]: | ||
| """Give the actual invocations the eval case's rubrics, as ADK's own eval | ||
| service does: case-level rubrics on every invocation, invocation-level |
There was a problem hiding this comment.
I tried this default against the pinned 0.9.10 CLI using real multi-turn traces (three-turn conversations: two clarification turns followed by a final draft, rubric_based_final_response_quality_v1, five samples). One behavior stood out: case-level rubrics describing the end state ("produces the labelled draft", "carries the operator's answers") failed on every intermediate turn. The metric evaluates each invocation in isolation, using only that turn's prompt, steps, and response. In this sample, two-thirds of the case mean came from those intermediate-turn penalties, while the same rubrics mostly passed on the final invocation. I also saw an answer from two turns earlier treated as "not supplied in the prompt" when the final invocation was graded.
Keeping case-level rubrics on every invocation matches ADK's eval service, so I left the implementation that way for now. For a per-invocation metric, though, I'm not sure it's the best default for a conversation. A few options we could consider:
- Keep ADK parity (the current implementation) and document the tradeoff; I've added a note to
docs/eval-set-format.md. - Apply case-level rubrics only to the final invocation for
rubric_based_*metrics, since those metrics score final response and tool use per turn and case-level properties often describe the overall outcome. - Add per-rubric
applies_to: final | every, defaulting to ADK'severy, so the eval set can opt into the intended behavior without duplicating the rubric.
My preference is 2 or 3, but I'm happy to follow the maintainers' preference. The id-keyed per-rubric verdicts in details—the piece downstream consumers rely on—stay the same in all three approaches.
|
|
||
| if metric_name in METRICS_NEEDING_RUBRICS: | ||
| try: | ||
| actual_invocations = attach_case_rubrics( |
There was a problem hiding this comment.
I noticed one related issue in the same measurement that had an even larger effect on scores. The rubric metrics aren't in _METRICS_NEEDING_INVOCATION_EVENTS, so they don't receive the app_details tool declarations synthesized by _enrich_app_details. ADK's rubric prompt only trusts output from "procedurally sound" tool calls and flags a call to a tool that doesn't exist. Without those declarations, the rater treated every tool call—including the clarification tool carrying the operator's answers—as invalid, and therefore treated the tool output as untrusted evidence. That led factual rubrics to fail with explanations such as "the values come only from the invalid tool call," even when the value also appeared in the user's prompt.
This feels adjacent to, rather than part of, the rubric input path covered by this PR. That said, it affects whether rubrics can pass on traces with tool calls. Would you prefer that I extend _enrich_app_details to METRICS_NEEDING_RUBRICS here, or open a focused follow-up issue? I'm happy with either direction.
|
|
||
| `rubrics` on an eval case apply to every invocation of that case; `rubrics` on an invocation apply to that invocation alone. Each rubric is an ADK `Rubric`: `rubric_id`, `rubric_content.text_property`, and an optional `type`. When a `rubric_based_*` metric runs, the runner adds the matched case's rubrics and each invocation's rubrics to the rubrics the metric's config declares, the way ADK's own eval service does, and a rubric without a `type` is given the metric's (`FINAL_RESPONSE_QUALITY` or `TOOL_USE_QUALITY`), since ADK applies invocation-level rubrics by type. Rubric ids must be distinct across the metric, the case and the invocation; a clash is reported as an error. With no rubric from any of the three, the metric reports an error. | ||
|
|
||
| The `rubric_based_*` metrics grade each invocation on its own, with that invocation's prompt, steps and final response, so a case-level rubric is judged once per turn. In a multi-turn case a rubric that describes the end state ("produces the completed draft") is therefore judged against the intermediate turns as well, where it fails, and lowers the case's mean. Write such a rubric on the final invocation instead, and keep case-level rubrics for properties every turn must have. |
There was a problem hiding this comment.
I added this note after testing the draft on multi-turn traces. Without it, someone using case-level rubrics on a conversation can see a low mean that doesn't really reflect the quality of the final answer. If we choose option 2 or 3 from the attach_case_rubrics comment, I'll update this paragraph to match the behavior we settle on.
Draft, opened to make the shape of #132 concrete; happy to reshape it to the design you have in mind.
What this does
builtinevaluator takesrubrics: a list of{id, text}mappings, or plain strings that get positional ids (rubric_0, ...). An entry may carrytype(an ADK rubric type); unset, the metric's own type is used. Duplicate ids are refused at load time. The field survives run-level overrides._find_eval_case;_find_expected_invocationsstays as a wrapper). For arubric_based_*metric, the matched case'srubricsare attached to every actual invocation and each expected invocation'srubricsto the actual invocation at the same position, the way ADK'sLocalEvalServicedoes it, with the metric's type given to untyped rubrics because ADK's evaluators filter invocation rubrics by type. A duplicate id across metric, case and invocation is reported as an error.detailscarriesper_invocation[].rubric_scoresandoverall_rubric_scores, each entryrubric_id,score,rationale./api/metricslists the rubric metrics asworking;requires_rubricsis unchanged.type: builtinreference indocs/custom-evaluators.md, and a Rubrics section indocs/eval-set-format.md.Example:
Not in this draft
--rubricoption can follow if wanted.evaluatetool already accepts an eval config path, so it gets rubrics from that.Tests
tests/test_rubrics.py: config parsing (mappings, strings, refusals), overrides preserving rubrics, YAML loading, strings and objects reaching the criterion, case and invocation attachment with type defaulting and duplicate detection, the error without rubrics before any judge call, the dispatch passing both sources, case matching, and the details shape. The unit suite passes (765 passed, 7 skipped).Closes #132 if the shape is acceptable.