Visitar URL original
`AgentEvaluator.evaluate()` fails a lower-is-better metric whose evaluator reports PASSED · Issue #7414 · google/adk-python · GitHub
Skip to content

AgentEvaluator.evaluate() fails a lower-is-better metric whose evaluator reports PASSED #7414

Description

@gaurav-gandhi-2411

🔴 Required Information

Describe the Bug:
AgentEvaluator._process_metrics_and_get_failures decides pass/fail with overall_score >= threshold (agent_evaluator.py:901), after the INFORMATIONAL branch and without consulting the metric's own eval_status. A custom evaluator for a lower-is-better metric (error rate, latency, cost) that correctly reports PASSED is marked FAILED whenever its score is below the threshold, and evaluate() raises AssertionError. Setting the metric to EvalStatus.INFORMATIONAL avoids the false failure but also removes gating: a value over the limit (e.g. 0.9) then passes too, so a lower-is-better metric still cannot fail a run when it should.

Steps to Reproduce:

  1. Create the three files below in one directory.
  2. Run python driver.py with google-adk from origin/main.

lowerbetter/__init__.py

from . import agent

lowerbetter/agent.py

from __future__ import annotations

from google.adk.agents.llm_agent import LlmAgent
from google.adk.evaluation.eval_metrics import (
    EvalMetric,
    EvalStatus,
    Interval,
    MetricInfo,
    MetricValueInfo,
)
from google.adk.evaluation.evaluator import EvaluationResult, Evaluator, PerInvocationResult
from google.adk.evaluation.metric_evaluator_registry import DEFAULT_METRIC_EVALUATOR_REGISTRY
from google.adk.models.base_llm import BaseLlm
from google.adk.models.llm_response import LlmResponse
from google.genai import types as genai_types

METRIC_NAME = "error_rate"  # lower is better; 0.10 against a 0.50 limit is a PASS
LIMIT = 0.5
VALUE = 0.1


class ErrorRateEvaluator(Evaluator):
    def __init__(self, *, eval_metric: EvalMetric) -> None:
        self._eval_metric = eval_metric

    def evaluate_invocations(
        self, actual_invocations, expected_invocations=None, conversation_scenario=None
    ) -> EvaluationResult:
        results = [
            PerInvocationResult(
                actual_invocation=inv,
                score=VALUE,
                # The evaluator's own verdict: lower is better, 0.1 <= 0.5.
                eval_status=EvalStatus.PASSED if VALUE <= LIMIT else EvalStatus.FAILED,
            )
            for inv in actual_invocations
        ]
        return EvaluationResult(
            overall_score=VALUE,
            overall_eval_status=EvalStatus.PASSED,
            per_invocation_results=results,
        )


DEFAULT_METRIC_EVALUATOR_REGISTRY.register_evaluator(
    metric_info=MetricInfo(
        metric_name=METRIC_NAME,
        description="Toy lower-is-better metric.",
        metric_value_info=MetricValueInfo(interval=Interval(min_value=0.0, max_value=1.0)),
    ),
    evaluator=ErrorRateEvaluator,
)


class _FakeLlm(BaseLlm):
    model: str = "fake-model"

    @classmethod
    def supported_models(cls) -> list[str]:
        return ["fake-model"]

    async def generate_content_async(self, llm_request, stream: bool = False):
        yield LlmResponse(
            content=genai_types.Content(parts=[genai_types.Part(text="ok")], role="model")
        )


root_agent = LlmAgent(name="toy_agent", model=_FakeLlm(), instruction="Answer briefly.")

driver.py

from __future__ import annotations

import asyncio
import json
from pathlib import Path

from google.adk.evaluation.eval_case import EvalCase, Invocation
from google.adk.evaluation.eval_set import EvalSet
from google.genai import types as genai_types

case = EvalCase(
    eval_id="case_1",
    conversation=[
        Invocation(user_content=genai_types.Content(parts=[genai_types.Part(text="hi")], role="user"))
    ],
)
Path("eval_set.json").write_text(EvalSet(eval_set_id="s", eval_cases=[case]).model_dump_json())
Path("test_config.json").write_text(json.dumps({"criteria": {"error_rate": 0.5}}))


async def main() -> None:
    from google.adk.evaluation.agent_evaluator import AgentEvaluator

    await AgentEvaluator.evaluate(
        agent_module="lowerbetter",
        eval_dataset_file_path_or_dir="eval_set.json",
        num_runs=1,
        print_detailed_results=True,
    )


asyncio.run(main())

Expected Behavior:
The evaluator's own verdict decides (0.1 against a limit of 0.5 passes), or the metric can declare its direction. A minimal fix: when the evaluator sets eval_status, use it, and fall back to overall_score >= threshold only when it doesn't.

Observed Behavior:

Summary: `EvalStatus.FAILED` for Metric: `error_rate`. Expected threshold: `0.5`, actual value: `0.1`.
|    | eval_status       |   score |   threshold | prompt   | ...   (remaining columns trimmed)
|  0 | EvalStatus.PASSED |     0.1 |         0.5 | hi       | ...
AssertionError: Following are all the test failures.
error_rate for lowerbetter Failed. Expected 0.5, but got 0.1.

Environment Details:

  • ADK: origin/main 6be386d8939c4898aa2c0fc8a1f7cb34e7df7904; Python 3.11; Windows 11
  • Also reproduces on the released google-adk 2.11.0.

🟡 Optional Information

Additional Context:
Searched issues and PRs (lower-is-better, metric polarity, threshold direction, _process_metrics_and_get_failures, "honor each metric's own eval_status"): no duplicate. PR #6739 implements a fix and can be rebased if a maintainer wants it.

How often has this issue occurred?: Always (100%)

Activity

  1. self-assigned this
    on Oct 6, 2026
  2. added theissue type on Oct 7, 2026
  3. added
    eval[Component] This issue is related to evaluation
    on Oct 7, 2026
  4. llalitkumarrr commented on Oct 7, 2026

    @llalitkumarrr
    Collaborator

    Hello @gaurav-gandhi-2411,

    Thank you for sharing the detailed report and for your contribution. Your PR is currently under review and once it is approved, it can be merged. However it looks like there are some merge conflicts, could you please take a look and resolve them? We will reach out if we need any additional information. We appreciate your patience in the meantime.

  5. gaurav-gandhi-2411 commented on Oct 7, 2026

    @gaurav-gandhi-2411
    ContributorAuthor

    Thanks @llalitkumarrr — resolved and pushed. The conflict was with 2d428ea, which added the INFORMATIONAL early-exit and the threshold lookup to the same loop; I kept both and applied the polarity inference and the higher_is_better verdict after them. The changed test file passes locally against current main, as does the rest of tests/unittests/evaluation/.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

eval[Component] This issue is related to evaluation

Type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions