🔴 Required Information
Describe the Bug:
AgentEvaluator._process_metrics_and_get_failures decides pass/fail with overall_score >= threshold (agent_evaluator.py:901), after the INFORMATIONAL branch and without consulting the metric's own eval_status. A custom evaluator for a lower-is-better metric (error rate, latency, cost) that correctly reports PASSED is marked FAILED whenever its score is below the threshold, and evaluate() raises AssertionError. Setting the metric to EvalStatus.INFORMATIONAL avoids the false failure but also removes gating: a value over the limit (e.g. 0.9) then passes too, so a lower-is-better metric still cannot fail a run when it should.
Steps to Reproduce:
- Create the three files below in one directory.
- Run
python driver.py with google-adk from origin/main.
lowerbetter/__init__.py
lowerbetter/agent.py
from __future__ import annotations
from google.adk.agents.llm_agent import LlmAgent
from google.adk.evaluation.eval_metrics import (
EvalMetric,
EvalStatus,
Interval,
MetricInfo,
MetricValueInfo,
)
from google.adk.evaluation.evaluator import EvaluationResult, Evaluator, PerInvocationResult
from google.adk.evaluation.metric_evaluator_registry import DEFAULT_METRIC_EVALUATOR_REGISTRY
from google.adk.models.base_llm import BaseLlm
from google.adk.models.llm_response import LlmResponse
from google.genai import types as genai_types
METRIC_NAME = "error_rate" # lower is better; 0.10 against a 0.50 limit is a PASS
LIMIT = 0.5
VALUE = 0.1
class ErrorRateEvaluator(Evaluator):
def __init__(self, *, eval_metric: EvalMetric) -> None:
self._eval_metric = eval_metric
def evaluate_invocations(
self, actual_invocations, expected_invocations=None, conversation_scenario=None
) -> EvaluationResult:
results = [
PerInvocationResult(
actual_invocation=inv,
score=VALUE,
# The evaluator's own verdict: lower is better, 0.1 <= 0.5.
eval_status=EvalStatus.PASSED if VALUE <= LIMIT else EvalStatus.FAILED,
)
for inv in actual_invocations
]
return EvaluationResult(
overall_score=VALUE,
overall_eval_status=EvalStatus.PASSED,
per_invocation_results=results,
)
DEFAULT_METRIC_EVALUATOR_REGISTRY.register_evaluator(
metric_info=MetricInfo(
metric_name=METRIC_NAME,
description="Toy lower-is-better metric.",
metric_value_info=MetricValueInfo(interval=Interval(min_value=0.0, max_value=1.0)),
),
evaluator=ErrorRateEvaluator,
)
class _FakeLlm(BaseLlm):
model: str = "fake-model"
@classmethod
def supported_models(cls) -> list[str]:
return ["fake-model"]
async def generate_content_async(self, llm_request, stream: bool = False):
yield LlmResponse(
content=genai_types.Content(parts=[genai_types.Part(text="ok")], role="model")
)
root_agent = LlmAgent(name="toy_agent", model=_FakeLlm(), instruction="Answer briefly.")
driver.py
from __future__ import annotations
import asyncio
import json
from pathlib import Path
from google.adk.evaluation.eval_case import EvalCase, Invocation
from google.adk.evaluation.eval_set import EvalSet
from google.genai import types as genai_types
case = EvalCase(
eval_id="case_1",
conversation=[
Invocation(user_content=genai_types.Content(parts=[genai_types.Part(text="hi")], role="user"))
],
)
Path("eval_set.json").write_text(EvalSet(eval_set_id="s", eval_cases=[case]).model_dump_json())
Path("test_config.json").write_text(json.dumps({"criteria": {"error_rate": 0.5}}))
async def main() -> None:
from google.adk.evaluation.agent_evaluator import AgentEvaluator
await AgentEvaluator.evaluate(
agent_module="lowerbetter",
eval_dataset_file_path_or_dir="eval_set.json",
num_runs=1,
print_detailed_results=True,
)
asyncio.run(main())
Expected Behavior:
The evaluator's own verdict decides (0.1 against a limit of 0.5 passes), or the metric can declare its direction. A minimal fix: when the evaluator sets eval_status, use it, and fall back to overall_score >= threshold only when it doesn't.
Observed Behavior:
Summary: `EvalStatus.FAILED` for Metric: `error_rate`. Expected threshold: `0.5`, actual value: `0.1`.
| | eval_status | score | threshold | prompt | ... (remaining columns trimmed)
| 0 | EvalStatus.PASSED | 0.1 | 0.5 | hi | ...
AssertionError: Following are all the test failures.
error_rate for lowerbetter Failed. Expected 0.5, but got 0.1.
Environment Details:
- ADK:
origin/main 6be386d8939c4898aa2c0fc8a1f7cb34e7df7904; Python 3.11; Windows 11
- Also reproduces on the released google-adk 2.11.0.
🟡 Optional Information
Additional Context:
Searched issues and PRs (lower-is-better, metric polarity, threshold direction, _process_metrics_and_get_failures, "honor each metric's own eval_status"): no duplicate. PR #6739 implements a fix and can be rebased if a maintainer wants it.
How often has this issue occurred?: Always (100%)
🔴 Required Information
Describe the Bug:
AgentEvaluator._process_metrics_and_get_failuresdecides pass/fail withoverall_score >= threshold(agent_evaluator.py:901), after theINFORMATIONALbranch and without consulting the metric's owneval_status. A custom evaluator for a lower-is-better metric (error rate, latency, cost) that correctly reportsPASSEDis markedFAILEDwhenever its score is below the threshold, andevaluate()raisesAssertionError. Setting the metric toEvalStatus.INFORMATIONALavoids the false failure but also removes gating: a value over the limit (e.g. 0.9) then passes too, so a lower-is-better metric still cannot fail a run when it should.Steps to Reproduce:
python driver.pywith google-adk fromorigin/main.lowerbetter/__init__.pylowerbetter/agent.pydriver.pyExpected Behavior:
The evaluator's own verdict decides (0.1 against a limit of 0.5 passes), or the metric can declare its direction. A minimal fix: when the evaluator sets
eval_status, use it, and fall back tooverall_score >= thresholdonly when it doesn't.Observed Behavior:
Environment Details:
origin/main6be386d8939c4898aa2c0fc8a1f7cb34e7df7904; Python 3.11; Windows 11🟡 Optional Information
Additional Context:
Searched issues and PRs (lower-is-better, metric polarity, threshold direction,
_process_metrics_and_get_failures, "honor each metric's own eval_status"): no duplicate. PR #6739 implements a fix and can be rebased if a maintainer wants it.How often has this issue occurred?: Always (100%)