🔴 Required Information
Is your feature request related to a specific problem?
This is the follow-up proposed in #7035: discuss streaming-lifecycle conformance and LlmResponse field ownership separately from the resolved bug, so maintainers can weigh in on the contract before we choose an implementation.
In #7035, an after_model_callback could return a rebuilt LlmResponse without partial / turn_complete. Intermediate SSE fragments then appeared final to downstream consumers and were persisted as separate session events. #7036 fixed this by inheriting those fields when unset for both plugin and agent callbacks. That regression is fixed and shipped; this RFC is not asking to reopen it.
The systemic question remains: what must stay true as a streamed model response crosses the adapter → flow → callback → emitted event → Runner/session boundary, and who owns the fields that express those truths? Related resolved incidents involving partial tool-call consumption (#6583) and event-persistence decisions (#7184) show why a cross-boundary contract can be more useful than another isolated fix. These are historical examples, not claims of unresolved bugs.
Existing coverage matters. The flow regression tests already cover the #7036 inheritance behavior, and progressive SSE tests exercise Runner behavior. ADK also already has adk conformance test: it replays recorded SSE model responses and compares the persisted events/session against recordings. Its current run loop skips partial events rather than independently checking the complete emitted partial/final sequence. That distinction is why I think a semantic lifecycle contract is worth discussing alongside transcript comparisons—not a reason to build another CLI or claim there is no conformance infrastructure.
Describe the Solution You'd Like
I'd like maintainer guidance on a small, explicit contract with two complementary test boundaries and a corresponding response-field policy.
1. Streaming conformance: which guarantees should be invariant?
| Boundary |
Candidate semantic checks |
| Model adapter → Flow |
For providers that expose incremental SSE output, intermediate responses remain fragments; the completed response represents the assembled model output. Incomplete FunctionCall arguments must not trigger tool execution or become completed/persistable calls. Correlation IDs must survive partial-to-final assembly where relevant. |
| Callback replacement → Event → Runner |
When a plugin or agent after_model_callback rebuilds a response, an unset lifecycle field should not silently change the stream's meaning. Check both emitted partial/final events and the actual stored session, not just the transformed object. Preserve intentional overrides under the existing API contract, and avoid mutating the callback-owned replacement merely to inherit lifecycle fields. |
For a controlled one-model-call SSE fixture (e.g. two partial chunks plus one aggregate), the expected result is partial, partial, completed, with no intermediate model fragments persisted and one completed model event. This is not a global “one final event per agent invocation” rule: an invocation may contain tool loops/multiple model calls, and Live/bidirectional interactions have different completion semantics. Adapter cases should be capability-aware, not force identical chunk counts across providers.
2. LlmResponse field ownership: what does replacement mean?
Today _inherit_unset_streaming_fields treats None in a replacement as unset for partial and turn_complete, creates a copy only when inheritance is needed, and respects explicit False. That is a useful precedent but not an ownership policy for the whole response.
| Candidate field class |
Examples |
Decision needed |
| Protocol / lifecycle |
partial, turn_complete, Live interaction markers |
Which values should normally survive content replacement? Which explicit overrides remain valid, and does the rule vary by streaming mode? |
| Callback-replaceable payload |
content, callback-generated metadata |
Which fields may a callback freely replace or intentionally clear? |
| Model-call provenance / context-dependent metadata |
usage_metadata, model_version, finish_reason, grounding and error details |
Which fields describe the original model call versus the replacement result? What should be inherited, overwritten or preserved separately? |
One concrete API question: should omitting a field differ from explicitly setting it to None? The current helper checks the value, not whether the field was explicitly set. Changing that behavior could break existing callbacks, so I am asking for a policy decision—not proposing a blanket “inherit every missing field” rule.
The separate usage_metadata issue #7451 and PR #7449 already address one specific instance. This RFC is about the general ownership rule, not duplicating that implementation.
3. How should ADK express and enforce the contract?
I see three plausible paths rather than a new framework proposal:
- Focused pytest contracts: use an in-process fake model and real Runner/session service to check agent/plugin callback replacement, explicit/unset fields, emitted fragments and persisted history. This is likely the smallest useful regression slice.
- Existing conformance replay: reuse recorded model streams and existing event/session comparison where they fit; consider additional semantic checks on emitted fragments only if maintainers want those guarantees enforced through
adk conformance test.
- Documented field rules first: agree on the ownership/override matrix before extending testing to more fields or modes.
These can be combined, but I am not suggesting an implementation order before the owners of these APIs weigh in. Lightweight debug traces of an explicit lifecycle change (for example partial=True becoming False at a callback boundary) might help diagnose incidents; warnings or hard runtime enforcement require separate judgment because intentional overrides exist.
Impact on your work
I contributed #7036. The original failure surfaced as a user-visible streaming/session-history problem rather than an early contract violation. This proposal aims to make that class of boundary error easier to prevent and diagnose across callback and adapter changes, using existing test infrastructure as far as possible. No breaking API change or large refactor is requested here.
Willingness to contribute
Yes. I can provide a focused regression/test slice or a field-ownership contract draft after maintainers identify the useful boundary and intended semantics. I would prefer to agree on that direction before submitting code.
🟡 Recommended Information
Describe Alternatives You've Considered
Proposed API / Implementation
No public API change proposed yet. The deliverable I'm asking to agree on is (a) a bounded lifecycle-invariant checklist per applicable streaming mode/boundary, and (b) an ownership/override policy for LlmResponse replacements. The corresponding checks can then be landed in small PRs using existing fixtures/tools.
Additional Context / Questions for Maintainers
- Contract: Are the two boundaries above the right scope for streaming lifecycle conformance? Which guarantees are genuinely shared, and which are provider- or Live-specific?
- Ownership: Should
LlmResponse replacement distinguish framework/protocol fields from callback-owned payload and original-model provenance? In particular, is None-means-unset intended, or should omission and explicit clearing differ?
- Test home / first contribution: Would you prefer a focused Runner/pytest contract, extensions to the existing conformance replay, or an ownership note first? I can adapt the initial PR to whichever direction is useful.
Related: #7035, #7036, #7451.
🔴 Required Information
Is your feature request related to a specific problem?
This is the follow-up proposed in #7035: discuss streaming-lifecycle conformance and
LlmResponsefield ownership separately from the resolved bug, so maintainers can weigh in on the contract before we choose an implementation.In #7035, an
after_model_callbackcould return a rebuiltLlmResponsewithoutpartial/turn_complete. Intermediate SSE fragments then appeared final to downstream consumers and were persisted as separate session events. #7036 fixed this by inheriting those fields when unset for both plugin and agent callbacks. That regression is fixed and shipped; this RFC is not asking to reopen it.The systemic question remains: what must stay true as a streamed model response crosses the adapter → flow → callback → emitted event → Runner/session boundary, and who owns the fields that express those truths? Related resolved incidents involving partial tool-call consumption (#6583) and event-persistence decisions (#7184) show why a cross-boundary contract can be more useful than another isolated fix. These are historical examples, not claims of unresolved bugs.
Existing coverage matters. The flow regression tests already cover the #7036 inheritance behavior, and progressive SSE tests exercise Runner behavior. ADK also already has
adk conformance test: it replays recorded SSE model responses and compares the persisted events/session against recordings. Its current run loop skips partial events rather than independently checking the complete emitted partial/final sequence. That distinction is why I think a semantic lifecycle contract is worth discussing alongside transcript comparisons—not a reason to build another CLI or claim there is no conformance infrastructure.Describe the Solution You'd Like
I'd like maintainer guidance on a small, explicit contract with two complementary test boundaries and a corresponding response-field policy.
1. Streaming conformance: which guarantees should be invariant?
after_model_callbackrebuilds a response, an unset lifecycle field should not silently change the stream's meaning. Check both emitted partial/final events and the actual stored session, not just the transformed object. Preserve intentional overrides under the existing API contract, and avoid mutating the callback-owned replacement merely to inherit lifecycle fields.For a controlled one-model-call SSE fixture (e.g. two partial chunks plus one aggregate), the expected result is partial, partial, completed, with no intermediate model fragments persisted and one completed model event. This is not a global “one final event per agent invocation” rule: an invocation may contain tool loops/multiple model calls, and Live/bidirectional interactions have different completion semantics. Adapter cases should be capability-aware, not force identical chunk counts across providers.
2.
LlmResponsefield ownership: what does replacement mean?Today
_inherit_unset_streaming_fieldstreatsNonein a replacement as unset forpartialandturn_complete, creates a copy only when inheritance is needed, and respects explicitFalse. That is a useful precedent but not an ownership policy for the whole response.partial,turn_complete, Live interaction markerscontent, callback-generated metadatausage_metadata,model_version,finish_reason, grounding and error detailsOne concrete API question: should omitting a field differ from explicitly setting it to
None? The current helper checks the value, not whether the field was explicitly set. Changing that behavior could break existing callbacks, so I am asking for a policy decision—not proposing a blanket “inherit every missing field” rule.The separate
usage_metadataissue #7451 and PR #7449 already address one specific instance. This RFC is about the general ownership rule, not duplicating that implementation.3. How should ADK express and enforce the contract?
I see three plausible paths rather than a new framework proposal:
adk conformance test.These can be combined, but I am not suggesting an implementation order before the owners of these APIs weigh in. Lightweight debug traces of an explicit lifecycle change (for example
partial=TruebecomingFalseat a callback boundary) might help diagnose incidents; warnings or hard runtime enforcement require separate judgment because intentional overrides exist.Impact on your work
I contributed #7036. The original failure surfaced as a user-visible streaming/session-history problem rather than an early contract violation. This proposal aims to make that class of boundary error easier to prevent and diagnose across callback and adapter changes, using existing test infrastructure as far as possible. No breaking API change or large refactor is requested here.
Willingness to contribute
Yes. I can provide a focused regression/test slice or a field-ownership contract draft after maintainers identify the useful boundary and intended semantics. I would prefer to agree on that direction before submitting code.
🟡 Recommended Information
Describe Alternatives You've Considered
partialduring SSE streaming - every delta is persisted to the session as a final event #7035.Proposed API / Implementation
No public API change proposed yet. The deliverable I'm asking to agree on is (a) a bounded lifecycle-invariant checklist per applicable streaming mode/boundary, and (b) an ownership/override policy for
LlmResponsereplacements. The corresponding checks can then be landed in small PRs using existing fixtures/tools.Additional Context / Questions for Maintainers
LlmResponsereplacement distinguish framework/protocol fields from callback-owned payload and original-model provenance? In particular, isNone-means-unset intended, or should omission and explicit clearing differ?Related: #7035, #7036, #7451.