Visitar URL original
Lost-response retry on a mutating `tools/call` re-executes external effects — reproduction evidence and a spec-clarification question · Issue #3394 · modelcontextprotocol/modelcontextprotocol · GitHub
Skip to content

Lost-response retry on a mutating tools/call re-executes external effects — reproduction evidence and a spec-clarification question #3394

Description

@arjun2075

Problem-question

On the tested MCP SDK path, a mutating tools/call commits its external
effect and then loses its response; the client perceives a timeout. When the
application retries with a fresh JSON-RPC request ID, the tool executes again
and a second external effect occurs. MCP 2025-11-25 defines no normative
retry or duplicate-effect semantics for tools/call, so this is
spec-permitted, implementation-specific behavior — not a protocol defect.

Question for maintainers: "When a mutating tools/call may have completed
but its response is lost, should the MCP specification explicitly describe
the call outcome as ambiguous and warn clients that retrying can re-execute
external effects unless the tool/application provides its own stable
deduplication mechanism?"

Minimal reproduction

Mutating tool, one ledger row per execution (baseline tool is
intentionally non-idempotent — no dedup, so the ledger authoritatively
counts executions): (1) deterministic fault hook drops the attempt-1
tools/call response after the handler returned (effect committed first);
(2) client perceives a timeout via read_timeout_seconds (MCPError) — the
SDK does not automatically retry, so the retry is application policy:
exactly one retry, identical arguments, fresh JSON-RPC id, as the spec
requires (request IDs "MUST NOT have been previously used by the requestor
within the same session"); (3) count ledger rows.

Observed behavior

Attempt 1 (jsonrpc_id=2): effect committed, response dropped; client sends
notifications/cancelled for the timed-out id. Attempt 2 (jsonrpc_id=3,
application retry): tool re-executes, second effect committed. Ledger:
2 external effects for 1 logical operation, deterministic (20/20 on
clean checkouts).

Positive control

Same sequence with a stable logical-operation key and dedup at the
application effect layer: attempt 2 returns deduped, 1 external effect
recorded. Application-level idempotency successfully mitigated the
reproduced behavior. MCP did not dedup; the application effect layer did —
one mitigation, not a prescribed solution.

Specification context

  • server/tools (2025-11-25): tools/call execution semantics
    (execute-once, dedup, effect guarantees), retry after a timeout, and
    duplicate detection are all UNDEFINED — the section is normative about
    wire shapes, not about what a completed tool call did.
  • idempotentHint (since 2025-03-26, PR ToolAnnotations #185) is NON-NORMATIVE: a tool
    description, not a retry protocol; no RFC 2119 obligation either side.
  • basic/utilities/cancellation covers races only (MUST handle gracefully;
    SHOULD stop processing; MAY ignore already-completed) — not
    retry-after-loss.
  • The spec's fresh-ID requirement actively prevents ID reuse by which a
    server might otherwise correlate a retry — the retry is structurally
    unrecognizable to the server as a duplicate.

Why this matters

The tested MCP SDK path permits a retrying application to execute a
mutating tool more than once after an ambiguous lost-response outcome.
tools/call currently does not define normative retry or duplicate-effect
semantics for this case. Applications can mitigate this using stable
idempotency at the effect boundary. The experiment motivates clarifying the
expected responsibility split between client, server, tool, and downstream
effect provider, so independent implementations converge on documented
expectations rather than ad-hoc retry conventions. Related prior art
(reference only, not revival): PR #3182 (closed 2026-08-23 on policy
grounds, not merits) proposed a normative idempotencyKey; this issue asks
only whether the spec should document the responsibility boundary.

Possible documentation-conformance options

Which vehicle would maintainers prefer, if any?

  1. Non-normative spec guidance (e.g. on server/tools): the call
    outcome after a lost response is ambiguous; retrying MAY re-execute
    external effects; stable deduplication lives outside the protocol.
    Candidate text: "The protocol defines no idempotency,
    duplicate-detection, or exactly-once semantics for tools/call. A
    client that retries a tool call after a timeout or lost response MAY
    cause the tool to execute more than once; the server is not required to
    recognize a retry as a duplicate. Implementations and application layers
    that require bounded effect multiplicity should coordinate retry
    identity and duplicate suppression outside the protocol (e.g.,
    application-level idempotency keys or durable result lookup)."
  2. Tool-author guidance on declaring and handling retry-relevant
    properties.
  3. Client retry guidance for effectful calls (caution before retrying a
    timed-out mutating call).
  4. A conformance vector: lost response after effect commit → retry with
    new request ID → count effects; positive control with a stable
    application-level key.
  5. Future normative retry semantics — explicitly out of scope here.

Environment

mcp Python SDK 2.2.0 / mcp-types 2.2.0, real ClientSession →
MCPServer over memory streams, negotiated protocolVersion 2025-11-25,
Python 3.12.3.

Reproduction link-commands

Self-contained scripts (local, not posted), ~7 s per run:

cd ~/workspace/agent-continuity-conformance
.venv/bin/python upstream/mcp-b1/reproduction/b1_lost_response_repro.py
.venv/bin/python upstream/mcp-b1/reproduction/b1_positive_control.py

Baseline prints authoritative effect count: 2; the control prints 1
(see evidence-summary.md for wire excerpts and frozen evidence hashes).

Limitations

  • In-memory transport only; single deterministic fault (response dropped
    after commit); other fault shapes not tested.
  • The retry is scripted application policy, not an SDK retry policy — the
    SDK performs no automatic retry.
  • No protocol defect is claimed; generalization is limited to the tested
    path.

This is implementation evidence plus a spec-clarification question. No
normative change is proposed; no option above is prescribed.

Activity

  1. arjun2075 commented on Sep 27, 2026

    @arjun2075
    Author

    Reproduction is now published so it can be run independently: arjun2075/mcp-b1-lost-response-repro @ 6373793.

    Run: pip install -r requirements.txt (mcp==2.2.0, Python 3.12), then python <script> (~7 s each). The scripts are self-contained — only the public mcp package is imported — and the reproduction logic matches the reviewed scripts; only the run instructions were adapted for standalone use.

  2. jarrettdustinqq commented on Sep 27, 2026

    @jarrettdustinqq

    Independent reproduction at the published commit 6373793aed3c98c281d6fa695afe9fb6fa541586:

    • Clean temporary virtual environment on Python 3.13.5 with mcp==2.2.0 and mcp-types==2.2.0.
    • Ran the baseline and positive-control scripts twice each.
    • Both baseline runs negotiated 2025-11-25, used fresh tools/call IDs 2 then 3, and recorded authoritative effect count: 2.
    • Both positive-control runs used the same lost-response/retry shape and recorded authoritative effect count: 1, with attempt 2 returning deduped.

    I also reviewed the pinned scripts before execution. The relay drops only the response matching the first recorded tools/call ID; that response is observed after the handler has returned, while the single retry is explicit application policy. This independently confirms the narrow claim: on this SDK/transport path, a lost response leaves the effect outcome ambiguous and a fresh-ID application retry can execute the effect again. It does not show an automatic SDK retry or a protocol violation.

    One process point: under the current specification, I would treat option 4 as an illustrative/interoperability fixture rather than a pass/fail conformance requirement. Both duplicate execution and application-layer dedup are currently permitted, so a conformance test cannot require one effect count without first adopting normative semantics. The reproduced evidence does support non-normative guidance that the outcome is ambiguous and that callers should not infer retry safety from idempotentHint alone.

    Limitations: in-memory transport only; the same deterministic after-commit response-drop shape; Python 3.13.5 rather than the documented 3.12. No code was modified.

    AI assistance disclosure: Codex inspected and executed the pinned scripts and helped draft this report; the outputs and claim boundary were checked against the published source.

  3. ndreno commented on Sep 28, 2026

    @ndreno

    Yes, and the spec should say so plainly.

    This is the standard at-least-once ambiguity rather than an MCP defect. HTTP has it for POST: RFC 9110 allows a client to automatically retry only on connection failure, and only for idempotent methods, because a lost response is indistinguishable from a lost request. No retry policy resolves this. Only deduplication does.

    The MCP-specific part is worth stating explicitly. The spec requires a fresh id on retry, since request IDs "MUST NOT have been previously used by the requestor within the same session", so the JSON-RPC id cannot serve as a deduplication key. Today a stable key can only live inside the tool's own arguments, which means every tool author reinvents it and no client can rely on it.

    The established answer is a client-supplied idempotency key that the server persists alongside the result and replays on repeat. Stripe's Idempotency-Key is the reference implementation. The IETF draft that tried to standardise the header expired in 2025, so there is no specification to defer to.

    Two concrete suggestions:

    1. A normative note on tools/call: a lost response leaves the outcome ambiguous, and clients MUST NOT automatically retry a mutating call unless the tool documents deduplication.
    2. An optional standard field carrying an idempotency key, so deduplication stops living in tool arguments.

    For context, the gateway I maintain retries across providers on timeout. That is only safe because the upstream call has no external effect beyond billing. tools/call is exactly where that assumption breaks.

  4. arjun2075 commented on Sep 28, 2026

    @arjun2075
    Author

    Thanks — this aligns with the ambiguity the reproduction is intended to isolate. I agree that the current request ID is not a stable logical-operation identity across retries, and that the specification would benefit from making the lost-response outcome explicit.

    I’d separate the immediate clarification from the normative mechanism, though. My narrower ask here is whether tools/call should document that a completed-but-unobserved response leaves the outcome ambiguous and that retry may re-execute external effects. A standardized idempotency field and normative retry requirements seem like a useful follow-on design question, but they would require decisions about scope, persistence, replay semantics, key/request mismatch, retention, and capability negotiation.

    So I’d be comfortable starting with non-normative guidance plus an illustrative interoperability fixture, and treating standardized retry identity as separate normative work if maintainers want to pursue it.

  5. sattyamjjain commented on Sep 28, 2026

    @sattyamjjain

    +1 on starting with non-normative text.

    One thing that makes it more pressing: the newest revision already treats a retry as a fresh request. In the 2026-07-28 MRTR flow, the retry id MUST be different "as they are independent requests", and the only linking field is requestState:

    1. The JSON-RPC `id` **MUST** be different between the initial request and the retry, as they are independent requests.

    A retry after a lost response doesn't even have requestState. So there's nothing at the protocol level a server could use to tie the two calls together, only an app-level key.

    The existing pieces already point to a safe default:

    • idempotentHint defaults to false and only means something when readOnlyHint is false, and destructiveHint defaults to true
      * (This property is meaningful only when `readOnlyHint == false`)
      *
      * Default: true
      */
      destructiveHint?: boolean;
      /**
      * If true, calling the tool repeatedly with the same arguments
      * will have no additional effect on its environment.
      *
      * (This property is meaningful only when `readOnlyHint == false`)
      *
      * Default: false
      */
      idempotentHint?: boolean;
    • annotations MUST be treated as untrusted unless the server is trusted
      For trust & safety and security, clients **MUST** consider tool annotations to
      be untrusted unless they come from trusted servers.
    • a cancel after completion MAY be ignored
      1. Receivers **MAY** ignore cancellation notifications if:
      - The referenced request is unknown
      - Processing has already completed
      - The request cannot be cancelled
      1. The sender of the cancellation notification **SHOULD** ignore any response to the

    So I feel the guidance can be one sentence next to the timeout/cancellation text:
    "After a timeout, the outcome of a tools/call is unknown. Clients SHOULD NOT automatically retry a tool unless it is readOnlyHint: true, or idempotentHint: true from a trusted server, and SHOULD surface the unknown outcome instead."

    Idempotency keys can be a separate SEP, like you said.

  6. arjun2075 commented on Sep 28, 2026

    @arjun2075
    Author

    Thanks — the MRTR comparison is useful because it reinforces the distinction between request identity and logical-operation identity.
    One small wording point: if the immediate change is intended to remain non-normative, I’d avoid RFC 2119 “SHOULD NOT” language. I’d also keep idempotentHint framed as application/tool metadata rather than a protocol-level retry guarantee.
    Something like: “After a timeout or lost response, the outcome of a tools/call may be unknown. Clients should not assume that retrying is safe unless the tool/application semantics establish that repeated execution has no additional external effect.”
    That keeps the immediate clarification narrow while leaving standardized retry identity/deduplication to a separate SEP.

  7. ndreno commented on Sep 29, 2026

    @ndreno

    Agreed on the split: non-normative guidance now, retry identity as a separate SEP. I am happy to draft that SEP.

    Several of the open questions already have answers in HTTP APIs that it can start from. Stripe's documented behaviour settles three of them:

    • Key reuse with different parameters is an error, not a replay: the server compares incoming parameters to those of the original request.
    • Retention is bounded: keys can be pruned after 24 hours, and a key reused after pruning starts a new request. That caps the persistence cost and defines the replay window.
    • What is stored is the first result, failures included, once execution has started. A request rejected before execution, or conflicting with one still running, is not stored and can be retried.

    What MCP has to decide for itself is how the key travels in tools/call, whether support is negotiated as a capability, and its scope. Scope is the one with a security edge: a key must be bound to the authorization context, otherwise a colliding key replays one caller's result to another.

  8. PetrefiedThunder commented on Sep 30, 2026

    @PetrefiedThunder

    Thank you for the careful reproduction and for keeping the claim boundary narrow. The split you and others are converging on—document ambiguity now; standardize retry identity later—matches what practitioners need. One product-facing note from building middleware for mutating tools: same logical key should return the original receipt; a fresh key is a new attempt and must not be silently collapsed. Looking forward to an SEP that covers key transport, retention, parameter mismatch, and auth-context binding.

  9. modsuperagi commented on Oct 6, 2026

    @modsuperagi

    Independent reproduction related to #3394 (#3394).

    Thanks for the write-up. I tried to reproduce this independently and the result matches yours.

    Setup: a non-idempotent tools/call that records one row per execution. The response to the first call is dropped after the handler has returned, the client times out and retries once (the SDK assigns a new JSON-RPC id).

    • No ledger: 2 effects in 20/20 runs, in-memory transport and also streamable HTTP with protocol 2026-07-28 (stateless, no session id).
    • Ledger keyed by a caller-chosen key passed as a tool argument (unique constraint in Postgres): 1 effect in 20/20 runs, including 3 concurrent calls with the same key, and the same key with different arguments is rejected.
    • Limit: if the failure happens between the effect and the ledger write, the server cannot tell whether the effect happened. Check-then-act with the record written afterwards gave 2 effects; reserving the key first gave 1 effect but the retry never got a result. That case needs the external system's cooperation (lookup by key or its own idempotency key).

    This supports asking for documentation/conformance guidance rather than anything normative: tool authors who perform side effects should expose a caller-supplied key and treat the retry as a new request.

    Caveats: mcp Python SDK 2.2.0 only; the message loss is injected by my own code (transport wrapper / ASGI middleware), not a real network failure; the failure points in the last bullet are simulated in a toy server; only the no-ledger and ledger cases were run over HTTP. Code and raw outputs are reproducible with one command per scenario: https://github.com/modsuperagi/ModSuper-C.Dr-cula-SoS/tree/main/mcp-lost-response-retry

    Written with AI assistance (Claude); please re-run rather than trust this text.


    Generated by Claude Code

  10. ndreno commented on Oct 6, 2026

    @ndreno

    Thanks, the third bullet is the useful one for the SEP. It is the case a key alone cannot close: deduplication gives exactly-once only inside the server's own transaction. Once the effect leaves that boundary, the downstream system has to cooperate. The known pattern is atomic phases with recovery points (as in Brandur Leach's write-up on Stripe-style keys in Postgres): record the key and its phase in the same transaction as each local step, and pass a key derived from the caller's key to every external call so the downstream system deduplicates too.

    Your reserve-first run also shows what the protocol needs beyond the key: a distinct outcome for "key reserved, result unknown", so a retry gets an answer it can act on (retry later or reconcile) instead of hanging. I'll cover both in the SEP: key propagation as tool-author guidance, and the unknown-outcome result as part of the wire format.

  11. rossbuckley1990-hash commented on Oct 7, 2026

    @rossbuckley1990-hash

    There is another implementation pattern that may complement the idempotency-key work here: reconciliation before retry.

    Disclosure: I maintain RIGHTCLICK. We deliberately do not collapse “provider accepted the call” into “the requested outcome happened”. The runtime has separate execution/verification states, so a successful transport/provider response can remain accepted-but-unverified, and an unobservable/lost outcome can remain unknown rather than being promoted to success.

    For mutations where the resulting state is independently queryable, that gives a safer recovery path:

    mutation sent
    → response lost / outcome ambiguous
    → do NOT immediately repeat mutation
    → query an independent observable
    → if intended state is present: reconcile as success
    → if non-effect is proven: retry may be considered
    → if still unknowable: keep outcome unknown
    

    We tested the positive form with a durable-state fixture: a POST returned an object ID, but the POST response itself was not treated as proof of durable state. A separate reflected GET read the object back and only that observation established the semantic outcome:
    https://github.com/rossbuckley1990-hash/rightclick/tree/main/evidence/moat-002-durable-readback-2026-10-06

    That does not solve the hard case already identified in this thread: an external side effect that has no lookup/reconciliation surface still needs downstream idempotency/cooperation, and a lost response remains genuinely ambiguous.

    But it may be worth keeping “retry identity” and “outcome reconciliation” as two complementary protocol/application patterns:

    • idempotency key prevents duplicate execution when a retry is necessary;
    • reconciliation/read-back can sometimes prove that no retry is necessary at all.

    If a future SEP introduces an explicit unknown-outcome result, I think an optional reconciliation hint/selector (or simply guidance for tools to expose a queryable operation keyed by the mutation result/logical operation) would make that state much more actionable for clients.

  12. vladzdev commented on Oct 8, 2026

    @vladzdev

    The new request ID is what makes this particularly tricky: it gives the server no protocol-level way to know that the second call is a retry of an already-committed effect. I’d be inclined to document the failure state as explicitly indeterminate and put the dedup requirement at the logical-operation boundary rather than trying to make JSON-RPC request IDs carry that meaning. That also makes the behavior clearer for tools backed by APIs, queues, or payment systems where the actual side effect happens downstream.

  13. Ashutosh2308Bhardwaj commented on Oct 10, 2026

    @Ashutosh2308Bhardwaj

    Some ecosystem-level data that may help the SEP discussion (@ndreno):

    I checked 24 popular MCP servers (official and widely used community ones) for any way to make a mutating tools/call retry-safe:

    • 489 tools in total; 266 not marked read-only
    • 0 of those 266 take an idempotency key as an argument (I searched every schema by name and by hand)
    • GitHub's server, run for real against a scratch repo: each write made exactly one effect per call. The duplicate comes entirely from the retry: one create-issue call sent twice made two issues.

    In many cases the API underneath has no key to expose (GitHub, Slack, Jira), so this isn't servers being careless. It's the gap this issue describes, across the ecosystem. Without a protocol-level identity for the operation, servers have nothing to deduplicate on.

    Method, raw per-tool results and limits (declared vs actual behaviour; servers I couldn't scan): https://ashutosh2308bhardwaj.github.io/agentsafe/mcp-retry-scan.html

    Disclosure: I maintain agentsafe, an MCP proxy for this failure mode, which is why I looked into it.

  14. arjun2075 commented on Oct 10, 2026

    @arjun2075
    Author

    Thanks @Ashutosh2308Bhardwaj — the scan adds useful ecosystem context to the lost-response case in this issue.

    I’d keep the conclusion scoped to what was measured: no explicit idempotency-key argument was found in the scanned tools that were not marked read-only. That classification does not establish that every tool mutates state, and schema inspection alone cannot rule out other application-level retry protections.

    The GitHub example illustrates the duplicate-effect risk when a caller repeats a non-idempotent operation. Our reproduction isolates the additional uncertainty: the first effect has committed, but its response is lost, so the caller does not know whether repeating the operation is safe.

    For the SEP, a stable logical-operation identity would give cooperating implementations a way to correlate retries. The key alone would not guarantee exactly-once external effects: that still depends on how the deduplication record and effect are coordinated, and on downstream idempotency or reconciliation where the effect crosses a transaction boundary.

    This supports the split discussed above: document ambiguous outcomes and retry risks now, while addressing standardized retry identity and its semantics in a separate SEP.

  15. ryjen commented on Oct 11, 2026

    @ryjen

    One security boundary for a future retry-identity SEP that seems distinct from key/argument mismatch and cross-principal key collisions: a deduplicated result is historical execution evidence, not continuing authorization to disclose that result or invoke another effect.

    Proposed negative vector (not a verified MCP implementation result):

    1. Principal P is authorized at policy/delegation epoch A. P invokes a mutating tools/call with logical-operation key K and arguments X. The effect commits, but the response is lost.
    2. P's relevant authorization is revoked (epoch B). P retries K/X with a fresh JSON-RPC request ID.
    3. The implementation must not repeat the external effect; equally, finding a cached K/X result must not bypass the current authorization check on whether P can receive that result. A previous ALLOW/approval cannot independently authorize the retry. If current disclosure is denied, retain the historical receipt for authorized audit/reconciliation without disclosing it to P.

    A second vector isolates the identity boundary: same K/X, but a different tenant/delegation context. This must not retrieve the first caller's result or execute under the first caller's authority. Bind dedup identity to the relevant principal/tenant, tool identity, canonical request and authorization context, with explicit behavior on authority changes.

    This does not request new normative semantics in #3394 or imply the current MCP protocol promises any such guarantee. The immediate non-normative ambiguity guidance can proceed independently. For a later SEP, it would be useful to specify separately: (a) duplicate-effect suppression, (b) result visibility on retry, and (c) fresh authorization for any newly attempted effect. Failure to establish (b)/(c) must not silently convert the prior receipt into authority.

    Context: this is a proposed conformance/counterexample shape based on effect-time authority and evidence separation; I have not run it against an MCP server. No protocol-level defect is claimed.

    AI assistance disclosure: ChatGPT was used to inspect the current issue discussion, prior SEP, public contribution policy, and draft this comment.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions