Repository navigation
Lost-response retry on a mutating tools/call re-executes external effects — reproduction evidence and a spec-clarification question #3394
Description
Activity
- added a commit that references this issue
on Sep 27, 2026 Reproduction is now published so it can be run independently: arjun2075/mcp-b1-lost-response-repro @
6373793.- Baseline: b1_lost_response_repro.py — prints
authoritative effect count: 2 - Positive control: b1_positive_control.py — prints
authoritative effect count: 1
Run:
pip install -r requirements.txt(mcp==2.2.0, Python 3.12), thenpython <script>(~7 s each). The scripts are self-contained — only the publicmcppackage is imported — and the reproduction logic matches the reviewed scripts; only the run instructions were adapted for standalone use.- Baseline: b1_lost_response_repro.py — prints
jarrettdustinqq commented
on Sep 27, 2026 More actionsIndependent reproduction at the published commit
6373793aed3c98c281d6fa695afe9fb6fa541586:- Clean temporary virtual environment on Python 3.13.5 with
mcp==2.2.0andmcp-types==2.2.0. - Ran the baseline and positive-control scripts twice each.
- Both baseline runs negotiated
2025-11-25, used freshtools/callIDs 2 then 3, and recordedauthoritative effect count: 2. - Both positive-control runs used the same lost-response/retry shape and recorded
authoritative effect count: 1, with attempt 2 returningdeduped.
I also reviewed the pinned scripts before execution. The relay drops only the response matching the first recorded
tools/callID; that response is observed after the handler has returned, while the single retry is explicit application policy. This independently confirms the narrow claim: on this SDK/transport path, a lost response leaves the effect outcome ambiguous and a fresh-ID application retry can execute the effect again. It does not show an automatic SDK retry or a protocol violation.One process point: under the current specification, I would treat option 4 as an illustrative/interoperability fixture rather than a pass/fail conformance requirement. Both duplicate execution and application-layer dedup are currently permitted, so a conformance test cannot require one effect count without first adopting normative semantics. The reproduced evidence does support non-normative guidance that the outcome is ambiguous and that callers should not infer retry safety from
idempotentHintalone.Limitations: in-memory transport only; the same deterministic after-commit response-drop shape; Python 3.13.5 rather than the documented 3.12. No code was modified.
AI assistance disclosure: Codex inspected and executed the pinned scripts and helped draft this report; the outputs and claim boundary were checked against the published source.
- Clean temporary virtual environment on Python 3.13.5 with
Yes, and the spec should say so plainly.
This is the standard at-least-once ambiguity rather than an MCP defect. HTTP has it for POST: RFC 9110 allows a client to automatically retry only on connection failure, and only for idempotent methods, because a lost response is indistinguishable from a lost request. No retry policy resolves this. Only deduplication does.
The MCP-specific part is worth stating explicitly. The spec requires a fresh id on retry, since request IDs "MUST NOT have been previously used by the requestor within the same session", so the JSON-RPC id cannot serve as a deduplication key. Today a stable key can only live inside the tool's own arguments, which means every tool author reinvents it and no client can rely on it.
The established answer is a client-supplied idempotency key that the server persists alongside the result and replays on repeat. Stripe's
Idempotency-Keyis the reference implementation. The IETF draft that tried to standardise the header expired in 2025, so there is no specification to defer to.Two concrete suggestions:
- A normative note on
tools/call: a lost response leaves the outcome ambiguous, and clients MUST NOT automatically retry a mutating call unless the tool documents deduplication. - An optional standard field carrying an idempotency key, so deduplication stops living in tool arguments.
For context, the gateway I maintain retries across providers on timeout. That is only safe because the upstream call has no external effect beyond billing.
tools/callis exactly where that assumption breaks.- A normative note on
Thanks — this aligns with the ambiguity the reproduction is intended to isolate. I agree that the current request ID is not a stable logical-operation identity across retries, and that the specification would benefit from making the lost-response outcome explicit.
I’d separate the immediate clarification from the normative mechanism, though. My narrower ask here is whether tools/call should document that a completed-but-unobserved response leaves the outcome ambiguous and that retry may re-execute external effects. A standardized idempotency field and normative retry requirements seem like a useful follow-on design question, but they would require decisions about scope, persistence, replay semantics, key/request mismatch, retention, and capability negotiation.
So I’d be comfortable starting with non-normative guidance plus an illustrative interoperability fixture, and treating standardized retry identity as separate normative work if maintainers want to pursue it.
+1 on starting with non-normative text.
One thing that makes it more pressing: the newest revision already treats a retry as a fresh request. In the 2026-07-28 MRTR flow, the retry id MUST be different "as they are independent requests", and the only linking field is requestState:
1. The JSON-RPC `id` **MUST** be different between the initial request and the retry, as they are independent requests.
A retry after a lost response doesn't even have requestState. So there's nothing at the protocol level a server could use to tie the two calls together, only an app-level key.The existing pieces already point to a safe default:
- idempotentHint defaults to false and only means something when readOnlyHint is false, and destructiveHint defaults to true
modelcontextprotocol/schema/2025-11-25/schema.ts
Lines 1197 to 1211 in ab3a39c
* (This property is meaningful only when `readOnlyHint == false`) * * Default: true */ destructiveHint?: boolean; /** * If true, calling the tool repeatedly with the same arguments * will have no additional effect on its environment. * * (This property is meaningful only when `readOnlyHint == false`) * * Default: false */ idempotentHint?: boolean; - annotations MUST be treated as untrusted unless the server is trusted
modelcontextprotocol/docs/specification/2025-11-25/server/tools.mdx
Lines 213 to 214 in ab3a39c
For trust & safety and security, clients **MUST** consider tool annotations to be untrusted unless they come from trusted servers. - a cancel after completion MAY be ignored
modelcontextprotocol/docs/specification/2025-11-25/basic/utilities/cancellation.mdx
Lines 41 to 45 in ab3a39c
1. Receivers **MAY** ignore cancellation notifications if: - The referenced request is unknown - Processing has already completed - The request cannot be cancelled 1. The sender of the cancellation notification **SHOULD** ignore any response to the
So I feel the guidance can be one sentence next to the timeout/cancellation text:
"After a timeout, the outcome of a tools/call is unknown. Clients SHOULD NOT automatically retry a tool unless it is readOnlyHint: true, or idempotentHint: true from a trusted server, and SHOULD surface the unknown outcome instead."Idempotency keys can be a separate SEP, like you said.
- idempotentHint defaults to false and only means something when readOnlyHint is false, and destructiveHint defaults to true
Thanks — the MRTR comparison is useful because it reinforces the distinction between request identity and logical-operation identity.
One small wording point: if the immediate change is intended to remain non-normative, I’d avoid RFC 2119 “SHOULD NOT” language. I’d also keep idempotentHint framed as application/tool metadata rather than a protocol-level retry guarantee.
Something like: “After a timeout or lost response, the outcome of a tools/call may be unknown. Clients should not assume that retrying is safe unless the tool/application semantics establish that repeated execution has no additional external effect.”
That keeps the immediate clarification narrow while leaving standardized retry identity/deduplication to a separate SEP.Agreed on the split: non-normative guidance now, retry identity as a separate SEP. I am happy to draft that SEP.
Several of the open questions already have answers in HTTP APIs that it can start from. Stripe's documented behaviour settles three of them:
- Key reuse with different parameters is an error, not a replay: the server compares incoming parameters to those of the original request.
- Retention is bounded: keys can be pruned after 24 hours, and a key reused after pruning starts a new request. That caps the persistence cost and defines the replay window.
- What is stored is the first result, failures included, once execution has started. A request rejected before execution, or conflicting with one still running, is not stored and can be retried.
What MCP has to decide for itself is how the key travels in
tools/call, whether support is negotiated as a capability, and its scope. Scope is the one with a security edge: a key must be bound to the authorization context, otherwise a colliding key replays one caller's result to another.Thank you for the careful reproduction and for keeping the claim boundary narrow. The split you and others are converging on—document ambiguity now; standardize retry identity later—matches what practitioners need. One product-facing note from building middleware for mutating tools: same logical key should return the original receipt; a fresh key is a new attempt and must not be silently collapsed. Looking forward to an SEP that covers key transport, retention, parameter mismatch, and auth-context binding.
Independent reproduction related to #3394 (#3394).
Thanks for the write-up. I tried to reproduce this independently and the result matches yours.
Setup: a non-idempotent
tools/callthat records one row per execution. The response to the first call is dropped after the handler has returned, the client times out and retries once (the SDK assigns a new JSON-RPC id).- No ledger: 2 effects in 20/20 runs, in-memory transport and also streamable HTTP with protocol 2026-07-28 (stateless, no session id).
- Ledger keyed by a caller-chosen key passed as a tool argument (unique constraint in Postgres): 1 effect in 20/20 runs, including 3 concurrent calls with the same key, and the same key with different arguments is rejected.
- Limit: if the failure happens between the effect and the ledger write, the server cannot tell whether the effect happened. Check-then-act with the record written afterwards gave 2 effects; reserving the key first gave 1 effect but the retry never got a result. That case needs the external system's cooperation (lookup by key or its own idempotency key).
This supports asking for documentation/conformance guidance rather than anything normative: tool authors who perform side effects should expose a caller-supplied key and treat the retry as a new request.
Caveats: mcp Python SDK 2.2.0 only; the message loss is injected by my own code (transport wrapper / ASGI middleware), not a real network failure; the failure points in the last bullet are simulated in a toy server; only the no-ledger and ledger cases were run over HTTP. Code and raw outputs are reproducible with one command per scenario: https://github.com/modsuperagi/ModSuper-C.Dr-cula-SoS/tree/main/mcp-lost-response-retry
Written with AI assistance (Claude); please re-run rather than trust this text.
Generated by Claude Code
Reacted by modsuperagiThanks, the third bullet is the useful one for the SEP. It is the case a key alone cannot close: deduplication gives exactly-once only inside the server's own transaction. Once the effect leaves that boundary, the downstream system has to cooperate. The known pattern is atomic phases with recovery points (as in Brandur Leach's write-up on Stripe-style keys in Postgres): record the key and its phase in the same transaction as each local step, and pass a key derived from the caller's key to every external call so the downstream system deduplicates too.
Your reserve-first run also shows what the protocol needs beyond the key: a distinct outcome for "key reserved, result unknown", so a retry gets an answer it can act on (retry later or reconcile) instead of hanging. I'll cover both in the SEP: key propagation as tool-author guidance, and the unknown-outcome result as part of the wire format.
rossbuckley1990-hash commented
on Oct 7, 2026 More actionsThere is another implementation pattern that may complement the idempotency-key work here: reconciliation before retry.
Disclosure: I maintain RIGHTCLICK. We deliberately do not collapse “provider accepted the call” into “the requested outcome happened”. The runtime has separate execution/verification states, so a successful transport/provider response can remain accepted-but-unverified, and an unobservable/lost outcome can remain unknown rather than being promoted to success.
For mutations where the resulting state is independently queryable, that gives a safer recovery path:
mutation sent → response lost / outcome ambiguous → do NOT immediately repeat mutation → query an independent observable → if intended state is present: reconcile as success → if non-effect is proven: retry may be considered → if still unknowable: keep outcome unknownWe tested the positive form with a durable-state fixture: a POST returned an object ID, but the POST response itself was not treated as proof of durable state. A separate reflected GET read the object back and only that observation established the semantic outcome:
https://github.com/rossbuckley1990-hash/rightclick/tree/main/evidence/moat-002-durable-readback-2026-10-06That does not solve the hard case already identified in this thread: an external side effect that has no lookup/reconciliation surface still needs downstream idempotency/cooperation, and a lost response remains genuinely ambiguous.
But it may be worth keeping “retry identity” and “outcome reconciliation” as two complementary protocol/application patterns:
- idempotency key prevents duplicate execution when a retry is necessary;
- reconciliation/read-back can sometimes prove that no retry is necessary at all.
If a future SEP introduces an explicit unknown-outcome result, I think an optional reconciliation hint/selector (or simply guidance for tools to expose a queryable operation keyed by the mutation result/logical operation) would make that state much more actionable for clients.
The new request ID is what makes this particularly tricky: it gives the server no protocol-level way to know that the second call is a retry of an already-committed effect. I’d be inclined to document the failure state as explicitly indeterminate and put the dedup requirement at the logical-operation boundary rather than trying to make JSON-RPC request IDs carry that meaning. That also makes the behavior clearer for tools backed by APIs, queues, or payment systems where the actual side effect happens downstream.
Some ecosystem-level data that may help the SEP discussion (@ndreno):
I checked 24 popular MCP servers (official and widely used community ones) for any way to make a mutating
tools/callretry-safe:- 489 tools in total; 266 not marked read-only
- 0 of those 266 take an idempotency key as an argument (I searched every schema by name and by hand)
- GitHub's server, run for real against a scratch repo: each write made exactly one effect per call. The duplicate comes entirely from the retry: one create-issue call sent twice made two issues.
In many cases the API underneath has no key to expose (GitHub, Slack, Jira), so this isn't servers being careless. It's the gap this issue describes, across the ecosystem. Without a protocol-level identity for the operation, servers have nothing to deduplicate on.
Method, raw per-tool results and limits (declared vs actual behaviour; servers I couldn't scan): https://ashutosh2308bhardwaj.github.io/agentsafe/mcp-retry-scan.html
Disclosure: I maintain agentsafe, an MCP proxy for this failure mode, which is why I looked into it.
Thanks @Ashutosh2308Bhardwaj — the scan adds useful ecosystem context to the lost-response case in this issue.
I’d keep the conclusion scoped to what was measured: no explicit idempotency-key argument was found in the scanned tools that were not marked read-only. That classification does not establish that every tool mutates state, and schema inspection alone cannot rule out other application-level retry protections.
The GitHub example illustrates the duplicate-effect risk when a caller repeats a non-idempotent operation. Our reproduction isolates the additional uncertainty: the first effect has committed, but its response is lost, so the caller does not know whether repeating the operation is safe.
For the SEP, a stable logical-operation identity would give cooperating implementations a way to correlate retries. The key alone would not guarantee exactly-once external effects: that still depends on how the deduplication record and effect are coordinated, and on downstream idempotency or reconciliation where the effect crosses a transaction boundary.
This supports the split discussed above: document ambiguous outcomes and retry risks now, while addressing standardized retry identity and its semantics in a separate SEP.
One security boundary for a future retry-identity SEP that seems distinct from key/argument mismatch and cross-principal key collisions: a deduplicated result is historical execution evidence, not continuing authorization to disclose that result or invoke another effect.
Proposed negative vector (not a verified MCP implementation result):
- Principal P is authorized at policy/delegation epoch A. P invokes a mutating
tools/callwith logical-operation key K and arguments X. The effect commits, but the response is lost. - P's relevant authorization is revoked (epoch B). P retries K/X with a fresh JSON-RPC request ID.
- The implementation must not repeat the external effect; equally, finding a cached K/X result must not bypass the current authorization check on whether P can receive that result. A previous ALLOW/approval cannot independently authorize the retry. If current disclosure is denied, retain the historical receipt for authorized audit/reconciliation without disclosing it to P.
A second vector isolates the identity boundary: same K/X, but a different tenant/delegation context. This must not retrieve the first caller's result or execute under the first caller's authority. Bind dedup identity to the relevant principal/tenant, tool identity, canonical request and authorization context, with explicit behavior on authority changes.
This does not request new normative semantics in #3394 or imply the current MCP protocol promises any such guarantee. The immediate non-normative ambiguity guidance can proceed independently. For a later SEP, it would be useful to specify separately: (a) duplicate-effect suppression, (b) result visibility on retry, and (c) fresh authorization for any newly attempted effect. Failure to establish (b)/(c) must not silently convert the prior receipt into authority.
Context: this is a proposed conformance/counterexample shape based on effect-time authority and evidence separation; I have not run it against an MCP server. No protocol-level defect is claimed.
AI assistance disclosure: ChatGPT was used to inspect the current issue discussion, prior SEP, public contribution policy, and draft this comment.
- Principal P is authorized at policy/delegation epoch A. P invokes a mutating
Problem-question
On the tested MCP SDK path, a mutating
tools/callcommits its externaleffect and then loses its response; the client perceives a timeout. When the
application retries with a fresh JSON-RPC request ID, the tool executes again
and a second external effect occurs. MCP 2025-11-25 defines no normative
retry or duplicate-effect semantics for
tools/call, so this isspec-permitted, implementation-specific behavior — not a protocol defect.
Question for maintainers: "When a mutating tools/call may have completed
but its response is lost, should the MCP specification explicitly describe
the call outcome as ambiguous and warn clients that retrying can re-execute
external effects unless the tool/application provides its own stable
deduplication mechanism?"
Minimal reproduction
Mutating tool, one ledger row per execution (baseline tool is
intentionally non-idempotent — no dedup, so the ledger authoritatively
counts executions): (1) deterministic fault hook drops the attempt-1
tools/callresponse after the handler returned (effect committed first);(2) client perceives a timeout via
read_timeout_seconds(MCPError) — theSDK does not automatically retry, so the retry is application policy:
exactly one retry, identical arguments, fresh JSON-RPC id, as the spec
requires (request IDs "MUST NOT have been previously used by the requestor
within the same session"); (3) count ledger rows.
Observed behavior
Attempt 1 (
jsonrpc_id=2): effect committed, response dropped; client sendsnotifications/cancelledfor the timed-out id. Attempt 2 (jsonrpc_id=3,application retry): tool re-executes, second effect committed. Ledger:
2 external effects for 1 logical operation, deterministic (20/20 on
clean checkouts).
Positive control
Same sequence with a stable logical-operation key and dedup at the
application effect layer: attempt 2 returns
deduped, 1 external effectrecorded. Application-level idempotency successfully mitigated the
reproduced behavior. MCP did not dedup; the application effect layer did —
one mitigation, not a prescribed solution.
Specification context
server/tools(2025-11-25):tools/callexecution semantics(execute-once, dedup, effect guarantees), retry after a timeout, and
duplicate detection are all UNDEFINED — the section is normative about
wire shapes, not about what a completed tool call did.
idempotentHint(since 2025-03-26, PR ToolAnnotations #185) is NON-NORMATIVE: a tooldescription, not a retry protocol; no RFC 2119 obligation either side.
basic/utilities/cancellationcovers races only (MUST handle gracefully;SHOULD stop processing; MAY ignore already-completed) — not
retry-after-loss.
server might otherwise correlate a retry — the retry is structurally
unrecognizable to the server as a duplicate.
Why this matters
The tested MCP SDK path permits a retrying application to execute a
mutating tool more than once after an ambiguous lost-response outcome.
tools/callcurrently does not define normative retry or duplicate-effectsemantics for this case. Applications can mitigate this using stable
idempotency at the effect boundary. The experiment motivates clarifying the
expected responsibility split between client, server, tool, and downstream
effect provider, so independent implementations converge on documented
expectations rather than ad-hoc retry conventions. Related prior art
(reference only, not revival): PR #3182 (closed 2026-08-23 on policy
grounds, not merits) proposed a normative
idempotencyKey; this issue asksonly whether the spec should document the responsibility boundary.
Possible documentation-conformance options
Which vehicle would maintainers prefer, if any?
server/tools): the calloutcome after a lost response is ambiguous; retrying MAY re-execute
external effects; stable deduplication lives outside the protocol.
Candidate text: "The protocol defines no idempotency,
duplicate-detection, or exactly-once semantics for
tools/call. Aclient that retries a tool call after a timeout or lost response MAY
cause the tool to execute more than once; the server is not required to
recognize a retry as a duplicate. Implementations and application layers
that require bounded effect multiplicity should coordinate retry
identity and duplicate suppression outside the protocol (e.g.,
application-level idempotency keys or durable result lookup)."
properties.
timed-out mutating call).
new request ID → count effects; positive control with a stable
application-level key.
Environment
mcpPython SDK 2.2.0 /mcp-types2.2.0, realClientSession→MCPServerover memory streams, negotiated protocolVersion2025-11-25,Python 3.12.3.
Reproduction link-commands
Self-contained scripts (local, not posted), ~7 s per run:
Baseline prints
authoritative effect count: 2; the control prints1(see
evidence-summary.mdfor wire excerpts and frozen evidence hashes).Limitations
after commit); other fault shapes not tested.
SDK performs no automatic retry.
path.
This is implementation evidence plus a spec-clarification question. No
normative change is proposed; no option above is prescribed.