Visitar URL original
SEP-2848: Asynchronous Approval for Tool Calls by mcguinness · Pull Request #2848 · modelcontextprotocol/modelcontextprotocol · GitHub
Skip to content

SEP-2848: Asynchronous Approval for Tool Calls - #2848

Open
mcguinness wants to merge 2 commits into
modelcontextprotocol:mainfrom
mcguinness:sep-async-approval-tool-calls
Open

mcguinness wants to merge 2 commits into
modelcontextprotocol:mainfrom
mcguinness:sep-async-approval-tool-calls

Conversation

@mcguinness

@mcguinness mcguinness commented Jun 3, 2026 •

Copy link
Copy Markdown

Summary

Draft Extensions Track SEP introducing an experimental extension,
io.modelcontextprotocol/tool-approval, that lets an MCP server gate a tool call on an
out-of-band approval without keeping the original request open. When the server's
authorization decision is a denial that is requestable (something can still approve it),
the server returns a task handle (SEP-2663 Tasks) in place of the tool result instead of
failing the call. The task stays working while the approval resolves out of band; on
approval the server re-evaluates policy and executes the tool, and the result is retrieved
via tasks/get.

The approval backend is pluggable. What approves is up to the deployment (a human
reviewer, a supervising agent, a policy or risk engine, an external ITSM/IGA system); the
approval protocol runs entirely server-side and never crosses the MCP wire; the client
carries only a server-generated taskId. The OpenID AuthZEN Access Request and Approval
Profile (ARAP) is included as a non-normative example binding, not a required backend.

What it builds on

  • SEP-2663 (Tasks extension) — the durable, server-directed task primitive.
  • SEP-2322 (MRTR) — the pre-task request round-trip channel for submission input.
  • SEP-2643 (Structured Authorization Denials) — its
    io.modelcontextprotocol/authorization envelope is composed as the portable denial
    classification.
  • SEP-2133 (Extensions) — extension and capability framing.
  • AuthZEN ARAP — a worked example backend binding (non-normative).

Why this is worth a SEP

Tasks gives durable async execution but says nothing about why a call is pending or how a
decision resolves it, and no in-session primitive survives a disconnect. This SEP defines
the narrow, MCP-observable binding: a requestable denial returns a task, authorization
artifacts never cross the wire, a required io.modelcontextprotocol/tool-approval-disposition
distinguishes a denial (no side effect) from an execution error or an unconfirmed outcome,
and at-most-once execution under a durable claim.

What changed in this revision

Decoupled from AuthZEN (generic pluggable backend; ARAP demoted to an example binding;
io.modelcontextprotocol/* identifiers) and substantially hardened after review:

  • Crash-consistency — durable submission-intent, a per-submission idempotency key,
    persist-handle-before-return, restart reconciliation, and a generation fence so a stale
    chained-request callback cannot act.
  • Execution safety — one atomic execution claim (at-most-once, the cancellation boundary
    and authorization-enforcement point), outcome-unknown when a result cannot be confirmed,
    and an execution deadline so a claimed task always terminalizes observably.
  • Authoritative enforcement — the caller identity and intent are pinned by an immutable
    call binding, while dynamic authorization state (principal validity, resource, risk,
    environment, credential, policy) is re-read at execution, not replayed.
  • Composition — composes SEP-2643's denial envelope, defines the disposition wire
    contract, reserves failed for JSON-RPC errors, and adds a three-clock retention model.
  • Security — durable sensitive-state protection, authenticated callbacks with
    authoritative status retrieval, and an expiry fence.

Open questions for reviewers (Limitations section)

  • Cross-principal resumption is not portable at the MCP layer (tasks/list and sessions
    removed); same-principal resume across restart works.
  • Result-consumption orchestration is out of scope; deployments must ensure a durable
    consumer or an independently auditable effect.
  • Placement of the generic disposition — the execution-outcome values could move into
    SEP-2643 or Tasks; the approval-specific parts stay here.
  • Extension vs Informational — posed for sponsor/Core Maintainer input.

Status

draft. Working Group and Extension Maintainers are TBD, and an official-SDK reference
implementation is a prerequisite for review (SEP-2133) that is not yet met. The MCP
Fine-Grained Authorization WG is the proposed home given the authorization focus, while the
extension builds directly on Tasks. A prototype (an MCP server fronting an approval backend,
plus a tasks-only client) and a conformance scenario are required before Final.

References

@mcguinness
mcguinness force-pushed the sep-async-approval-tool-calls branch 3 times, most recently from f5101c1 to 8cca561 Compare June 3, 2026 00:21
@mcguinness
mcguinness requested review from a team as code owners June 3, 2026 00:21
@mcguinness
mcguinness force-pushed the sep-async-approval-tool-calls branch 2 times, most recently from 9d4fde8 to cd425b4 Compare June 3, 2026 05:18
@localden localden added the SEP label Jun 8, 2026
@localden localden added the proposal SEP proposal without a sponsor. label Jun 8, 2026
@rpelevin

rpelevin commented Jun 8, 2026

Copy link
Copy Markdown

This is a good shape for async approval because the MCP-visible part stays narrow: the client gets a task handle, while the approval system stays server-side.

One invariant I would make explicit is that the task handle is not authority by itself. It should be bound to the exact original call envelope and to the policy evaluation that produced the requestable denial.

The server-side record should preserve at least:

  • task id;
  • tool name;
  • canonical arguments digest;
  • requester principal/session;
  • subject or resource being acted on;
  • policy id/version;
  • approval request id;
  • disposition;
  • expiry/freshness window.

On approval, the server should re-evaluate policy against the same envelope before execution. If the tool, arguments, principal, resource, policy version, or freshness window no longer match, the task should resolve as stale/denied with no side effect.

Acceptance tests I would want:

  • requestable denial returns a task without executing the tool;
  • approval resumes the call at most once;
  • denial produces a terminal disposition, not an execution error;
  • cancelled or expired task cannot later execute;
  • changed arguments or principal require a new approval;
  • result retrieval does not expose backend approval artifacts.

That keeps async approval durable without turning taskId into a bearer approval token.

Draft Extensions Track SEP (experimental extension net.openid.authzen/tool-approval,
per SEP-2133) that lets an MCP server gate a tool call on out-of-session approval,
resolved out of band by a human, a supervising agent, a policy or risk engine, or an
external system.

Layered on the tasks extension (SEP-2663): a requestable denial returns a
CreateTaskResult (resultType "task") in place of the tool result; the server brokers
the access request server-side (OpenID AuthZEN Access Request and Approval Profile),
re-evaluates against the policy decision point on approval, and executes the tool
exactly once. Submission input flows through the task input channel (SEP-2322 MRTR),
answerable by a human or an autonomous agent. The client carries only a
server-generated taskId; all authorization artifacts stay server-side.

Key semantics:
- The denied-vs-executed distinction is carried in the CallToolResult body (not only
  _meta), with a net.openid.authzen/disposition companion, so a tasks-only client can
  act on it safely.
- A non-tasks client may receive a degraded actionable requestable CallToolResult in
  addition to the -32003 path; includes a client-maturity note.
- At-most-once execution; best-effort dedup scoped to originating auth context, tool,
  and arguments; a re-submission after a terminal denial starts a new request.
- tasks/cancel before execution is honored (skip-and-deny); once execution begins the
  cancel is ineffective, never both executed and cancelled.
- COAZ -> Access Request -> re-eval composition stated explicitly; COAZ flagged as a
  draft dependency.
- ttlMs covers and is extended to the learned approval window; in-task input requires
  a connected client (collect inputs pre-submission for long gaps).
- Limitations cover cross-principal and lost-taskId resumption; approval-amplification
  rate-limiting is MUST with an over-limit non-requestable denial; backward
  compatibility is additive for new tools but a behavior change for upgraded ones.
@mcguinness
mcguinness force-pushed the sep-async-approval-tool-calls branch from cd425b4 to 8d12c29 Compare June 9, 2026 03:19
@mcguinness

Copy link
Copy Markdown
Author

Strongly agree, and this is the invariant worth stating in normative text rather than leaving implied. The handle is a reference to a pending decision, never a grant. The SEP leaned this way in a few scattered places; per your comment I have consolidated it into one explicit rule plus a required binding record. Applied in 8d12c295.

Where it already lived:

  • Completion requires a fresh re-evaluation on approval (approval is an input to a new decision, not a standing grant; PDP authoritative at enforcement), plus an execution-time freshness/credential check.
  • Security / Confused deputy binds the task to the originally evaluated subject, resource, action, and context and forbids steering execution to a different operation.
  • At-most-once + best-effort dedup, and authorization artifacts stay server-side (taskId is not a bearer token; binding material never crosses the wire).

What I added (new subsection "The task handle is not authority", plus edits to Completion and Security):

  1. An explicit invariant: the taskId is not authority. Authority is the PDP decision, re-derived at execution against the original call envelope.
  2. A normative binding record the server MUST retain and MUST check before executing: task id, tool name, canonical arguments digest, requester principal/session, subject/resource, approval request id, disposition, and expiry/freshness window (your list, adopted close to verbatim).
  3. An explicit envelope-match-or-new-approval rule: if tool, arguments, principal, or subject/resource differ from the bound envelope at execution time, the task resolves denied-not-executed with no side effect and a changed call requires a new access request. This sharpens the dedup section, which only spoke to retries.

One refinement I flagged rather than adopting verbatim: policy id/version. I record it (audit, drift detection), but did not make a version change auto-resolve to stale. The SEP re-evaluates against current policy, so a policy change should produce a fresh allow-or-deny, not an automatic denial; pinning to the exact version that produced the requestable denial would also reject approvals still valid after an unrelated policy edit. So in the text: envelope fields (tool/args/principal/resource) are an exact-match binding; policy version is recorded and re-evaluation runs against current policy. If you specifically want version pinning for a class of high-assurance tools, that reads as a deployment policy on top rather than the default. Did you mean strict pinning or drift-detection?

Your acceptance tests are exactly the conformance scenarios this needs, and most are MCP-observable, so I folded them into the conformance clause (now a checklist):

  • requestable denial returns a task without executing — observable
  • approval resumes at most once — observable
  • denial is a terminal disposition, not an execution error — observable (the denied-not-executed vs execution-error distinction, carried in the CallToolResult body)
  • cancelled or expired task cannot later execute — observable (ties to the cancel-before-execution boundary rule)
  • changed arguments or principal require a new approval — observable (the envelope-match test above)
  • result retrieval does not expose backend approval artifacts — observable

Thanks, this tightens the "durable handle, not bearer token" line that is the whole point of keeping the approval system server-side.

@AgentGymLeader

Copy link
Copy Markdown

On the Limitations: you give two safety conditions for a side-effecting tool — a durable consumer, or that the effect is "independently auditable." The durable-consumer half has a protocol shape (poll the task). The independently-auditable half doesn't — it's a requirement with nothing in the task model to satisfy it against.

Is that intentional (left fully to deployments), or is leaving room for it in the task model in scope for this SEP? The two conditions read as parallel, but only one is actionable at the protocol level.

@rpelevin

rpelevin commented Jun 9, 2026

Copy link
Copy Markdown

I would split the answer in two.

The audit system itself can stay deployment-owned. I do not think this SEP needs to standardize the audit record schema, storage backend, retention policy, or verifier format.

But I would avoid leaving "independently auditable" as pure prose, because then the two safety conditions are not really parallel. The durable-consumer path has a protocol object to come back to. The independently-auditable path should at least have a task-visible hook that lets the final task outcome reconcile to some server-side evidence.

The smallest shape I would leave room for is:

  • the task is still bound to the original call envelope: task id, tool, arguments digest, requester/subject/resource, policy context, and approval request;
  • the terminal task result still distinguishes approved-executed, denied-not-executed, execution-error, and cancelled;
  • if the side effect can complete without a live consumer, the server records an outcome entry keyed by the task id / decision id / original call digest;
  • tasks/get can surface an opaque outcome or audit reference when the deployment exposes one;
  • a deployment with neither a durable consumer nor a later-verifiable outcome record should be called out as unsupported for side-effecting tools, not merely discouraged.

That keeps the audit evidence out of MCP core while making the requirement testable. A later verifier does not need the whole approval backend on the MCP wire, but it should be able to answer:

  1. did the tool execute or not;
  2. which bound call envelope did that outcome apply to;
  3. was the approval/denial/cancel terminal before execution;
  4. can a retry or lost consumer cause a second mutation.

So my read is: leave the evidence format to deployments, but leave an explicit task-model attachment point for an opaque outcome/audit reference. Otherwise "independently auditable" is true operationally, but not actionable for implementers reading the SEP.

@localden localden changed the title Asynchronous Approval for Tool Calls (Extensions Track SEP) SEP-2848: Asynchronous Approval for Tool Calls Jul 29, 2026
@KimHyeongRae0

Copy link
Copy Markdown

One thing I ran into building this that I don't see in Limitations: the PEP only enforces the calls that actually go through it.

The SEP says the server "is the only party that speaks the approval protocol", and the four-eyes case says the operation is "audited end to end by the approval service". Both are true of the approval path. Neither says anything about a second route to the same resource.

I hit this against live Postgres. Every write action gated, all working. Then I added a second Postgres MCP server pointed at the same database, which is a normal thing for someone to do and not an attack, and the agent deleted five rows through it. Zero approvals, nothing on my audit chain. The gate held fine on its own tools. It just wasn't on the path.

That runs into this bullet: "Deployments MUST ensure either a durable consumer that polls or that the effect is independently auditable." But auditability is scoped to the PEP. A second route produces an effect that no approval record ever mentions, so there's nothing for a deployment to independently audit. Feels like the requirement needs a second half: the PEP has to be the only route to whatever it's guarding, and MCP can't verify that.

Worth a bullet mostly because of how quietly it fails. Nothing errors, no task gets created, and the approval log stays perfectly consistent about the calls it did see. If you're reading the audit chain you have no signal that anything went around it.

Different thing, but a data point for something already decided here: keeping resolution out of band with the client limited to tasks/get is doing real work. An earlier version of mine had the approval check reachable as a tool on the same surface, and the first thing the agent did when it got "pending" back was call it itself. Not adversarial, it just polled its own pending work, which is what anyone would do if you tell them to wait. Moving the resolver off the tool surface kills that outright.

Happy to share the repro if it's useful.

@AasthaPJoshi

Copy link
Copy Markdown

One execution-time invariant I would make explicit is the boundary between call-envelope binding and state binding.

The SEP now does a good job binding the pending task to the original tool, arguments, principal, subject/resource, approval request, and freshness window. That prevents an approval from being steered toward a different call.

There is still a TOCTOU case where the call envelope is identical but the resource state that made the approval meaningful has changed while the task was pending.

For example:

approve update(resource=R, args=A)
R changes from version V1 to V2
approval arrives
tool=update, resource=R, args=A still match exactly

The bound envelope passes, but the approved transition may no longer be the transition that will execute.

I would avoid requiring MCP to standardize resource versioning, but I think the SEP should state the invariant:

If authorization or approval depended on mutable external state, execution MUST revalidate the relevant state assumptions before performing the side effect.

Implementations could satisfy that with an ETag/version, resource digest, optimistic concurrency token, transactional predicate, or a fresh policy evaluation that incorporates current state.

This also gives a useful conformance boundary: envelope equality is necessary, but not always sufficient, for approval freshness.

For high-consequence tools, it may also be useful for the server-side binding record to retain an opaque state_precondition or version reference when one exists, without exposing it on the MCP wire.

That keeps the SEP backend-agnostic while closing the case where the approved request is syntactically unchanged but semantically stale.

Recast the async-approval extension as a generic MCP capability with a
pluggable approval backend. AuthZEN ARAP is now a non-normative example
binding, and the extension uses the io.modelcontextprotocol/* namespace.

- Compose SEP-2643's denial envelope (io.modelcontextprotocol/authorization)
  as the portable denial classification, and define the required
  io.modelcontextprotocol/tool-approval-disposition contract (approved-executed,
  denied-not-executed, execution-error, outcome-unknown; placement and enum).
- Crash-consistent creation: durable submission-intent with a fresh
  per-submission idempotency key, persist the handle before CreateTaskResult,
  reconcile orphans on restart, and fence superseded generations on chained
  requests so a stale callback cannot act.
- Execution safety: a single atomic execution claim (at-most-once, the
  cancellation boundary and authorization-enforcement point), outcome-unknown
  when the result cannot be confirmed, and an execution deadline so a claimed
  task always terminalizes observably before deletion.
- Before any side effect, take the caller identity and intent from the immutable
  call binding and re-read all dynamic authorization state (resource, risk,
  environment, credential, policy); re-evaluation resolves allow/retry/request/deny.
- Split the immutable call binding (identity and intent, carried in requestState
  before the task exists) from mutable workflow and execution state, and
  distinguish MRTR-retry validation from tasks/update input validation.
- Reserve `failed` for JSON-RPC errors, including a backend/infra failure
  surfaced as an internal error; it is never an authorization outcome. Map the
  backend terminal states normatively.
- Add the lifecycle-and-status table, the Execution disposition section, and a
  three-clock retention model (request deadline, approval validity, MCP retention).
- Security: durable sensitive-state protection, authenticated callbacks with
  authoritative status retrieval, and an expiry fence.
- Use -32003 for Missing Required Client Capability; add SEP-2575/2567/2260 to
  References; align examples with SEP-2663 (nested result shape, required fields).
- Editorial consolidation and tightening; add the AuthZEN ARAP example-binding
  section and sequence diagram.
- Regenerate the docs MDX.
@sep-automation-bot

Copy link
Copy Markdown

Friendly Reminder

Hi @mcguinness!

This SEP proposal has been inactive for 90 days.

We wanted to check in:

  • Are you still working on this proposal?
  • Is there anything blocking progress?
  • Do you need help finding a sponsor?

If this proposal is no longer being pursued, please let us know and we can close it. Otherwise, any update on the current status would be appreciated!


This is an automated message from the SEP lifecycle bot.

Copy link
Copy Markdown

The draft already separates immutable call binding from mutable authorization state, requires fresh evaluation before execution, and distinguishes outcome-unknown from a known non-execution. Building on those rules and the resource-state revalidation discussion above, I would like to suggest an evidence-focused conformance case.

Could two fixtures share the same bound call and successful authorization decision, then diverge after the execution claim?

  1. A durable completion record supports approved-executed.
  2. A crash leaves execution uncertain, so recovery reports outcome-unknown and does not reinvoke.

The audit/reconstruction assertion would be that an ALLOW decision or approval record alone cannot turn the second fixture into proof of execution. Evidence should correlate the bound call, authorization decision, execution claim, and terminal disposition, while preserving the distinction between authorization and observed outcome. This could remain an internal evidence requirement without prescribing a new public receipt format.

For comparison, MCP2's receipt/reconstruction sections (§§13–14) and portable reconstruction example address reconstructing an authorization decision. That example does not prove protected execution, implement this SEP, or qualify atomic claim/crash recovery.

Would this fixture pair fit the draft's conformance plan, particularly around crash recovery and durable audit consumers?

AI assistance disclosure: Codex inspected the linked specifications and MCP2 sources and drafted and posted this comment at my direction. The proposed SEP-2848 fixtures have not been implemented or executed.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

proposal SEP proposal without a sponsor. SEP

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

7 participants