Repository navigation
SEP-2848: Asynchronous Approval for Tool Calls - #2848
mcguinness wants to merge 2 commits into
Conversation
f5101c1 to
8cca561
Compare
9d4fde8 to
cd425b4
Compare
|
This is a good shape for async approval because the MCP-visible part stays narrow: the client gets a task handle, while the approval system stays server-side. One invariant I would make explicit is that the task handle is not authority by itself. It should be bound to the exact original call envelope and to the policy evaluation that produced the requestable denial. The server-side record should preserve at least:
On approval, the server should re-evaluate policy against the same envelope before execution. If the tool, arguments, principal, resource, policy version, or freshness window no longer match, the task should resolve as stale/denied with no side effect. Acceptance tests I would want:
That keeps async approval durable without turning |
Draft Extensions Track SEP (experimental extension net.openid.authzen/tool-approval, per SEP-2133) that lets an MCP server gate a tool call on out-of-session approval, resolved out of band by a human, a supervising agent, a policy or risk engine, or an external system. Layered on the tasks extension (SEP-2663): a requestable denial returns a CreateTaskResult (resultType "task") in place of the tool result; the server brokers the access request server-side (OpenID AuthZEN Access Request and Approval Profile), re-evaluates against the policy decision point on approval, and executes the tool exactly once. Submission input flows through the task input channel (SEP-2322 MRTR), answerable by a human or an autonomous agent. The client carries only a server-generated taskId; all authorization artifacts stay server-side. Key semantics: - The denied-vs-executed distinction is carried in the CallToolResult body (not only _meta), with a net.openid.authzen/disposition companion, so a tasks-only client can act on it safely. - A non-tasks client may receive a degraded actionable requestable CallToolResult in addition to the -32003 path; includes a client-maturity note. - At-most-once execution; best-effort dedup scoped to originating auth context, tool, and arguments; a re-submission after a terminal denial starts a new request. - tasks/cancel before execution is honored (skip-and-deny); once execution begins the cancel is ineffective, never both executed and cancelled. - COAZ -> Access Request -> re-eval composition stated explicitly; COAZ flagged as a draft dependency. - ttlMs covers and is extended to the learned approval window; in-task input requires a connected client (collect inputs pre-submission for long gaps). - Limitations cover cross-principal and lost-taskId resumption; approval-amplification rate-limiting is MUST with an over-limit non-requestable denial; backward compatibility is additive for new tools but a behavior change for upgraded ones.
cd425b4 to
8d12c29
Compare
|
Strongly agree, and this is the invariant worth stating in normative text rather than leaving implied. The handle is a reference to a pending decision, never a grant. The SEP leaned this way in a few scattered places; per your comment I have consolidated it into one explicit rule plus a required binding record. Applied in Where it already lived:
What I added (new subsection "The task handle is not authority", plus edits to Completion and Security):
One refinement I flagged rather than adopting verbatim: policy id/version. I record it (audit, drift detection), but did not make a version change auto-resolve to stale. The SEP re-evaluates against current policy, so a policy change should produce a fresh allow-or-deny, not an automatic denial; pinning to the exact version that produced the requestable denial would also reject approvals still valid after an unrelated policy edit. So in the text: envelope fields (tool/args/principal/resource) are an exact-match binding; policy version is recorded and re-evaluation runs against current policy. If you specifically want version pinning for a class of high-assurance tools, that reads as a deployment policy on top rather than the default. Did you mean strict pinning or drift-detection? Your acceptance tests are exactly the conformance scenarios this needs, and most are MCP-observable, so I folded them into the conformance clause (now a checklist):
Thanks, this tightens the "durable handle, not bearer token" line that is the whole point of keeping the approval system server-side. |
|
On the Limitations: you give two safety conditions for a side-effecting tool — a durable consumer, or that the effect is "independently auditable." The durable-consumer half has a protocol shape (poll the task). The independently-auditable half doesn't — it's a requirement with nothing in the task model to satisfy it against. Is that intentional (left fully to deployments), or is leaving room for it in the task model in scope for this SEP? The two conditions read as parallel, but only one is actionable at the protocol level. |
|
I would split the answer in two. The audit system itself can stay deployment-owned. I do not think this SEP needs to standardize the audit record schema, storage backend, retention policy, or verifier format. But I would avoid leaving "independently auditable" as pure prose, because then the two safety conditions are not really parallel. The durable-consumer path has a protocol object to come back to. The independently-auditable path should at least have a task-visible hook that lets the final task outcome reconcile to some server-side evidence. The smallest shape I would leave room for is:
That keeps the audit evidence out of MCP core while making the requirement testable. A later verifier does not need the whole approval backend on the MCP wire, but it should be able to answer:
So my read is: leave the evidence format to deployments, but leave an explicit task-model attachment point for an opaque outcome/audit reference. Otherwise "independently auditable" is true operationally, but not actionable for implementers reading the SEP. |
|
One thing I ran into building this that I don't see in Limitations: the PEP only enforces the calls that actually go through it. The SEP says the server "is the only party that speaks the approval protocol", and the four-eyes case says the operation is "audited end to end by the approval service". Both are true of the approval path. Neither says anything about a second route to the same resource. I hit this against live Postgres. Every write action gated, all working. Then I added a second Postgres MCP server pointed at the same database, which is a normal thing for someone to do and not an attack, and the agent deleted five rows through it. Zero approvals, nothing on my audit chain. The gate held fine on its own tools. It just wasn't on the path. That runs into this bullet: "Deployments MUST ensure either a durable consumer that polls or that the effect is independently auditable." But auditability is scoped to the PEP. A second route produces an effect that no approval record ever mentions, so there's nothing for a deployment to independently audit. Feels like the requirement needs a second half: the PEP has to be the only route to whatever it's guarding, and MCP can't verify that. Worth a bullet mostly because of how quietly it fails. Nothing errors, no task gets created, and the approval log stays perfectly consistent about the calls it did see. If you're reading the audit chain you have no signal that anything went around it. Different thing, but a data point for something already decided here: keeping resolution out of band with the client limited to tasks/get is doing real work. An earlier version of mine had the approval check reachable as a tool on the same surface, and the first thing the agent did when it got "pending" back was call it itself. Not adversarial, it just polled its own pending work, which is what anyone would do if you tell them to wait. Moving the resolver off the tool surface kills that outright. Happy to share the repro if it's useful. |
|
One execution-time invariant I would make explicit is the boundary between call-envelope binding and state binding. The SEP now does a good job binding the pending task to the original tool, arguments, principal, subject/resource, approval request, and freshness window. That prevents an approval from being steered toward a different call. There is still a TOCTOU case where the call envelope is identical but the resource state that made the approval meaningful has changed while the task was pending. For example:
The bound envelope passes, but the approved transition may no longer be the transition that will execute. I would avoid requiring MCP to standardize resource versioning, but I think the SEP should state the invariant:
Implementations could satisfy that with an ETag/version, resource digest, optimistic concurrency token, transactional predicate, or a fresh policy evaluation that incorporates current state. This also gives a useful conformance boundary: envelope equality is necessary, but not always sufficient, for approval freshness. For high-consequence tools, it may also be useful for the server-side binding record to retain an opaque That keeps the SEP backend-agnostic while closing the case where the approved request is syntactically unchanged but semantically stale. |
Recast the async-approval extension as a generic MCP capability with a pluggable approval backend. AuthZEN ARAP is now a non-normative example binding, and the extension uses the io.modelcontextprotocol/* namespace. - Compose SEP-2643's denial envelope (io.modelcontextprotocol/authorization) as the portable denial classification, and define the required io.modelcontextprotocol/tool-approval-disposition contract (approved-executed, denied-not-executed, execution-error, outcome-unknown; placement and enum). - Crash-consistent creation: durable submission-intent with a fresh per-submission idempotency key, persist the handle before CreateTaskResult, reconcile orphans on restart, and fence superseded generations on chained requests so a stale callback cannot act. - Execution safety: a single atomic execution claim (at-most-once, the cancellation boundary and authorization-enforcement point), outcome-unknown when the result cannot be confirmed, and an execution deadline so a claimed task always terminalizes observably before deletion. - Before any side effect, take the caller identity and intent from the immutable call binding and re-read all dynamic authorization state (resource, risk, environment, credential, policy); re-evaluation resolves allow/retry/request/deny. - Split the immutable call binding (identity and intent, carried in requestState before the task exists) from mutable workflow and execution state, and distinguish MRTR-retry validation from tasks/update input validation. - Reserve `failed` for JSON-RPC errors, including a backend/infra failure surfaced as an internal error; it is never an authorization outcome. Map the backend terminal states normatively. - Add the lifecycle-and-status table, the Execution disposition section, and a three-clock retention model (request deadline, approval validity, MCP retention). - Security: durable sensitive-state protection, authenticated callbacks with authoritative status retrieval, and an expiry fence. - Use -32003 for Missing Required Client Capability; add SEP-2575/2567/2260 to References; align examples with SEP-2663 (nested result shape, required fields). - Editorial consolidation and tightening; add the AuthZEN ARAP example-binding section and sequence diagram. - Regenerate the docs MDX.
Friendly ReminderHi @mcguinness! This SEP proposal has been inactive for 90 days. We wanted to check in:
If this proposal is no longer being pursued, please let us know and we can close it. Otherwise, any update on the current status would be appreciated! This is an automated message from the SEP lifecycle bot. |
|
The draft already separates immutable call binding from mutable authorization state, requires fresh evaluation before execution, and distinguishes Could two fixtures share the same bound call and successful authorization decision, then diverge after the execution claim?
The audit/reconstruction assertion would be that an ALLOW decision or approval record alone cannot turn the second fixture into proof of execution. Evidence should correlate the bound call, authorization decision, execution claim, and terminal disposition, while preserving the distinction between authorization and observed outcome. This could remain an internal evidence requirement without prescribing a new public receipt format. For comparison, MCP2's receipt/reconstruction sections (§§13–14) and portable reconstruction example address reconstructing an authorization decision. That example does not prove protected execution, implement this SEP, or qualify atomic claim/crash recovery. Would this fixture pair fit the draft's conformance plan, particularly around crash recovery and durable audit consumers? AI assistance disclosure: Codex inspected the linked specifications and MCP2 sources and drafted and posted this comment at my direction. The proposed SEP-2848 fixtures have not been implemented or executed. |
Summary
Draft Extensions Track SEP introducing an experimental extension,
io.modelcontextprotocol/tool-approval, that lets an MCP server gate a tool call on anout-of-band approval without keeping the original request open. When the server's
authorization decision is a denial that is requestable (something can still approve it),
the server returns a task handle (SEP-2663 Tasks) in place of the tool result instead of
failing the call. The task stays
workingwhile the approval resolves out of band; onapproval the server re-evaluates policy and executes the tool, and the result is retrieved
via
tasks/get.The approval backend is pluggable. What approves is up to the deployment (a human
reviewer, a supervising agent, a policy or risk engine, an external ITSM/IGA system); the
approval protocol runs entirely server-side and never crosses the MCP wire; the client
carries only a server-generated
taskId. The OpenID AuthZEN Access Request and ApprovalProfile (ARAP) is included as a non-normative example binding, not a required backend.
What it builds on
io.modelcontextprotocol/authorizationenvelope is composed as the portable denialclassification.
Why this is worth a SEP
Tasks gives durable async execution but says nothing about why a call is pending or how a
decision resolves it, and no in-session primitive survives a disconnect. This SEP defines
the narrow, MCP-observable binding: a requestable denial returns a task, authorization
artifacts never cross the wire, a required
io.modelcontextprotocol/tool-approval-dispositiondistinguishes a denial (no side effect) from an execution error or an unconfirmed outcome,
and at-most-once execution under a durable claim.
What changed in this revision
Decoupled from AuthZEN (generic pluggable backend; ARAP demoted to an example binding;
io.modelcontextprotocol/*identifiers) and substantially hardened after review:persist-handle-before-return, restart reconciliation, and a generation fence so a stale
chained-request callback cannot act.
and authorization-enforcement point),
outcome-unknownwhen a result cannot be confirmed,and an execution deadline so a claimed task always terminalizes observably.
call binding, while dynamic authorization state (principal validity, resource, risk,
environment, credential, policy) is re-read at execution, not replayed.
contract, reserves
failedfor JSON-RPC errors, and adds a three-clock retention model.authoritative status retrieval, and an expiry fence.
Open questions for reviewers (Limitations section)
tasks/listand sessionsremoved); same-principal resume across restart works.
consumer or an independently auditable effect.
SEP-2643 or Tasks; the approval-specific parts stay here.
Status
draft. Working Group and Extension Maintainers are TBD, and an official-SDK referenceimplementation is a prerequisite for review (SEP-2133) that is not yet met. The MCP
Fine-Grained Authorization WG is the proposed home given the authorization focus, while the
extension builds directly on Tasks. A prototype (an MCP server fronting an approval backend,
plus a tasks-only client) and a conformance scenario are required before Final.
References