Repository navigation
feat: retry failed MCP servers on context resync - #30489
Open
jeremyruppel wants to merge 1 commit into
Open
jeremyruppel wants to merge 1 commit into
jeremyruppel wants to merge 1 commit into
Conversation
The agent only reconnected workspace MCP servers when .mcp.json changed, so a server that was down at startup (for example, a dev server started after the agent) stayed unavailable. Resync now re-runs the MCP reload first, so `coder exp chat context refresh` and the agent's /api/v0/context/resync retry failed servers. Connected servers keep their session. Generated by Coder Agents.
jeremyruppel
marked this pull request as ready for review
October 7, 2026 22:12
jeremyruppel
added this pull request to stack #30490
October 7, 2026 22:23
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #30484.
The agent only reconnects workspace MCP servers when
.mcp.jsonchanges, so a server that is down when the agent first connects (for example, a dev server like Storybook started after the workspace is up) stays unavailable until the config file is touched.agentcontext.Manager.Resyncnow calls a newagentmcp.Manager.Reconnectbefore resolving.Reconnectre-runs the last MCP reload even when the config is unchanged, so servers that failed to connect are retried; connected servers keep their session because their config is unchanged. This makescoder exp chat context refresh(no argument) and the agent'sPOST /api/v0/context/resyncretry failed servers. Resync is only called on explicit user request today, so no background traffic is added.Not covered here: a server that restarts after it connected keeps its stale session, since its config is unchanged.
Plan and decisions
Problem
failed to connectuntil the config file changes, becausestartReloadIfNeededshort-circuits whenSnapshotChangedis false.ClientSession;classifyServerskeeps it because its config is unchanged. Out of scope for this PR.Current wiring
agent/agent.go) and the.mcp.jsonfsnotify watcher.coder exp chat context refresh(no argument) calls the agent socketResyncContext, which runsagentcontext.Manager.Resync. Before this PR, Resync re-read the MCP catalog but did not reconnect.chatd.RefreshChatContext, which is DB-only (re-pin) and never reaches the agent.Options considered
Implementation choices
Reconnectclears the stored config file snapshot and callsReload, rather than adding a force flag to the reload path.SnapshotChangedthen reports a change, the singleflight reload runs, andinstallServerswrites a fresh snapshot. False-positive reloads were already documented as safe.Follow-up (C)
Make the existing "Refresh context" button run the same agent resync as the CLI (coderd calls the agent's
/api/v0/context/resync, then re-pins), and render it when MCP server issues exist.Generated by Coder Agents on behalf of @jeremyruppel.