[Bug] Java SDK 自动发现持续选择已断网的 DataNode,导致查询周期性等待连接超时 #18633
Description
Activity
I'm working on this.
I checked the current master (63ae0f3) and the path described above is unchanged:
Session.getQuerySessionConnection()falls back to the default connection whenconstructSessionConnectionfails, but the failed endpoint stays in theNodesSupplierround-robin list, so every rotation pays the full connection timeout again.ShowAvailableUrlsTaskstill only skipsRemovingDataNodes, soUnknownnodes are re-added on every node list refresh.
Proposed fix (client side):
- When a query endpoint fails to connect, isolate it for a short period.
getQueryEndPoint()skips isolated endpoints and only falls back to the default connection when none are left. Once the period expires the endpoint is eligible again, so a recovered node is picked up without restarting the client; a successful connection clears the mark. Keeping this state in the node supplier lets all sessions of aSessionPoolshare it. - Make
RoundRobinPolicythread-safe, since a single supplier is shared by every session in the pool. - Unit tests: an isolated endpoint is not retried within the period, is retried after it, and the all-isolated case still falls back to the default connection.
A few questions before I open the PR:
- Should the isolation period be a new builder option, or an internal default?
- Should
SHOW AVAILABLE URLSalso excludeUnknownnodes? I'd suggest a separate PR for that, since client-side isolation is still needed when a node isRunningbut unreachable from the client. SessionConnection.reconnect()also iterates over all available nodes. Should it skip isolated endpoints as well, or keep that for a follow-up?
Reacted by Alan ZhangThanks for the analysis and the proposed fix. Here are my suggestions:
1. Make the isolation period configurable through the Builder, with a reasonable default.
The default configuration should provide endpoint isolation without requiring users to tune it, while allowing adjustments for different network environments. All sessions in the same pool should share the isolation state. Refreshing the node list should not clear an unexpired isolation period, otherwise a failed endpoint could re-enter the rotation prematurely. Once the period expires, allow a probe: clear the isolation on a successful connection, or renew it on failure.
2. Excluding
Unknownnodes fromSHOW AVAILABLE URLScan be addressed in a separate PR.This involves server-side API semantics and compatibility with other clients, so it can be discussed independently without blocking the client-side fix. Even if the server filters out
Unknownnodes, a node marked asRunningmay still be unreachable from the client. SDK-side isolation is therefore still necessary.3.
SessionConnection.reconnect()should skip endpoints whose isolation period has not expired. I suggest covering this in the same PR.Otherwise, normal query endpoint selection would avoid a failed node, but reconnection would still attempt to connect to it and incur the timeout again. Query endpoint selection and reconnection should share the same isolation state.
When all endpoints are isolated, retain a bounded recovery-probing mechanism with backoff, concurrency limits, and an overall timeout budget, rather than unconditionally retrying every failed endpoint. A controlled early probe could help detect recovery sooner, but each request should not independently retry the entire list of failed endpoints.
Best wishes.
- added 3 commits that reference this issue
on Sep 16, 2026
现象与复现
IoTDB 服务端和 Java SDK 均为 2.0.10,Table model,三节点各部署 ConfigNode/DataNode;Schema 三副本,Data 两副本,数据共识为 IoTConsensus。
SHOW CLUSTER显示 A 为Unknown、B/C 为Running。SHOW AVAILABLE URLS仍返回 A/B/C 三个地址。TableSessionPoolBuilder,初始nodeUrls只配置 B/C,设置maxSize(1)、connectionTimeoutInMs(3000),开启enableAutoFetch(true)和enableRedirection(true),串行重复执行同一条SELECT … LIMIT 5,完整读取并关闭结果集。绕过业务适配层仍能复现,12 次查询均成功返回 5 行:
节点 A 在整个对照期间保持断网。即使初始地址仅填写 B/C,自动发现仍会将 A 加回查询候选列表。这不是故障发生瞬间的一次切换等待,而是持续的周期性延迟。
关键源码
以下链接固定到 v2.0.11:对比发现,本问题涉及的关键逻辑仍保留;尚未在完整 2.0.11 集群实测复现。
1. 服务端仅排除 Removing,没有排除 Unknown。
ShowAvailableUrlsTask.buildTsBlock():2. SDK 根据发现列表轮询端点。
NodesSupplier使用以下策略;getQueryEndPoint()调用policy.chooseOne(get()):3. 连接失败后回落默认连接,但没有在该路径隔离失败端点。
Session.getQuerySessionConnection(),该方法与 2.0.10 完全一致:computeIfAbsent返回 null 时不会保存映射;方法随后回落到默认连接。故障地址仍留在轮询列表,下次选中时再次尝试连接,和实测“每三次出现一次秒级等待”吻合。期望与希望确认
当一个节点持续不可达且其他节点可完成查询时,希望 SDK 能暂时隔离失败端点并探测恢复,避免后续查询反复承担完整连接超时。
请确认:
SHOW AVAILABLE URLS包含 Unknown 是否为预期?此场景是否应由 SDK 增加故障端点退避/隔离,或调整服务端过滤规则?是否已有对应修复或推荐配置?目前临时通过关闭自动发现、重定向并只配置健康节点规避,但这需要用户手动维护节点列表。