Visitar URL original
[Bug] Java SDK 自动发现持续选择已断网的 DataNode,导致查询周期性等待连接超时 · Issue #18633 · apache/iotdb · GitHub
Skip to content

[Bug] Java SDK 自动发现持续选择已断网的 DataNode,导致查询周期性等待连接超时 #18633

Description

@AlanDevise

现象与复现

IoTDB 服务端和 Java SDK 均为 2.0.10,Table model,三节点各部署 ConfigNode/DataNode;Schema 三副本,Data 两副本,数据共识为 IoTConsensus。

  1. 将节点 A 断网,保持 B/C 正常。等待 SHOW CLUSTER 显示 A 为 Unknown、B/C 为 Running。
  2. 此时 SHOW AVAILABLE URLS 仍返回 A/B/C 三个地址。
  3. 使用原生 TableSessionPoolBuilder,初始 nodeUrls 只配置 B/C,设置 maxSize(1)、connectionTimeoutInMs(3000),开启 enableAutoFetch(true) 和 enableRedirection(true),串行重复执行同一条 SELECT … LIMIT 5,完整读取并关闭结果集。

绕过业务适配层仍能复现,12 次查询均成功返回 5 行:

设置 12 次查询结果
自动发现、重定向均开启 第 1/4/7/10 次分别约 2423/2976/3029/3028 ms,其余约 22–29 ms
两项均关闭,其他配置及 SQL 不变 全部约 21–30 ms

节点 A 在整个对照期间保持断网。即使初始地址仅填写 B/C,自动发现仍会将 A 加回查询候选列表。这不是故障发生瞬间的一次切换等待,而是持续的周期性延迟。

关键源码

以下链接固定到 v2.0.11:对比发现,本问题涉及的关键逻辑仍保留;尚未在完整 2.0.11 集群实测复现。

1. 服务端仅排除 Removing,没有排除 Unknown。

ShowAvailableUrlsTask.buildTsBlock():

for (TDataNodeInfo dataNodeInfo : showDataNodesResp.getDataNodesInfoList()) {
  String status = dataNodeInfo.getStatus();
  if (RegionStatus.Removing.getStatus().equals(status)) {
    continue;
  }

2. SDK 根据发现列表轮询端点。

NodesSupplier 使用以下策略;getQueryEndPoint() 调用 policy.chooseOne(get()):

private final QueryEndPointPolicy policy = new RoundRobinPolicy();

3. 连接失败后回落默认连接,但没有在该路径隔离失败端点。

Session.getQuerySessionConnection(),该方法与 2.0.10 完全一致:

endPointToSessionConnection.computeIfAbsent(
    endPoint.get(),
    k -> {
      try {
        return constructSessionConnection(this, endPoint.get(), zoneId);
      } catch (IoTDBConnectionException ex) {
        return null;
      }
    });

computeIfAbsent 返回 null 时不会保存映射;方法随后回落到默认连接。故障地址仍留在轮询列表,下次选中时再次尝试连接,和实测“每三次出现一次秒级等待”吻合。

期望与希望确认

当一个节点持续不可达且其他节点可完成查询时,希望 SDK 能暂时隔离失败端点并探测恢复,避免后续查询反复承担完整连接超时。

请确认:SHOW AVAILABLE URLS 包含 Unknown 是否为预期?此场景是否应由 SDK 增加故障端点退避/隔离,或调整服务端过滤规则?是否已有对应修复或推荐配置?

目前临时通过关闭自动发现、重定向并只配置健康节点规避,但这需要用户手动维护节点列表。

Activity

  1. kianWei721 commented on Sep 14, 2026

    @kianWei721

    I'm working on this.

    I checked the current master (63ae0f3) and the path described above is unchanged:

    • Session.getQuerySessionConnection() falls back to the default connection when constructSessionConnection fails, but the failed endpoint stays in the NodesSupplier round-robin list, so every rotation pays the full connection timeout again.
    • ShowAvailableUrlsTask still only skips Removing DataNodes, so Unknown nodes are re-added on every node list refresh.

    Proposed fix (client side):

    1. When a query endpoint fails to connect, isolate it for a short period. getQueryEndPoint() skips isolated endpoints and only falls back to the default connection when none are left. Once the period expires the endpoint is eligible again, so a recovered node is picked up without restarting the client; a successful connection clears the mark. Keeping this state in the node supplier lets all sessions of a SessionPool share it.
    2. Make RoundRobinPolicy thread-safe, since a single supplier is shared by every session in the pool.
    3. Unit tests: an isolated endpoint is not retried within the period, is retried after it, and the all-isolated case still falls back to the default connection.

    A few questions before I open the PR:

    • Should the isolation period be a new builder option, or an internal default?
    • Should SHOW AVAILABLE URLS also exclude Unknown nodes? I'd suggest a separate PR for that, since client-side isolation is still needed when a node is Running but unreachable from the client.
    • SessionConnection.reconnect() also iterates over all available nodes. Should it skip isolated endpoints as well, or keep that for a follow-up?
  2. AlanDevise commented on Sep 15, 2026

    @AlanDevise
    Author

    Thanks for the analysis and the proposed fix. Here are my suggestions:

    1. Make the isolation period configurable through the Builder, with a reasonable default.

    The default configuration should provide endpoint isolation without requiring users to tune it, while allowing adjustments for different network environments. All sessions in the same pool should share the isolation state. Refreshing the node list should not clear an unexpired isolation period, otherwise a failed endpoint could re-enter the rotation prematurely. Once the period expires, allow a probe: clear the isolation on a successful connection, or renew it on failure.

    2. Excluding Unknown nodes from SHOW AVAILABLE URLS can be addressed in a separate PR.

    This involves server-side API semantics and compatibility with other clients, so it can be discussed independently without blocking the client-side fix. Even if the server filters out Unknown nodes, a node marked as Running may still be unreachable from the client. SDK-side isolation is therefore still necessary.

    3. SessionConnection.reconnect() should skip endpoints whose isolation period has not expired. I suggest covering this in the same PR.

    Otherwise, normal query endpoint selection would avoid a failed node, but reconnection would still attempt to connect to it and incur the timeout again. Query endpoint selection and reconnection should share the same isolation state.

    When all endpoints are isolated, retain a bounded recovery-probing mechanism with backoff, concurrency limits, and an overall timeout budget, rather than unconditionally retrying every failed endpoint. A controlled early probe could help detect recovery sooner, but each request should not independently retry the entire list of failed endpoints.

    Best wishes.

  3. added 3 commits that reference this issue on Sep 16, 2026
    0f04ed1
    71b0a71
    20e40e8
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions