Visitar URL original
[fix][docs] Correct entry-bucket, ordering and auto-scaling claims for checkpoint consumers by merlimat · Pull Request #1245 · apache/pulsar-site · GitHub
Skip to content

[fix][docs] Correct entry-bucket, ordering and auto-scaling claims for checkpoint consumers - #1245

Open
merlimat wants to merge 2 commits into
apache:mainfrom
merlimat:mmerli/checkpoint-consumer-docs
Open

merlimat wants to merge 2 commits into
apache:mainfrom
merlimat:mmerli/checkpoint-consumer-docs

Conversation

@merlimat

@merlimat merlimat commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Documents three V5 checkpoint consumer fixes:

  • whole-segment assignment for consumer groups;
  • auto-scaling that splits but never rebuckets for groups;
  • per-key ordering across splits and merges for ungrouped consumers.

The first two are merged in apache/pulsar#26838, which also carried the commit of apache/pulsar#26836. The ordering fix is apache/pulsar#26837, still open.

Problem. The scalable-topic pages describe checkpoint consumers as if they worked like stream consumers:

  • A checkpoint consumer group is said to divide "segments and entry buckets" among its members and to drive automatic rebucketing. But a checkpoint consumer reads each segment through its own Reader, so it can't share a segment by entry bucket; the client specification's conformance section says checkpoint consumers are unaffected by entry-bucketing. The controller now gives a group whole segments, one member each, and a group's member count no longer triggers a rebucket.
  • "Stream and checkpoint consumers preserve per-key ordering across layout changes." A checkpoint consumer group does not: after a split or merge, one member can receive a key's newer messages from a successor segment while another member is still reading the key's older messages from the predecessor.
  • "Automatic merges preserve the parallelism currently used by these coordinated consumers." The merge guard protects the entry-bucket parallelism of stream consumers. A checkpoint consumer group uses at most one member per segment, so a merge can leave one of its members idle.

Example. concepts-scalable-topics.md, "Entry buckets and consumer parallelism":

A segment can serve several stream or checkpoint consumers without creating more physical segments. [...] The controller assigns these buckets to consumers in the same group [...]

When more stream consumers or grouped checkpoint consumers join, automatic scaling chooses between splitting a segment and increasing its bucket count. [...]

Change.

  • docs/concepts-scalable-topics.md:
    • Entry buckets serve the stream consumers of a subscription; a checkpoint consumer group gives each segment to a single member.
    • A group with more members than segments can trigger only a split, under the same traffic conditions as stream consumers. On a lower-traffic topic or at the segment limit, its extra members stay idle.
    • Per-key ordering across layout changes holds for stream consumers and ungrouped checkpoint consumers, not for groups.
  • docs/admin-api-scalable-topics.md:
    • Below scalableTopicSplitVsRebucketMinMsgRateInThreshold, or at the segment ceiling, bucket capacity grows for stream consumers only, and grouped checkpoint consumers drive only splits.
    • Merges preserve stream consumers' entry-bucket parallelism and can leave a group member idle.
    • The broker defaults table says the same for that threshold. The generated configuration reference picks up the matching ServiceConfiguration description from [fix][broker] Stop checkpoint consumer groups from triggering rebucket rollovers pulsar#26838 on the next docs sync.
  • client-libraries/java-v5.md: a group distributes segments, not entry buckets. Each segment is read by one member, extra members stay idle, and a group does not preserve per-key order across splits and merges.
  • versioned_docs/version-5.0.x/concepts-scalable-topics.md: the entry-bucket and ordering corrections. The old text was wrong for every 5.0.x release, and the new text matches 5.0.x once the group and ordering fixes are backported to branch-5.0. The auto-scaling corrections are not mirrored there yet: a 5.0.x broker still rebuckets for a group's member count until #26838 is backported.

Testing. Docs-only prose edits inside existing paragraphs; no links, anchors or markup added. The statements follow the code in the PRs above and the client specification. The site was not built locally.

✅ Contribution Checklist

…onsumers

Checkpoint consumers do not use entry buckets: a checkpoint consumer group
gives each segment to a single member, and extra members stay idle (the
client specification, key-shared.md section 9). Only stream consumers and
ungrouped checkpoint consumers preserve per-key ordering across splits and
merges; a checkpoint consumer group has no way to wait for a predecessor
segment that another member is still reading.
Since apache/pulsar#26838, only stream subscriptions count toward an automatic
rebucket rollover. A checkpoint consumer group reads whole segments, one member
each, so more entry buckets would not serve it: its member count drives only a
split, and on a lower-traffic topic or at the segment limit its extra members
stay idle.

Also correct the claim that automatic merges preserve the parallelism of every
coordinated consumer. The merge guard protects the entry-bucket parallelism of
stream consumers; a checkpoint consumer group uses at most one member per
segment, so a merge can leave one of its members idle.

The versioned 5.0.x pages are unchanged: they still describe the 5.0.x broker,
which has neither fix.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant