This document specifies the canonical on-disk layout of FLOAT/DOUBLE TS_2DIFF
pages, derived from the Java reference implementation
(FloatEncoder, FloatDecoder, DeltaBinaryEncoder, BitMap).
The Java layout is the cross-language compatibility boundary. Other layouts
produced by earlier C++ writers (raw bit-cast, per-block wrapper metadata) are
implementation artifacts outside the compatibility scope; the C++ decoder
treats them as a format error.
TS_2DIFF encodes integers. Floating-point values go through a wrapper that
converts each value to an integer, encodes the integers with
IntDeltaEncoder (FLOAT) or LongDeltaEncoder (DOUBLE), and emits page-wide
conversion metadata.
Given maxPointNumber = mpn and maxPointValue = 10^mpn (mpn <= 0 implies
maxPointValue = 1), each value maps to one of three stored forms:
| Condition | Stored bits | Decoder action |
|---|---|---|
round(v * 10^mpn) fits the int type |
round(v * 10^mpn) |
divide by 10^mpn |
scaled overflows but v itself fits |
round(v) |
divide by 1 |
v out of int range, or NaN |
floatToIntBits(v) / doubleToLongBits(v) |
restore raw bits |
The three forms are tracked per page as a tri-state flag list
(underflowFlags in Java):
true-> scaled formfalse-> rounded form (scale overflow)null-> raw IEEE 754 bits (value overflow or NaN)
round is Java Math.round semantics, floor(x + 0.5) — ties go towards
+infinity (-2.5 -> -2), unlike C lround's ties-away-from-zero.
# Form 1: every value stored in scaled form (no bitmap at all)
[maxPointNumber varint]
[TS_2DIFF block 1][TS_2DIFF block 2]...[final block]
# Form 2: at least one value is 'false' (scale overflow), none is 'null'
[Integer.MAX_VALUE varint] # 0xFF 0xFF 0xFF 0xFF 0x07
[pageValueCount varint]
[scaled-bitmap, pageValueCount/8+1 bytes] # marks 'true' entries
[maxPointNumber varint]
[TS_2DIFF block 1]...[final block]
# Form 3: at least one value is 'null' (raw bits)
[Integer.MAX_VALUE-1 varint] # 0xFF 0xFF 0xFF 0xFF 0x06
[pageValueCount varint]
[scaled-bitmap, pageValueCount/8+1 bytes] # marks 'true' entries
[raw-bitmap, pageValueCount/8+1 bytes] # marks 'null' entries
[maxPointNumber varint]
[TS_2DIFF block 1]...[final block]
Key invariants:
maxPointNumberappears exactly once per page, before the first integer block (for Forms 2/3 it appears after the bitmaps).- The bitmaps cover the entire page, not individual TS_2DIFF blocks. The
decoder keeps one page-wide
positionthat never resets between blocks; only a page-levelreset()clears it. - Bitmap byte length is always
size/8 + 1, even whensize % 8 == 0(BitMap.getSizeOfBytes). - Bitmap bit order is LSB-first within each byte: position
pmaps tobits[p / 8] & (1 << (p % 8)). pageValueCountcounts all values of the page (across blocks).- A first-page byte of
0x00is the normal encoding ofmaxPointNumber = 0(an explicitmax_point_number=0property on the Java side; reachable but not the default — see Encoder Construction below). It is not a legacy marker. - NaN handling is writer-dependent: Java
floatToIntBits/doubleToLongBitscanonicalize any NaN to0x7fc00000/0x7ff8000000000000, while the C++ encoder preserves the payload bits. Both are valid raw-bit entries; readers restore the bits as stored.
Identical to the integer TS_2DIFF format (DeltaBinaryEncoder):
[writeIndex int32 BE] # number of deltas in this block
[bitWidth int32 BE]
[block-specific header] # first value; min delta
[packed data] # writeIndex * bitWidth bits
A block stores writeIndex + 1 values (first value + writeIndex deltas).
BLOCK_DEFAULT_SIZE = 128 is only DeltaBinaryEncoder's default buffer
size, not a wire-format limit: Java exposes block-size constructors
(IntDeltaEncoder(int)), so writeIndex is bounded only by the declared
page value count and the packed bytes available. A 300-value page from the
default encoder produces blocks of 129, 129, 42 values.
The Java TSEncodingBuilder.Ts2Diff field initializes maxPointNumber = 0,
but the standard schema write path (MeasurementSchema.getValueEncoder)
always calls initFromProps(), which replaces it with the schema's
max_point_number property or, when the property is absent, with
TSFileConfig.floatPrecision (current default 2). A writer that
explicitly sets max_point_number = 0 produces Form 1 pages starting with
0x00. The C++ FloatTS2DIFFEncoder / DoubleTS2DIFFEncoder default to
2, matching the standard Java schema path. The value stored in the
stream is self-describing, so files written with other maxPointNumber
values remain readable.
Per page, exactly once, the decoder reads the leading marker:
- Read varint
tag. tag == Integer.MAX_VALUE-> readcountvarint,count/8+1bytes scaled-bitmap, then varintmaxPointNumber(Form 2).tag == Integer.MAX_VALUE-1-> additionally read a secondcount/8+1bytes raw-bitmap (Form 3).- Otherwise
tagitself ismaxPointNumber(Form 1);mpn <= 0meansmaxPointValue = 1.
Then values are decoded from the integer blocks. For value at page position
p:
- raw-bitmap (if present) marks
p->intBitsToFloat/longBitsToDouble - else scaled-bitmap (if present) marks
p->value / 10^mpn - else ->
value / 1
Any input that does not conform to this grammar (for example, an integer TS_2DIFF block header where the page metadata is expected) is a format error.