Repository navigation
Redis online store: HGETALL for wide reads, and an opt-in GIL-free client (valkey-glide) for get_online_features #6856
Description
Activity
Hey @chandlerok, thanks for the writeup. I reproduced the GIL half closely (181 ms idle, 749 ms with one background thread, 8438 ms with eight), so that one clearly holds up.
I couldn't reproduce the HGETALL half I measured 1.3x–2.0x server-side rather than 4.6x, and end-to-end it's usually a net loss, since Feast splits reads per feature view so each HMGET asks for ~22 fields rather than 90. Could you share how you measured the 5.6 vs 25.8 µs — fields per HMGET, hash width, and whether it was raw Redis or through get_online_features?
Please do take the valkey-glide PR if you'd like it you have the production context for it, and I'd be glad to help however is most useful. On HGETALL, would it be reasonable to hold off until we've worked out where our measurements diverge? Entirely your call, and happy to go whichever way you prefer.
hey @patelchaitany! Thanks for taking a look!
I'm guessing you probably measured a narrower shape than I did. My number is for one 123-field listpack hash:
• HMGET asks for ~94 fields
• HGETALL reads the whole hash onceIf you benchmark a single ~20-field HMGET, or swap each per-view HMGET for its own HGETALL, the win mostly disappears because you either aren’t measuring the wide scan or you’re returning the same full hash multiple times.
To reproduce: seed 500 listpack hashes with ~123 fields, reset INFO commandstats, run 500 HMGETs for ~94 fields, reset again, run 500 HGETALLs, then compare cmdstat_hmget.usec_per_call vs cmdstat_hgetall.usec_per_call.
Here's the exact setup behind those numbers, plus the two things that decide whether you see a large ratio or a small one.
Your three questions: raw Redis (not through
get_online_features), 94 named fields perHMGET, 123-field hashes.The setup
500 hashes x 123 fields, then 500
HMGETs of 94 fields, then 500HGETALLs, readingcmdstat_*.usec_per_callafter each phase. This script is runnable as-is.It seeds into database 9 and calls
FLUSHDBon that database, so pointDBat a scratch database before running it:import random import redis HOST, DB = "127.0.0.1", 9 N_KEYS, HASH_FIELDS, HMGET_FIELDS, ROUNDS = 500, 123, 94, 500 client = redis.Redis(host=HOST, port=6379, db=DB) client.flushdb() random.seed(0) # Values must stay under hash-max-listpack-value (64 bytes by default) or the # hash converts to a hashtable and the effect disappears. value = "v" * 4 fields = [f"f{i}" for i in range(HASH_FIELDS)] pipe = client.pipeline(transaction=False) for k in range(N_KEYS): pipe.hset(f"repro:{k}", mapping={f: value for f in fields}) pipe.execute() print("encoding:", client.object("encoding", "repro:0")) # must print listpack # Real Feast field names are mmh3 hashes, so their position in the listpack is # arbitrary. Requesting the first N is the cheapest case and understates this. requested = random.sample(fields, HMGET_FIELDS) def phase(queue, stat): client.execute_command("CONFIG", "RESETSTAT") pipe = client.pipeline(transaction=False) for k in range(N_KEYS): queue(pipe, k) pipe.execute() stats = client.info("commandstats")[f"cmdstat_{stat}"] assert stats["calls"] == ROUNDS, "another client is issuing this command" return stats["usec_per_call"] hmget_us = phase(lambda pipe, k: pipe.hmget(f"repro:{k}", requested), "hmget") hgetall_us = phase(lambda pipe, k: pipe.hgetall(f"repro:{k}"), "hgetall") print(f"HMGET 94 of 123: {hmget_us:6.1f} us/call") print(f"HGETALL 123: {hgetall_us:6.1f} us/call") print(f"ratio: {hmget_us / hgetall_us:.2f}x")
The
assert stats["calls"] == ROUNDSmatters: if anything else on the instance issues the same command,INFO commandstatsis server-wide and the number is polluted. Also worth checking the limits on your instance, since they gate everything below:CONFIG GET hash-max-listpack-entries hash-max-listpack-valueOn the Valkey 9 box I measured on that returned
entries=512, value=64, so a 123-field hash stays listpack either way.What actually decides the ratio
HMGETover a listpack hash does one linear scan from the front per requested field, so 94 fields is ~94 scans.HGETALLis a single scan. Two conditions control how large that gap is.1. The hash has to still be listpack-encoded. If any value exceeds
hash-max-listpack-value(64 bytes by default), that entity's hash converts to a hashtable,HMGETbecomes O(1) per field, andHGETALLstill reads everything, so the advantage inverts. Median of 3 runs, call counts verified at exactly 500:value bytes encoding fields requested HMGET HGETALL ratio 4 listpack first 94 of 123 45.6 us 10.1 us 4.36x 4 listpack random 94 of 123 52.1 us 10.1 us 5.17x 4 listpack last 94 of 123 63.4 us 10.0 us 6.36x 60 listpack random 94 of 123 57.0 us 10.4 us 5.41x 100 hashtable random 94 of 123 10.1 us 11.8 us 0.85x The absolute microseconds are hardware-dependent. The ratio is the portable part, and it is the part that reproduces.
2. How deep in the listpack the requested fields sit. The scan starts at the front, so the first 94 fields is the cheapest case (4.4x) and the last 94 is the most expensive (6.4x). Feast's field names are mmh3 hashes, so their positions are effectively arbitrary, which is the middle row (~5.2x).
The shape that reproduces ~4.6x is therefore: listpack hash, ~123 fields, ~94 fields requested per
HMGETscattered through the hash, measured raw viaINFO commandstats.Why you likely measured 1.3-2.0x
Three things, all of which pull the number toward 1x:
- Mixed encodings within one dataset. Encoding is per hash, so if some entities carry a
ValueProtoblob that pushes a value past 64 bytes, those entities become hashtables while the rest stay listpacks. A real dataset blends both and lands between the 0.85x and 5.4x rows above. This is my best guess at what you hit, and it is easy to check:OBJECT ENCODINGon a sample of entity keys will show listpack vs hashtable directly. - Where the requested fields sit in the listpack. Fields near the front scan cheaply.
- Measuring through
get_online_features. Feast splits reads per feature view, so eachHMGETasks for ~22 fields rather than ~94. That is a genuinely cheaper shape, and it is why the end-to-end number can be a net loss even when the per-command server time improves.
One caveat on my own numbers: the table above is from a different box than the figures in my original comment, so the absolute microseconds differ. I would compare the ratio.
- Mixed encodings within one dataset. Encoding is per hash, so if some entities carry a
Is your feature request related to a problem? Please describe.
The Python
get_online_featurespath against the Redis online store is slow for wide reads, and the cost is on the client, not the server. Our shape is 4 FeatureViews, ~90 fields, 350 to 500 entities per call, served from a FastAPI process that does other work at the same time.Two separate problems show up:
HMGET on listpack-encoded hashes scans once per requested field. Redis and Valkey keep a hash as a listpack up to
hash-max-listpack-entries(default 128 on Redis 7 and Valkey 8). Our hashes have ~120 fields, so every HMGET of ~90 fields does ~90 linear scans. Measured on Valkey 8 withINFO commandstats, same 500 keys:HGETALL returns about a third more bytes and is still 4.6x cheaper for the server. This is the optimization HGETALL optimization for Redis retrieval in Python #3337 asked for in 2022 (the Java server switches to HGETALL above 50 features); that issue was closed by the stale bot without a change.
The Python client work holds the GIL in thousands of short bursts per read. redis-py packs every command, hiredis parses every reply, and each socket receive releases and reacquires the GIL. Then the SDK turns every
ValueProtoblob into a Python object. In a process with other threads running Python, every one of those reacquires waits out the switch interval, so the read time depends on how busy the rest of the process is.Same machine, same data, Feast 0.66 (which already has the single-pipeline read from feat: Addresses performance issues in the Redis online store #6337), 500 entities, ~90 fields across 4 views, p50 of 40 runs:
get_online_features(redis-py + hiredis)The 6x gap under contention is the part that matters in production. It does not show up in a single-threaded benchmark.
Describe the solution you'd like
Two changes to
RedisOnlineStore, independent of each other:In
_read_features_per_fv/ the batchedget_online_features, issue HGETALL instead of HMGET when the number of requested hash fields crosses a threshold (the Java server uses 50), and pick the requested fields out of the reply. Everything else stays the same: same_redis_key, same_mmh3field names, same_ts:<view>presence check.An opt-in client for the Redis online store that runs the pipeline off the GIL. valkey-glide (
valkey-glide-syncon PyPI) is an official Valkey client with a Rust core; a non-atomicBatchis one FFI call, so the whole fetch runs with the GIL released and returns once. It speaks RESP to Redis and Valkey and supports TLS and cluster mode. Aconnection_string-compatibleclient: glideoption onRedisOnlineStoreConfigwould keepfeature_store.yamlas the single place the store is configured.Describe alternatives you've considered
get_online_features_async. Same Python parsing on the same GIL; measured within a few percent of the sync path.hash-max-listpack-entrieson the server so hashes become hashtables. That does remove most of the HMGET cost (7.4 ms to 1.1 ms server time in our test) but costs ~40% more memory and is an operator-side change; HGETALL gets most of the same win from the client.OnlineResponse.protodirectly instead ofto_dict(). Worth about 20%, but the redis-py fetch is still the part that stalls under contention.We have both changes running in a fork of the read path and are happy to contribute either as a PR if there is interest in the direction.
Additional context
Related: #3337 (HGETALL request, closed stale), #4711 and #6337 (single pipeline across feature views, merged; the numbers above are with that in place), #3649 (registry
from_protooverhead, a separate cost).