Repository navigation
fix: Map narrow and unsigned Arrow integer types to Feast value types - #6960
raashish1601 wants to merge 2 commits into
Conversation
|
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #6960 +/- ##
=======================================
Coverage 49.32% 49.32%
=======================================
Files 443 443
Lines 55094 55094
Branches 8017 8017
=======================================
Hits 27176 27176
Misses 26035 26035
Partials 1883 1883
Continue to review full report in Codecov by Harness.
🚀 New features to boost your workflow:
|
f437d04 to
5d31929
Compare
|
One thing worth resolving before merge:
https://github.com/feast-dev/feast/blob/master/sdk/python/feast/type_map.py#L397 but this PR maps To be clear, I think the |
|
Agreed, thanks. I changed |
|
Please resolve conflicts |
Signed-off-by: Raashish Aggarwal <94279692+raashish1601@users.noreply.github.com>
Signed-off-by: Raashish Aggarwal <94279692+raashish1601@users.noreply.github.com>
|
Rebased on master and resolved the conflict in |
ec11d88 to
e6c9c93
Compare
What this PR does / why we need it:
pa_to_feast_value_typeonly knowsint32andint64, so a Parquet file with a narrower or unsigned integer column breaks schema inference for aFileSource(and the other sources that use this decoder):python_type_to_feast_value_typealready maps these numpy dtypes (int8,int16,uint8,uint16toINT32), so the same data works from pandas but not from its Arrow schema.This adds the Arrow names to the map:
int8,int16,uint8anduint16decode toINT32, anduint32anduint64toINT64(foruint32I pickedINT64so every value fits). Lists of them decode to the matching_LISTtype through the existing list handling.Which issue(s) this PR fixes:
No issue, found while checking Arrow type decoding.
Checks
git commit -s)Testing Strategy
Added
test_pa_to_feast_value_type_small_and_unsigned_intsinsdk/python/tests/unit/test_type_map.py(8 cases, all fail on master).pytest tests/unit/test_type_map.pypasses apart from theTestSparkNativeTypeValidationcases, which fail the same way on master here because pyspark isn't installed.ruff checkandruff format --checkare clean on the changed files.