Visitar URL original
ci: Pin CPU-only torch and bound floating dependencies by patelchaitany · Pull Request #6825 · feast-dev/feast · GitHub
Skip to content

ci: Pin CPU-only torch and bound floating dependencies - #6825

Merged
ntkathole merged 4 commits into
feast-dev:masterfrom
patelchaitany:ci/pytorch-cpu-index
Sep 15, 2026
Merged

ntkathole merged 4 commits into
feast-dev:masterfrom
patelchaitany:ci/pytorch-cpu-index

Conversation

@patelchaitany

@patelchaitany patelchaitany commented Sep 10, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it:

Every CI job currently fails in the install step:

× Failed to download `torchvision==0.28.0+cpu`
╰─▶ Hash mismatch for `torchvision==0.28.0+cpu`

The CI requirement files are compiled without a torch backend, so their hashes describe PyPI wheels, but install-python-dependencies-ci synced them on Linux with --torch-backend cpu, which serves the +cpu builds from download.pytorch.org. Those are different artifacts with different hashes.

The divergence has existed since #6588 added the flag to sync only; what changed is verification. uv 0.12.11 began applying hashes from public-version pins to matching local versions when no exact local-version hash is provided, so the PyPI hashes recorded for torchvision==0.28.0 are now checked against the +cpu wheel. Earlier releases found no hash for the local version and skipped the check. uv is unpinned in the workflows, so CI moved from 0.12.7 to 0.12.11 and every job began failing.

Approach

Compile and sync now both pass --torch-backend cpu, so they can no longer disagree about where torch comes from.

The CI locks are compiled --universal so one file serves both the Linux and macOS runners, pinning each variant with its own hashes:

torch==2.13.0           ; sys_platform == 'darwin'
torch==2.13.0+cpu       ; sys_platform != 'darwin'
torchvision==0.28.0     ; sys_platform == 'darwin'
torchvision==0.28.0+cpu ; sys_platform != 'darwin'

Without --universal the lock pins an unmarked torch==2.13.0 carrying PyPI hashes. Linux still resolves that to 2.13.0+cpu, because ==2.13.0 matches the local version under PEP 440, and the download then fails verification — the original breakage in a different guise.

An explicit index with [tool.uv.sources] does not serve here. The pip interface applies sources only at compile time; uv pip sync ignores them, so the +cpu builds would not be found at all.

No nvidia-* package remains in any CI requirements file, which is what #6588 set out to achieve.

pixi raises its uv floor to 0.6.9, the first release with --torch-backend. The uv already locked in infra/scripts/pixi/pixi.lock satisfies it, so that lock is untouched.

Dependency bounds

Recompiling exposed five dependencies with no upper bound, which float into regressions on every recompile. Two of them are already breaking jobs on this PR's own CI:

  • mcp — 1.30 stops returning the mcp-session-id response header, so the handshake in mcp-feature-server-runtime fails (the server itself starts fine). 98e5bca already pinned 1.29.0 in the lock files; the bound makes that durable instead of resettable on the next recompile.
  • ray — 2.56+ returns its own extension dtypes from ray.data, failing 17 assertions in test_universal_types.py (is_object_dtype(TensorDtype(...)) → False). feast declares ray directly only on 3.10, so the bound is repeated for 3.11+ where ray arrives transitively through codeflare-sdk.
  • pyarrow, mlflow — capped at the versions 3.10 and 3.11 already lock to, so all three interpreters agree.

Every cap holds a package at a version already present on master. The bounds also bring two interpreters back in line with the others: py3.10 was at ray==2.56.1, and py3.12 at pyarrow==24.0.0 / mlflow==3.14.0, where 3.10 and 3.11 sat at 21.0.0 / 3.2.0.

pixi.lock moves pyarrow 24.0.0 → 21.0.0 in all four environments and nothing else — pyproject.toml is the root pixi manifest, so the bound invalidates the lock and setup-pixi checks it.

The locks were recompiled in place over the existing files, which uv uses as preference anchors; that is what keeps the diff at ~1,200 lines rather than tens of thousands.

Which issue(s) this PR fixes:

None filed; this repairs CI on master.

Checks

  • I've made sure the tests are passing.
  • My commits are signed off (git commit -s)
  • My PR title follows conventional commits format

Testing Strategy

  • Unit tests
  • Integration tests

Verified locally on a macOS host, under uv 0.12.13 — the release line whose stricter hash checking broke CI — so the full Linux install was not exercised end to end:

  • uv pip sync --dry-run --python-platform x86_64-unknown-linux-gnu --torch-backend cpu against py3.11-ci-requirements.txt resolves all 445 packages and selects torch==2.13.0+cpu and torchvision==0.28.0+cpu.
  • A real hash-checked install of the CPU torchvision wheel for the Linux platform downloads from download.pytorch.org and verifies clean — the specific artifact that was failing.
  • The same install against master's lock entry reproduces the failure, confirming --universal is load-bearing rather than cosmetic.
  • pixi lock --check (pixi 0.75.0, the version CI pins) is clean on master and reports pyarrow as the only movement under the new bounds.

CI on this PR is the actual proof.

🤖 Generated with Claude Code

@patelchaitany
patelchaitany requested a review from a team as a code owner September 10, 2026 10:38
@codecov-commenter

codecov-commenter commented Sep 10, 2026 •

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

❌ Patch coverage is 0% with 6 lines in your changes missing coverage. Please review.
✅ Project coverage is 47.09%. Comparing base (81e1546) to head (113f31a).
⚠️ Report is 3 commits behind head on master.

Files with missing lines Patch % Lines
sdk/python/feast/infra/ray_initializer.py 0.00% 6 Missing ⚠️
❗ Your organization needs to install the Codecov GitHub app to enable full functionality.
Additional details and impacted files

Impacted file tree graph

@@           Coverage Diff           @@
##           master    #6825   +/-   ##
=======================================
  Coverage   47.08%   47.09%           
=======================================
  Files         419      419           
  Lines       51877    51885    +8     
  Branches     7525     7528    +3     
=======================================
+ Hits        24428    24436    +8     
+ Misses      25700    25699    -1     
- Partials     1749     1750    +1     
Flag Coverage Δ
go-feature-server 30.58% <ø> (ø)
python-unit 48.40% <0.00%> (+<0.01%) ⬆️
Files with missing lines Coverage Δ
sdk/python/feast/infra/ray_initializer.py 0.00% <0.00%> (ø)

... and 3 files with indirect coverage changes


Continue to review full report in Codecov by Harness.

Legend - Click here to learn more
Δ = absolute <relative> (impact), ø = not affected, ? = missing data
Powered by Codecov. Last update 9f7b503...113f31a. Read the comment docs.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

@haoxu0 haoxu0 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The core fix is sound, but this PR currently regenerates roughly 64,000 lines across the dependency locks and upgrades many packages unrelated to the torchvision hash failure, including mcp 1.29.0 -> 1.30.0, mlflow 3.2.0 -> 3.16.0, pyarrow 21.0.0 -> 25.0.1, and Ray-related dependencies.

The remaining CI failures show why this is risky: the MCP runtime no longer satisfies the workflow's expected response contract, and the Ray integration suite now has 17 timestamp/dtype failures. These regressions are outside the intended Torch hash repair.

Please keep the compile/sync CPU-backend alignment and universal CI lock, but preserve the existing non-Torch dependency versions where possible. In particular, restrict regenerated requirements to the three CI lock files, revert the non-CI/minimal requirement files and unrelated root pixi.lock churn, and avoid globally changing every uv resolution path if passing --torch-backend cpu consistently at CI compile and sync can provide the same guarantee.

@patelchaitany patelchaitany changed the title ci: Pin CPU-only torch through the uv torch-backend setting ci: Pin CPU-only torch and bound floating dependencies Sep 12, 2026
@patelchaitany
patelchaitany force-pushed the ci/pytorch-cpu-index branch 2 times, most recently from ba5e35a to 5698480 Compare September 12, 2026 17:23
Comment thread pyproject.toml Outdated
# ray arrives transitively via codeflare-sdk above 3.10; bound it there too so the
# cap applies on every interpreter.
'ray<2.56; python_version > "3.10"',
'codeflare-sdk>=0.31.1,<0.39; python_version > "3.10"',

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

any reason to add upper bound for ray and codeflare-sdk versions ?

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can we fix test_universal_types instead?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixing in the store, not the test public API regression.

Comment thread pyproject.toml Outdated
mcp = ["fastapi_mcp", "mcp>=1.0,<2"]
mlflow = ["mlflow>=2.10.0"]
mcp = ["fastapi_mcp", "mcp>=1.0,<1.30"]
mlflow = ["mlflow>=2.10.0,<3.3"]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

same here for mlflow ?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

for the MLflow Versions diverged across 3.10/3.11/3.12.

patelchaitany and others added 3 commits September 14, 2026 21:05
Every CI job fails in the install step:

    x Failed to download `torchvision==0.28.0+cpu`
    `-> Hash mismatch for `torchvision==0.28.0+cpu`

The CI locks are compiled without a torch backend, so their hashes
describe PyPI wheels, but `install-python-dependencies-ci` synced them on
Linux with `--torch-backend cpu`, which serves the `+cpu` builds from
download.pytorch.org. Different artifacts, different hashes.

The divergence has existed since feast-dev#6588 added the flag to sync only; what
changed is verification. uv 0.12.11 began applying hashes from public
version pins to matching local versions, so the PyPI hashes recorded for
`torchvision==0.28.0` are now checked against the `+cpu` wheel. uv is
unpinned in the workflows, so CI moved from 0.12.7 to 0.12.11 and every
job started failing.

Compile and sync now both pass `--torch-backend cpu`, so they can no
longer disagree about where torch comes from. The CI locks are compiled
`--universal` so one file serves both the Linux and macOS runners,
pinning each variant with its own hashes:

    torch==2.13.0           ; sys_platform == 'darwin'
    torch==2.13.0+cpu       ; sys_platform != 'darwin'
    torchvision==0.28.0     ; sys_platform == 'darwin'
    torchvision==0.28.0+cpu ; sys_platform != 'darwin'

Without `--universal` the lock pins an unmarked `torch==2.13.0` carrying
PyPI hashes. Linux still resolves that to `2.13.0+cpu`, because
`==2.13.0` matches the local version under PEP 440, and the download then
fails verification -- the original breakage in a different guise.

No `nvidia-*` package remains in any CI requirements file, which is what
feast-dev#6588 set out to achieve.

Recompiling the locks also exposed five dependencies with no upper bound,
which float into regressions on every recompile:

  * mcp -- 1.30 drops the `mcp-session-id` response header, so the
    feature server handshake in mcp-feature-server-runtime fails.
    98e5bca already pinned 1.29.0 in the lock files; the bound makes
    that durable instead of resettable on the next recompile.
  * ray -- 2.56+ returns its own extension dtypes from ray.data, failing
    17 type assertions in test_universal_types.py. feast declares ray
    directly only on 3.10, so the bound is repeated for 3.11+ where ray
    arrives transitively through codeflare-sdk.
  * pyarrow, mlflow -- capped at the versions 3.10 and 3.11 already lock
    to, so all three interpreters agree.

The caps hold every package at a version already present on master, and
bring py3.10 (ray 2.56.1) and py3.12 (pyarrow 24.0.0, mlflow 3.14.0) back
in line with the other interpreters. pixi.lock moves pyarrow 24.0.0 ->
21.0.0 and nothing else.

pixi raises its uv floor to 0.6.9, the first release with
`--torch-backend`. The uv already locked in infra/scripts/pixi/pixi.lock
satisfies it, so that lock is untouched.

Verified on a macOS host, so the full Linux install was not exercised
end to end:

  * `uv pip sync --dry-run --python-platform x86_64-unknown-linux-gnu
    --torch-backend cpu` against py3.11-ci-requirements.txt resolves all
    445 packages and selects torch==2.13.0+cpu and torchvision==0.28.0+cpu.
  * A real hash-checked install of the CPU torchvision wheel for the
    Linux platform, under uv 0.12.13, downloads from download.pytorch.org
    and verifies clean -- the specific artifact that was failing.
  * The same install against master's lock entry reproduces the failure,
    confirming the universal marker split is load-bearing.
  * `pixi lock --check` is clean on master and reports pyarrow as the
    only movement under the new bounds.

CI on the PR is the actual proof.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com>
mlflow 3.2.0 already requires `pyarrow<22,>=4.0.0`, so the `mlflow<3.3`
bound implies the pyarrow cap wherever mlflow takes part in the
resolution. The explicit `pyarrow>=16.1.0,<22` added nothing, and it
broke test_flink_extra_constrains_shared_pyarrow_dependency, which
compares the core pyarrow spec by exact string.

Dropping it leaves all three CI requirement files byte-identical -- the
resolution is unchanged, including py3.12 coming down to pyarrow 21.0.0,
which mlflow drives rather than the removed bound.

pixi.lock shrinks to a metadata-only update: the recorded requires_dist
for the editable feast package now lists the four remaining bounds, and
no package version moves. It keeps pyarrow 24.0.0 in the pixi
environments, as master does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com>
Ray 2.56 changed BlockAccessor.to_pandas to map Arrow types onto
pandas Arrow-backed dtypes, so RayRetrievalJob.to_df() began returning
pd.ArrowDtype columns where every other offline store returns
numpy/object. That fails 17 assertions in test_universal_types, and
more importantly changes the dtypes users get back from
get_historical_features(...).to_df() on the Ray store alone.

Ray 2.56.1 added DataContext.enable_arrow_backed_pandas_conversion as
the opt-out. Set it at each of the three initialization sites, next to
the enable_tensor_extension_casting already pinned there for the same
reason. The hasattr guard doubles as a version check: the attribute
exists only from 2.56.1, and below that the behaviour it disables does
not exist either.

Fixing the store is what lets the dependency caps come off, rather than
loosening the universal type assertions to accommodate one store:

- ray<2.56 and codeflare-sdk<0.39 are dropped; codeflare-sdk pins ray
  exactly, so the two moved in lockstep anyway.
- mlflow relaxes from <3.3 to <4, matching the major-version bound every
  other extra in the file carries. The <3.3 cap held mlflow 14 minor
  releases back on no demonstrated breakage.

The CI locks were recompiled in place with --upgrade-package limited to
those three, so all three interpreters now agree on ray 2.58.0,
codeflare-sdk 0.39.0 and mlflow 3.16.0. The torch and torchvision pins
and their hashes are untouched, and no nvidia-* package reappears.

Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com>
Comment thread pyproject.toml Outdated
mcp = ["fastapi_mcp", "mcp>=1.0,<2"]
mlflow = ["mlflow>=2.10.0"]
mcp = ["fastapi_mcp", "mcp>=1.0,<1.30"]
mlflow = ["mlflow>=2.10.0,<4"]

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

<4 constraint can also be removed, adds future maintenance

Review feedback on feast-dev#6825: the `<4` cap adds maintenance for no benefit.

mlflow's latest release is 3.16.0, so the bound never constrained
resolution. Recompiling all three CI locks with it removed leaves them
byte-identical, and the pyarrow cap that originally motivated aligning
mlflow across interpreters is already gone.

pixi.lock only mirrors the requires-dist metadata; `pixi lock --check`
(pixi 0.75.0, the version CI pins) is clean.

Signed-off-by: Chaitany Patel <patelchaitany93@gmail.com>
@ntkathole
ntkathole merged commit b5090c6 into feast-dev:master Sep 15, 2026
29 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants