Repository navigation
Tags: apache/tvm
Tags
[FIX][TIRx][CUDA] Support low-precision, tensor-map and cluster kerne… …ls in CUDA-host bundles (#20489) A CUDA-host bundle (`tvm.backend.cuda.export_cuda_host`, #20395) compiles device and host code as one NVCC C++ translation unit. Building the native kernels of mlc-ai/TIRx-kernels this way exposed several patterns that the bundle could not compile or launch: - **Low-precision types in the host pass.** The CUDA device header included `cuda_fp16.h`, `cuda_bf16.h`, `cuda_fp8.h`, `cuda_fp6.h` and `cuda_fp4.h` (and defined the `fp8_e4_t`-style aliases) only under `defined(__CUDA_ARCH__)`. NVCC's host pass then fails on kernel signatures that use `half` or `nv_bfloat16`. The guards now also admit the host pass (`!defined(__CUDA_ARCH__) || __CUDA_ARCH__ >= N`); NVRTC and device-only NVCC compilation always define `__CUDA_ARCH__` and are unchanged. - **Tensor-map parameters.** The host wrapper bound a `T.TensorMap()` parameter as `CUtensorMap* x = ((void*)...)`, which C accepts and C++ rejects. It now casts explicitly. - **Argument types at the launch.** Host C code spells some device types differently (bfloat16 is `uint16_t*` on the host, `nv_bfloat16*` in the kernel), so the direct `<<<>>>` launch did not type-check. Launches now go through a small helper that converts each argument to the kernel's own parameter type and calls `cudaLaunchKernelEx`. - **Launch attributes.** `clusterCtaIdx.*`, `preferredClusterCtaIdx.*`, programmatic dependent launch and cooperative launch were rejected. They now become the corresponding `cudaLaunchAttribute`s, following `CUDAWrappedFunc` in the CUDA runtime module (including `cudaFuncAttributeNonPortableClusterSizeAllowed` for cluster launches and omitting a unit preferred cluster). Required block dimensions remain unsupported and are still diagnosed.
[FIX][TIRx] Fix boolean bitwise-not codegen for C-like targets (#20445) Fix incorrect code generation for boolean `BitwiseNot`. For boolean values, bitwise NOT should flip `false` to `true` and `true` to `false`. The current C code generator emits `~value`, but C/C++ first converts the boolean to an integer: - `~false` becomes `~0`, which is `-1`. - `~true` becomes `~1`, which is `-2`. Both results are nonzero, so converting them back to boolean produces `true` in both cases. CUDA boolean vectors have a related problem: their generated vector types do not support the `~` operator, causing compilation to fail. This change emits `value == false` for boolean operands. It produces the correct result for scalar booleans and uses the existing vector comparison code to apply the operation to each lane. Integer operands continue to use `~value`..
[FIX][TIRx] Fix boolean bitwise-not codegen for C-like targets (#20445) Fix incorrect code generation for boolean `BitwiseNot`. For boolean values, bitwise NOT should flip `false` to `true` and `true` to `false`. The current C code generator emits `~value`, but C/C++ first converts the boolean to an integer: - `~false` becomes `~0`, which is `-1`. - `~true` becomes `~1`, which is `-2`. Both results are nonzero, so converting them back to boolean produces `true` in both cases. CUDA boolean vectors have a related problem: their generated vector types do not support the `~` operator, causing compilation to fail. This change emits `value == false` for boolean operands. It produces the correct result for scalar booleans and uses the existing vector comparison code to apply the operation to each lane. Integer operands continue to use `~value`..
[FIX] Update v0.27.0 versions and FFI floor (#20390) Raise the apache-tvm-ffi minimum to 0.1.14.post0 and update the matching CI dependency assertion. Set the Python, C++, and package metadata fallback versions to 0.27.0 so source archives without Git metadata no longer report 0.26.dev0. Tagged builds continue to derive their version from Git.
tirx: represent buffer parameters with BufferType (#20086) This phases out PrimFunc.buffer_map entirely. Buffer parameters now carry their BufferType annotation directly in PrimFunc.params, with no compatibility constructor or derived PrimFunc buffer map. The migration updates construction and visitor paths across TIRx, TE, S-TIR, Relax, printers/parsers, packed-ABI lowering, specialization, storage rewriting, and related tests and documentation. BufferType signature shapes use left-to-right match-scope binding: the first quoted shape expression defines missing symbolic variables, and later dimensions, parameters, and matching scalar declarations reuse those exact variables.
[Fix][TIRx] Fix MSVC build of IndexDataTypeNormalizer (#20096) The Windows wheel build for `v0.26.0.rc0` fails to compile: ``` src\tirx\ir\data_type_rewriter.cc(696,47): error C2352: 'tvm::tirx::ExprMutator::VisitPrimExpr': a call of a non-static member function requires an object ``` `StmtExprMutator` derives from both `ExprMutator` and `StmtMutator`, and re-exports the name via `using ExprMutator::VisitPrimExpr;`. MSVC resolves the class-qualified `IndexDataTypeNormalizer::VisitPrimExpr` down to `ExprMutator::VisitPrimExpr` and then fails to form the implicit object conversion. GCC and Clang accept the same expression, so only the Windows leg broke — the macOS and both Linux wheels built fine. The fix calls it through `this` instead, matching every other `VisitPrimExpr` call site in this file. `VisitPrimExpr` is a non-virtual inline helper, so the qualification was suppressing nothing and behavior is unchanged. The other class-qualified call sites in the tree name `StmtExprMutator` directly — that is where the using-declaration lives, so they resolve fine and are left alone. The call was introduced in #19931, which changed `IndexDataTypeNormalizer::VisitExpr` to `IndexDataTypeNormalizer::VisitPrimExpr`; the former resolved unambiguously. Targeting the release branch first to unblock the `v0.26.0.rc0` Windows wheel; it will be ported to `main` separately. Failing job: https://github.com/apache/tvm/actions/runs/31033249253/job/92400739674
[Fix][TIRx] Fix MSVC build of IndexDataTypeNormalizer (#20096) The Windows wheel build for `v0.26.0.rc0` fails to compile: ``` src\tirx\ir\data_type_rewriter.cc(696,47): error C2352: 'tvm::tirx::ExprMutator::VisitPrimExpr': a call of a non-static member function requires an object ``` `StmtExprMutator` derives from both `ExprMutator` and `StmtMutator`, and re-exports the name via `using ExprMutator::VisitPrimExpr;`. MSVC resolves the class-qualified `IndexDataTypeNormalizer::VisitPrimExpr` down to `ExprMutator::VisitPrimExpr` and then fails to form the implicit object conversion. GCC and Clang accept the same expression, so only the Windows leg broke — the macOS and both Linux wheels built fine. The fix calls it through `this` instead, matching every other `VisitPrimExpr` call site in this file. `VisitPrimExpr` is a non-virtual inline helper, so the qualification was suppressing nothing and behavior is unchanged. The other class-qualified call sites in the tree name `StmtExprMutator` directly — that is where the using-declaration lives, so they resolve fine and are left alone. The call was introduced in #19931, which changed `IndexDataTypeNormalizer::VisitExpr` to `IndexDataTypeNormalizer::VisitPrimExpr`; the former resolved unambiguously. Targeting the release branch first to unblock the `v0.26.0.rc0` Windows wheel; it will be ported to `main` separately. Failing job: https://github.com/apache/tvm/actions/runs/31033249253/job/92400739674
[CI] Build CUDA sidecar on plain manylinux, install CUDA manually (#1… …9882) #19754 (besides bumping cibuildwheel to 4.1.0) switched the CUDA runtime sidecar build onto the preinstalled quay.io/manylinux_cuda image and dropped the in-script CUDA toolkit install. On that image the built libtvm_runtime_cuda.so silently lost its libcuda.so.1 dependency: the libcuda.so stub is not where CMake's FindCUDA looks, so CUDA_CUDA_LIBRARY resolved empty and ldd no longer lists libcuda, breaking downstream runtime use. This restores only the sidecar pieces -- the plain manylinux_2_28 image and the curl/rpm/dnf CUDA toolkit install. The cibw 4.1.0 bump is kept (confirmed working). The driver stub is back where FindCUDA finds it, so the sidecar links against libcuda.so.1 again, as ldd confirms. This is a hotfix; a follow-up can re-adopt the CUDA image with the stub resolved explicitly plus a post-build linkage check.
[Fix] Revert C++20-only lambda captures for C++17 build Revert four explicit-this lambda captures ([=, this, ...] -> [=, ...]) in run_codegen.cc, inject_software_pipeline.cc, and cudnn_json_runtime.cc. These were introduced by the C++20 baseline upgrade (#19734); after reverting the baseline to C++17, MSVC rejects capturing 'this' explicitly under a '=' capture-default (error C3791) -- that form is C++20-only. GCC and Clang accept it as an extension in C++17, which is why only the Windows wheel build failed. Dropping the explicit 'this' ('=' captures it implicitly in C++17) restores the MSVC build with no behavior change.
PreviousNext