Skip to content

Add block_bucketize_sparse_features operator - #76

Draft
flezaalv wants to merge 34 commits into
intel:mainfrom
aagalleg:flezaalv/feat/block_bucketize_sparse_features_operator
Draft

Add block_bucketize_sparse_features operator#76
flezaalv wants to merge 34 commits into
intel:mainfrom
aagalleg:flezaalv/feat/block_bucketize_sparse_features_operator

Conversation

@flezaalv

@flezaalv flezaalv commented Jun 25, 2026

Copy link
Copy Markdown
Contributor

Summary

Completes the integration of the block_bucketize_sparse_features operator by fixing CMake configuration and resolving dynamic library loading issues.

Depends on: #74

Changes

Module Initialization (src/fbgemm_xpu/__init__.py)

  • Import torch before loading _C extension to ensure PyTorch shared libraries are loaded in the process address space
  • Resolves undefined symbol errors when loading the compiled SYCL extension

FBGEMM Test Patch (packages/fbgemm-xpu/patches/0001-Add-XPU-support-to-fbgemm-tests.patch)

  • Enable XPU in the upstream block_bucketize_test.py so block_bucketize_sparse_features is covered by the FBGEMM sparse test suite

cc: @manuelhsantana, @aagalleg

@flezaalv flezaalv changed the title Flezaalv/feat/block bucketize sparse features operator Add block_bucketize_sparse_features operator Jun 25, 2026
@aagalleg
aagalleg force-pushed the flezaalv/feat/block_bucketize_sparse_features_operator branch from 2bcec61 to ecde2a4 Compare July 17, 2026 23:08
aagalleg and others added 27 commits August 5, 2026 18:22
Add complete test coverage for invert_permute operator on XPU
devices, covering correctness, validation, parity, and performance.

Test coverage includes:
- Correctness tests for int32/int64 with edge cases (empty, single
  element, identity, reverse, random permutations)
- Input validation tests for invalid dimensions and dtypes
- Meta function tests for torch.compile compatibility
- PyTorch opcheck validation for operator conventions
- Parametric tests with varying sizes (1 to 1M elements)
- CPU-XPU parity tests to ensure consistent results
- Performance benchmarks measuring execution time and bandwidth
Replace the custom standalone test_invert_permute.py with a git-am
patch applied to upstream FBGEMM v1.7.0 misc_ops_test.py, following the
torchcodec-xpu convention. The patch makes test_invert_permute run on
XPU (permute.xpu(), gated on torch.xpu.is_available()) and skips the
remaining operator tests that are not implemented on XPU.
…cture

- Fix test patches to match FBGEMM v1.8.0 tests.
- Move test patches from test/patches/ subdirectory to patches/ at the
package root level for better organization. Remove now-unnecessary
.gitkeep file and update patch with correct base commit reference.
Replaced test for upstream patched FBGEMM tests that enables testing XPU. This file is no longer needed.
Add SYCL port of FBGEMM's asynchronous_complete_cumsum operator for
Intel XPU devices. The operator computes a complete cumulative sum
with a leading zero (e.g., [a, b, c] → [0, a, a+b, a+b+c]).
Delete the accidentally tracked submodule reference to FBGEMM-v1.7.0.
Rename asynchronous_complete_cumsum files to sparse_async_cumsum.
Update 0001-Add-XPU-support-to-fbgemm-tests.patch to enable XPU testing
for asynchronous cumsum operators in cumsum_test.py:

- Add XPU device to test_cumsum (tests exclusive, inclusive, and
  complete cumsum)
- Add XPU device to test_asynchronous_complete_cumsum_2d
- Skip test_batched_complete_cumsum (operator not implemented on XPU)
Add SYCL infrastructure headers from intel/torch-xpu-ops/
to support advanced kernel implementations:
- DeviceProperties.h: Device capability queries and work group sizing
- SYCLContext.h: SYCL context management and namespace aliases
- SYCLHelpers.h: SYCL kernel submission and utility functions
- TensorInfo.h: Tensor metadata and dimension handling structures
- TensorOptions.h: Tensor configuration and options management
- Runtime.h: SYCL runtime utilities
- Macros.h: Common macro definitions
- Scalar.h: Scalar type conversion utilities

These headers provide the foundation for implementing 2D sparse data
permutation and other complex SYCL operations on XPU devices.
Add foundational utility headers and implementations to support
complex SYCL kernel operations:
- utils.h/cpp: Core constants, type definitions, kernel launch
  helpers, and device property queries
- dispatch_macros.h: Type dispatch macros for handling multiple
  data types (int32, int64, float, etc.)
- tensor_utils.h: Tensor manipulation and metadata utilities
- function_types.h: Symbol visibility definitions for shared
  library exports

These utilities provide essential infrastructure for implementing
2D sparse data permutation and other advanced operators on XPU
devices, including work group sizing, kernel launch helpers, and
type-safe dispatching mechanisms.
Add SYCL port of FBGEMM's permute_2D_sparse_data operator for
Intel XPU devices. This operator permutes 2D sparse data including
lengths [T, B], indices, and optional weights according to a
permutation vector, commonly used for reordering embedding table
features.

Implementation includes:
- SYCL kernels: permute_2D_lengths_kernel and permute_2D_data_kernel
- Host function: permute_2D_sparse_data_xpu
Integrate permute_2D_sparse_data operator into fbgemm-xpu:
- Add Python wrapper with type hints and documentation
- Register operator schema in torch library
- Include implementation files in CMake build (utils.cpp, SYCL
  kernels, and operator implementation)
Signed-off-by: Felipe Leza Alvarez <felipe.leza.alvarez@intel.com>
The permute_2d_sparse_data_op.cpp file was incorrectly emptied.
Restore the SYCL implementation.
…mpatible

Signed-off-by: Felipe Leza Alvarez <felipe.leza.alvarez@intel.com>
Signed-off-by: Felipe Leza Alvarez <felipe.leza.alvarez@intel.com>
…rators

Signed-off-by: Felipe Leza Alvarez <felipe.leza.alvarez@intel.com>
Signed-off-by: Felipe Leza Alvarez <felipe.leza.alvarez@intel.com>
@flezaalv
flezaalv force-pushed the flezaalv/feat/block_bucketize_sparse_features_operator branch from 0a1aaec to 672c921 Compare August 8, 2026 00:42
…tests

Integrate block_bucketize_sparse_features SYCL kernel implementation and
comprehensive test suite from experimentation_prs_integration branch.

- Add SYCL kernel implementation for block_bucketize_sparse_features
- Add block_bucketize_sparse_features_inference variant
- Add populate_bucketized_permute helper function
- Include comprehensive test suite with 18 test cases covering:
  * Variable bucket sizes and batch sizes
  * Long indices and keep_orig_idx modes
  * Total num blocks variations
  * Float64 weights support
  * Edge cases and error handling

All tests pass successfully on XPU hardware.
…t in rebase

Signed-off-by: Felipe Leza Alvarez <felipe.leza.alvarez@intel.com>
… coverage

Update block_bucketize_test.py hunk to use the accelerator_unavailable,
xpu_available, and fbgemm_xpu registration helpers introduced by earlier
patches (misc_ops_test.py, permute_sparse_features_test.py, test_utils.py).
Signed-off-by: Felipe Leza Alvarez <felipe.leza.alvarez@intel.com>
@aagalleg
aagalleg force-pushed the flezaalv/feat/block_bucketize_sparse_features_operator branch from 672c921 to 926a135 Compare August 10, 2026 18:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants