composable_kernel icon indicating copy to clipboard operation
composable_kernel copied to clipboard

(do not merge) use clang builtin for compile-time sequence indexing

Open tenpercent opened this issue 8 months ago • 0 comments

Proposed changes

As titled

Checklist

Please put an x into the boxes that apply. You can also fill these out after creating the PR. If you're not sure, please don't hesitate to ask.

  • [ ] I have added tests relevant to the introduced functionality, and the unit tests are passing locally
  • [ ] I have added the test to REGRESSION_TESTS list defined at the top of CMakeLists.txt in tests/CMakeLists.txt, IF the test takes more than 30 seconds to run.
  • [x] I have added inline documentation which enables the maintainers with understanding the motivation
  • [ ] I have removed the stale documentation which is no longer relevant after this pull request
  • [ ] (If this change is user-facing) I have added release notes which provide the end users with a brief summary of the improvement from this pull request
  • [x] I have run clang-format on all changed files
  • [ ] Any dependent changes have been merged

Discussion

Local benchmarking shows that non-builtin version is faster

tenpercent avatar Mar 20 '25 21:03 tenpercent