composable_kernel
composable_kernel copied to clipboard
[CK] Added s_prefetch unit test.
-added s_buffer_load_b32/64 assembly -added amd_s_buffer_load_impl
Proposed changes
Please describe the motivation behind the pull request, whether it enables a new feature or fixes a bug. If there are associated pull requests or issues, please link them to the pull request.
Checklist
Please put an x into the boxes that apply. You can also fill these out after creating the PR. If you're not sure, please don't hesitate to ask.
- [ ] I have added tests relevant to the introduced functionality, and the unit tests are passing locally
- [ ] I have added the test to REGRESSION_TESTS list defined at the top of CMakeLists.txt in tests/CMakeLists.txt, IF the test takes more than 30 seconds to run.
- [ ] I have added inline documentation which enables the maintainers with understanding the motivation
- [ ] I have removed the stale documentation which is no longer relevant after this pull request
- [ ] (If this change is user-facing) I have added release notes which provide the end users with a brief summary of the improvement from this pull request
- [x] I have run
clang-formaton all changed files - [x] Any dependent changes have been merged
Discussion
IS IT A GOOD PLACE TO PUT ASM? OR SHOULD IT BE PUT IN AMD_BUFFER_ADDRESSING.HPP to avoid #include AMD_INLINE_ASM.HPP in AMD_BUFFER_ADDRESSING.HPP?