vllm
vllm copied to clipboard
[Misc] Update `gptq_marlin` to use new vLLMParameters
Summary
- Updates the
gptq_marlinparameters to usevLLMParametersto simplify linear layer weight loading - Updates to add
PackedColumnParameterto support packed parameters without row parallelism
👋 Hi! Thank you for contributing to the vLLM project.
Just a reminder: PRs would not trigger full CI run by default. Instead, it would only run fastcheck CI which consists a small and essential subset of CI tests to quickly catch errors. You can run other CI tests on top of default ones by unblocking the steps in your fast-check build on Buildkite UI.
Once the PR is approved and ready to go, please make sure to run full CI as it is required to merge (or just use auto-merge).
To run full CI, you can do one of these:
- Comment
/readyon the PR - Add
readylabel to the PR - Enable auto-merge.
🚀
/ready