vllm/quantization at a1242324c99ff8b1e29981006dfb504da198c7c3 - vllm

History

Dipika Sikka a1242324c9 [Kernel] Initial Activation Quantization Support (#4525 ) Co-authored-by: Varun Sundar Rabindranath <varunsundar08@gmail.com> Co-authored-by: Varun Sundar Rabindranath <varun@neuralmagic.com>		2024-05-23 21:29:18 +00:00
..
aqlm	[CI/Build] Enforce style for C++ and CUDA code with `clang-format` (#4722 )	2024-05-22 07:18:41 +00:00
awq	[CI/Build] Enforce style for C++ and CUDA code with `clang-format` (#4722 )	2024-05-22 07:18:41 +00:00
compressed_tensors	[Kernel] Initial Activation Quantization Support (#4525 )	2024-05-23 21:29:18 +00:00
cutlass_w8a8	[Kernel] Fixup for CUTLASS kernels in CUDA graphs (#4954 )	2024-05-22 14:10:43 +00:00
fp8	[CI/Build] Enforce style for C++ and CUDA code with `clang-format` (#4722 )	2024-05-22 07:18:41 +00:00
gptq	[CI/Build] Enforce style for C++ and CUDA code with `clang-format` (#4722 )	2024-05-22 07:18:41 +00:00
gptq_marlin	[CI/Build] Enforce style for C++ and CUDA code with `clang-format` (#4722 )	2024-05-22 07:18:41 +00:00
marlin	Marlin 24 prefill performance improvement (about 25% better on average) (#4983 )	2024-05-23 02:39:27 -04:00
squeezellm	[CI/Build] Enforce style for C++ and CUDA code with `clang-format` (#4722 )	2024-05-22 07:18:41 +00:00