vllm/quantization at 260d119e864edbf023b1be7fa446a08bbea11f80 - vllm

History

Tyler Michael Smith 260d119e86 [Kernel] Refactor CUTLASS kernels to always take scales that reside on the GPU (#5137 )		2024-06-01 06:45:32 +00:00
..
aqlm	[CI/Build] Enforce style for C++ and CUDA code with `clang-format` (#4722 )	2024-05-22 07:18:41 +00:00
awq	[CI/Build] Enforce style for C++ and CUDA code with `clang-format` (#4722 )	2024-05-22 07:18:41 +00:00
compressed_tensors	[Kernel] Initial Activation Quantization Support (#4525 )	2024-05-23 21:29:18 +00:00
cutlass_w8a8	[Kernel] Refactor CUTLASS kernels to always take scales that reside on the GPU (#5137 )	2024-06-01 06:45:32 +00:00
fp8	[CI/Build] Enforce style for C++ and CUDA code with `clang-format` (#4722 )	2024-05-22 07:18:41 +00:00
gptq	[CI/Build] Enforce style for C++ and CUDA code with `clang-format` (#4722 )	2024-05-22 07:18:41 +00:00
gptq_marlin	[CI/Build] Enforce style for C++ and CUDA code with `clang-format` (#4722 )	2024-05-22 07:18:41 +00:00
marlin	Revert "[Kernel] Marlin_24: Ensure the mma.sp instruction is using the ::ordered_metadata modifier (introduced with PTX 8.5)" (#5149 )	2024-05-30 22:00:26 -07:00
squeezellm	[CI/Build] Enforce style for C++ and CUDA code with `clang-format` (#4722 )	2024-05-22 07:18:41 +00:00