vllm/csrc at 07feecde1a69859d565786a7ad64c0f604f17b28 - vllm

sergey-tinkoff 07feecde1a [Model] LoRA support added for command-r (#5178 )	2024-06-18 11:01:21 -07:00
..
attention	[Kernel][Misc] Use TORCH_LIBRARY instead of PYBIND11_MODULE for custom ops (#5047 )	2024-06-09 16:23:30 -04:00
cpu	[Hardware][Intel] Support CPU inference with AVX2 ISA (#5452 )	2024-06-13 17:22:24 -06:00
moe	[Kernel][Misc] Use TORCH_LIBRARY instead of PYBIND11_MODULE for custom ops (#5047 )	2024-06-09 16:23:30 -04:00
punica	[Model] LoRA support added for command-r (#5178 )	2024-06-18 11:01:21 -07:00
quantization	[Kernel] Suppress mma.sp warning on CUDA 12.5 and later (#5401 )	2024-06-14 10:02:00 -07:00
activation_kernels.cu	[Kernel][Misc] Use TORCH_LIBRARY instead of PYBIND11_MODULE for custom ops (#5047 )	2024-06-09 16:23:30 -04:00
cache_kernels.cu	[Kernel][Misc] Use TORCH_LIBRARY instead of PYBIND11_MODULE for custom ops (#5047 )	2024-06-09 16:23:30 -04:00
cache.h	[Kernel][Misc] Use TORCH_LIBRARY instead of PYBIND11_MODULE for custom ops (#5047 )	2024-06-09 16:23:30 -04:00
cuda_compat.h	[Kernel][ROCm][AMD] enable fused topk_softmax kernel for moe layer (#4927 )	2024-06-02 14:13:26 -07:00
cuda_utils_kernels.cu	[Kernel][Misc] Use TORCH_LIBRARY instead of PYBIND11_MODULE for custom ops (#5047 )	2024-06-09 16:23:30 -04:00
cuda_utils.h	[Kernel][Misc] Use TORCH_LIBRARY instead of PYBIND11_MODULE for custom ops (#5047 )	2024-06-09 16:23:30 -04:00
custom_all_reduce_test.cu	[CI/Build] Enforce style for C++ and CUDA code with `clang-format` (#4722 )	2024-05-22 07:18:41 +00:00
custom_all_reduce.cu	[Kernel][Misc] Use TORCH_LIBRARY instead of PYBIND11_MODULE for custom ops (#5047 )	2024-06-09 16:23:30 -04:00
custom_all_reduce.cuh	[CI/Build] Enforce style for C++ and CUDA code with `clang-format` (#4722 )	2024-05-22 07:18:41 +00:00
dispatch_utils.h	[Kernel][Misc] Use TORCH_LIBRARY instead of PYBIND11_MODULE for custom ops (#5047 )	2024-06-09 16:23:30 -04:00
layernorm_kernels.cu	[Kernel][Misc] Use TORCH_LIBRARY instead of PYBIND11_MODULE for custom ops (#5047 )	2024-06-09 16:23:30 -04:00
moe_align_block_size_kernels.cu	[Kernel][Misc] Use TORCH_LIBRARY instead of PYBIND11_MODULE for custom ops (#5047 )	2024-06-09 16:23:30 -04:00
ops.h	[Kernel] Factor out epilogues from cutlass kernels (#5391 )	2024-06-13 11:22:19 -07:00
pos_encoding_kernels.cu	[Kernel][Misc] Use TORCH_LIBRARY instead of PYBIND11_MODULE for custom ops (#5047 )	2024-06-09 16:23:30 -04:00
reduction_utils.cuh	[Kernel] Dynamic Per-Token Activation Quantization (#5037 )	2024-06-07 09:36:26 -07:00
registration.h	[Kernel][Misc] Use TORCH_LIBRARY instead of PYBIND11_MODULE for custom ops (#5047 )	2024-06-09 16:23:30 -04:00
torch_bindings.cpp	[Kernel] Factor out epilogues from cutlass kernels (#5391 )	2024-06-13 11:22:19 -07:00