Woosuk Kwon
|
76a7983b23
|
[BugFix] Fix RoPE kernel on long sequences(#2164)
|
2023-12-17 17:09:10 -08:00 |
|
Woosuk Kwon
|
8041b7305e
|
[BugFix] Raise error when max_model_len is larger than KV cache (#2163)
|
2023-12-17 17:08:23 -08:00 |
|
Suhong Moon
|
3ec8c25cd0
|
[Docs] Update documentation for gpu-memory-utilization option (#2162)
|
2023-12-17 10:51:57 -08:00 |
|
Woosuk Kwon
|
671af2b1c0
|
Bump up to v0.2.6 (#2157)
|
2023-12-17 10:34:56 -08:00 |
|
Woosuk Kwon
|
6f41f0e377
|
Disable CUDA graph for SqueezeLLM (#2161)
|
2023-12-17 10:24:25 -08:00 |
|
Woosuk Kwon
|
2c9b638065
|
[Minor] Fix a typo in .pt weight support (#2160)
|
2023-12-17 10:12:44 -08:00 |
|
Antoni Baum
|
a7347d9a6d
|
Make sampler less blocking (#1889)
|
2023-12-17 23:03:49 +08:00 |
|
Woosuk Kwon
|
f8c688d746
|
[Minor] Add Phi 2 to supported models (#2159)
|
2023-12-17 02:54:57 -08:00 |
|
Woosuk Kwon
|
c9fadda543
|
[Minor] Fix xformers version (#2158)
|
2023-12-17 02:28:02 -08:00 |
|
Woosuk Kwon
|
30fb0956df
|
[Minor] Add more detailed explanation on quantization argument (#2145)
|
2023-12-17 01:56:16 -08:00 |
|
Woosuk Kwon
|
3a765bd5e1
|
Temporarily enforce eager mode for GPTQ models (#2154)
|
2023-12-17 01:51:12 -08:00 |
|
Woosuk Kwon
|
26c52a5ea6
|
[Docs] Add CUDA graph support to docs (#2148)
|
2023-12-17 01:49:20 -08:00 |
|
Woosuk Kwon
|
c3372e87be
|
Remove dependency on CuPy (#2152)
|
2023-12-17 01:49:07 -08:00 |
|
Woosuk Kwon
|
b0a1d667b0
|
Pin PyTorch & xformers versions (#2155)
|
2023-12-17 01:46:54 -08:00 |
|
Woosuk Kwon
|
e1d5402238
|
Fix all-reduce memory usage (#2151)
|
2023-12-17 01:44:45 -08:00 |
|
Woosuk Kwon
|
3d1cfbfc74
|
[Minor] Delete Llama tokenizer warnings (#2146)
|
2023-12-16 22:05:18 -08:00 |
|
Woosuk Kwon
|
37ca558103
|
Optimize model execution with CUDA graph (#1926)
Co-authored-by: Chen Shen <scv119@gmail.com>
Co-authored-by: Antoni Baum <antoni.baum@protonmail.com>
|
2023-12-16 21:12:08 -08:00 |
|
Roy
|
eed74a558f
|
Simplify weight loading logic (#2133)
|
2023-12-16 12:41:23 -08:00 |
|
Woosuk Kwon
|
2acd76f346
|
[ROCm] Temporarily remove GPTQ ROCm support (#2138)
|
2023-12-15 17:13:58 -08:00 |
|
Woosuk Kwon
|
b81a6a6bb3
|
[Docs] Add supported quantization methods to docs (#2135)
|
2023-12-15 13:29:22 -08:00 |
|
CHU Tianxiang
|
0fbfc4b81b
|
Add GPTQ support (#916)
|
2023-12-15 03:04:22 -08:00 |
|
Yunfeng Bai
|
c06170cc8e
|
Add a flag to include stop string in output text (#1976)
|
2023-12-15 00:45:58 -08:00 |
|
Mingcan Xiang
|
614856da25
|
Avoid multiple redefinition (#1817)
|
2023-12-14 09:35:58 -08:00 |
|
TJian
|
05bdf4eaf3
|
Fix Dockerfile.rocm (#2101)
Co-authored-by: miloice <jeffaw99@hotmail.com>
|
2023-12-14 00:45:58 -08:00 |
|
mezuzza
|
6774bd50b0
|
Fix typing in AsyncLLMEngine & add toml to requirements-dev (#2100)
|
2023-12-14 00:19:41 -08:00 |
|
Woosuk Kwon
|
31c1f3255e
|
Bump up to v0.2.5 (#2095)
|
2023-12-13 23:56:15 -08:00 |
|
Antoni Baum
|
21d93c140d
|
Optimize Mixtral with expert parallelism (#2090)
|
2023-12-13 23:55:07 -08:00 |
|
Woosuk Kwon
|
f1c8520146
|
[BugFix] Fix input positions for long context with sliding window (#2088)
|
2023-12-13 12:28:13 -08:00 |
|
Woosuk Kwon
|
096827c284
|
[Docs] Add notes on ROCm-supported models (#2087)
|
2023-12-13 09:45:34 -08:00 |
|
Woosuk Kwon
|
6565d9e33e
|
Update installation instruction for vLLM + CUDA 11.8 (#2086)
|
2023-12-13 09:25:59 -08:00 |
|
TJian
|
f375ec8440
|
[ROCm] Upgrade xformers version for ROCm & update doc (#2079)
Co-authored-by: miloice <jeffaw99@hotmail.com>
|
2023-12-13 00:56:05 -08:00 |
|
Woosuk Kwon
|
518369d78c
|
Implement lazy model loader (#2044)
|
2023-12-12 22:21:45 -08:00 |
|
Woosuk Kwon
|
30bad5c492
|
Fix peak memory profiling (#2031)
|
2023-12-12 22:01:53 -08:00 |
|
Simon Mo
|
3fefe271ec
|
Update Dockerfile to build Megablocks (#2042)
|
2023-12-12 17:34:17 -08:00 |
|
Megha Agarwal
|
6428f1d051
|
Support MPT with GQA (#1938)
Co-authored-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>
|
2023-12-12 10:16:05 -08:00 |
|
Woosuk Kwon
|
7e1b21daac
|
Remove einops from requirements (#2049)
|
2023-12-12 09:34:09 -08:00 |
|
Woosuk Kwon
|
cb3f30c600
|
Upgrade transformers version to 4.36.0 (#2046)
|
2023-12-11 18:39:14 -08:00 |
|
Woosuk Kwon
|
f3e024bece
|
[CI/CD] Upgrade PyTorch version to v2.1.1 (#2045)
|
2023-12-11 17:48:11 -08:00 |
|
Woosuk Kwon
|
31d2ab4aff
|
Remove python 3.10 requirement (#2040)
|
2023-12-11 12:26:42 -08:00 |
|
Simon Mo
|
eb17212858
|
Update Dockerfile to support Mixtral (#2027)
|
2023-12-11 11:59:08 -08:00 |
|
Woosuk Kwon
|
4dd4b5c538
|
Bump up to v0.2.4 (#2034)
|
2023-12-11 11:49:39 -08:00 |
|
Woosuk Kwon
|
6120e5aaea
|
Fix import error msg for megablocks (#2038)
|
2023-12-11 11:40:56 -08:00 |
|
Ram
|
2eaa81b236
|
Update README.md to add megablocks requirement for mixtral (#2033)
|
2023-12-11 11:37:34 -08:00 |
|
Woosuk Kwon
|
81ce2a4b26
|
[Minor] Fix type annotation in Mixtral (#2036)
|
2023-12-11 11:32:39 -08:00 |
|
Woosuk Kwon
|
5dd80d3777
|
Fix latency benchmark script (#2035)
|
2023-12-11 11:19:08 -08:00 |
|
Woosuk Kwon
|
beeee69bc9
|
Revert adding Megablocks (#2030)
|
2023-12-11 10:49:00 -08:00 |
|
Ram
|
9bf28d0b69
|
Update requirements.txt for mixtral (#2029)
|
2023-12-11 10:39:29 -08:00 |
|
Ikko Eltociear Ashimine
|
c0ce15dfb2
|
Update run_on_sky.rst (#2025)
sharable -> shareable
|
2023-12-11 10:32:58 -08:00 |
|
Woosuk Kwon
|
b9bcdc7158
|
Change the load format to pt for Mixtral (#2028)
|
2023-12-11 10:32:17 -08:00 |
|
Woosuk Kwon
|
4ff0203987
|
Minor fixes for Mixtral (#2015)
|
2023-12-11 09:16:15 -08:00 |
|