Lily Liu
|
fe6d09ae61
|
[Minor] More fix of test_cache.py CI test failure (#2750)
|
2024-02-06 11:38:38 -08:00 |
|
liuyhwangyh
|
ed70c70ea3
|
modelscope: fix issue when model parameter is not a model id but path of the model. (#2489)
|
2024-02-06 09:57:15 -08:00 |
|
Woosuk Kwon
|
f0d4e14557
|
Add fused top-K softmax kernel for MoE (#2769)
|
2024-02-05 17:38:02 -08:00 |
|
Lukas
|
b92adec8e8
|
Set local logging level via env variable (#2774)
|
2024-02-05 14:26:50 -08:00 |
|
Rex
|
5a6c81b051
|
Remove eos tokens from output by default (#2611)
|
2024-02-04 14:32:42 -08:00 |
|
dancingpipi
|
51cd22ce56
|
set&get llm internal tokenizer instead of the TokenizerGroup (#2741)
Co-authored-by: shujunhua1 <shujunhua1@jd.com>
|
2024-02-04 14:25:36 -08:00 |
|
zspo
|
0e163fce18
|
Fix default length_penalty to 1.0 (#2667)
|
2024-02-01 15:59:39 -08:00 |
|
Kunshang Ji
|
96b6f475dd
|
Remove hardcoded device="cuda" to support more devices (#2503)
Co-authored-by: Jiang Li <jiang1.li@intel.com>
Co-authored-by: Kunshang Ji <kunshang.ji@intel.com>
|
2024-02-01 15:46:39 -08:00 |
|
Pernekhan Utemuratov
|
c410f5d020
|
Use revision when downloading the quantization config file (#2697)
Co-authored-by: Pernekhan Utemuratov <pernekhan@deepinfra.com>
|
2024-02-01 15:41:58 -08:00 |
|
Simon Mo
|
b9e96b17de
|
fix python 3.8 syntax (#2716)
|
2024-02-01 14:00:58 -08:00 |
|
Fengzhe Zhou
|
cd9e60c76c
|
Add Internlm2 (#2666)
|
2024-02-01 09:27:40 -08:00 |
|
Robert Shaw
|
93b38bea5d
|
Refactor Prometheus and Add Request Level Metrics (#2316)
|
2024-01-31 14:58:07 -08:00 |
|
Philipp Moritz
|
d0d93b92b1
|
Add unit test for Mixtral MoE layer (#2677)
|
2024-01-31 14:34:17 -08:00 |
|
zspo
|
c664b0e683
|
fix some bugs (#2689)
|
2024-01-31 10:09:23 -08:00 |
|
Tao He
|
d69ff0cbbb
|
Fixes assertion failure in prefix caching: the lora index mapping should respect prefix_len (#2688)
Signed-off-by: Tao He <sighingnow@gmail.com>
|
2024-01-31 18:00:13 +01:00 |
|
Zhuohan Li
|
1af090b57d
|
Bump up version to v0.3.0 (#2656)
|
2024-01-31 00:07:07 -08:00 |
|
Woosuk Kwon
|
3dad944485
|
Add quantized mixtral support (#2673)
|
2024-01-30 16:34:10 -08:00 |
|
Woosuk Kwon
|
105a40f53a
|
[Minor] Fix false warning when TP=1 (#2674)
|
2024-01-30 14:39:40 -08:00 |
|
Philipp Moritz
|
bbe9bd9684
|
[Minor] Fix a small typo (#2672)
|
2024-01-30 13:40:37 -08:00 |
|
Wen Sun
|
d79ced3292
|
Fix 'Actor methods cannot be called directly' when using --engine-use-ray (#2664)
* fix: engine-useray complain
* fix: typo
|
2024-01-30 17:17:05 +01:00 |
|
Philipp Moritz
|
ab40644669
|
Fused MOE for Mixtral (#2542)
Co-authored-by: chen shen <scv119@gmail.com>
|
2024-01-29 22:43:37 -08:00 |
|
wangding zeng
|
5d60def02c
|
DeepseekMoE support with Fused MoE kernel (#2453)
Co-authored-by: roy <jasonailu87@gmail.com>
|
2024-01-29 21:19:48 -08:00 |
|
zhaoyang-star
|
b72af8f1ed
|
Fix error when tp > 1 (#2644)
Co-authored-by: zhaoyang-star <zhao.yang16@zte.com.cn>
|
2024-01-28 22:47:39 -08:00 |
|
zhaoyang-star
|
9090bf02e7
|
Support FP8-E5M2 KV Cache (#2279)
Co-authored-by: zhaoyang <zhao.yang16@zte.com.cn>
Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
|
2024-01-28 16:43:54 -08:00 |
|
Murali Andoorveedu
|
89be30fa7d
|
Small async_llm_engine refactor (#2618)
|
2024-01-27 23:28:37 -08:00 |
|
Woosuk Kwon
|
5f036d2bcc
|
[Minor] Fix warning on Ray dependencies (#2630)
|
2024-01-27 15:43:40 -08:00 |
|
Hanzhi Zhou
|
380170038e
|
Implement custom all reduce kernels (#2192)
|
2024-01-27 12:46:35 -08:00 |
|
Xiang Xu
|
220a47627b
|
Use head_dim in config if exists (#2622)
|
2024-01-27 10:30:49 -08:00 |
|
Casper
|
beb89f68b4
|
AWQ: Up to 2.66x higher throughput (#2566)
|
2024-01-26 23:53:17 -08:00 |
|
Philipp Moritz
|
390b495ff3
|
Don't build punica kernels by default (#2605)
|
2024-01-26 15:19:19 -08:00 |
|
dakotamahan-stability
|
3a0e1fc070
|
Support for Stable LM 2 (#2598)
Co-authored-by: Zhuohan Li <zhuohan123@gmail.com>
|
2024-01-26 12:45:19 -08:00 |
|
Hongxia Yang
|
6b7de1a030
|
[ROCm] add support to ROCm 6.0 and MI300 (#2274)
|
2024-01-26 12:41:10 -08:00 |
|
Junyang Lin
|
2832e7b9f9
|
fix names and license for Qwen2 (#2589)
|
2024-01-24 22:37:51 -08:00 |
|
Simon Mo
|
3a7dd7e367
|
Support Batch Completion in Server (#2529)
|
2024-01-24 17:11:07 -08:00 |
|
Federico Galatolo
|
f1f6cc10c7
|
Added include_stop_str_in_output and length_penalty parameters to OpenAI API (#2562)
|
2024-01-24 10:21:56 -08:00 |
|
Nikola Borisov
|
3209b49033
|
[Bugfix] fix crash if max_tokens=None (#2570)
|
2024-01-23 22:38:55 -08:00 |
|
Antoni Baum
|
9b945daaf1
|
[Experimental] Add multi-LoRA support (#1804)
Co-authored-by: Chen Shen <scv119@gmail.com>
Co-authored-by: Shreyas Krishnaswamy <shrekris@anyscale.com>
Co-authored-by: Avnish Narayan <avnish@anyscale.com>
|
2024-01-23 15:26:37 -08:00 |
|
Erfan Al-Hossami
|
9c1352eb57
|
[Feature] Simple API token authentication and pluggable middlewares (#1106)
|
2024-01-23 15:13:00 -08:00 |
|
Junyang Lin
|
94b5edeb53
|
Add qwen2 (#2495)
|
2024-01-22 14:34:21 -08:00 |
|
Philipp Moritz
|
ab7e6006d6
|
Fix https://github.com/vllm-project/vllm/issues/2540 (#2545)
|
2024-01-22 19:02:38 +01:00 |
|
Cade Daniel
|
18bfcdd05c
|
[Speculative decoding 2/9] Multi-step worker for draft model (#2424)
|
2024-01-21 16:31:47 -08:00 |
|
Jannis Schönleber
|
71d63ed72e
|
migrate pydantic from v1 to v2 (#2531)
|
2024-01-21 16:05:56 -08:00 |
|
Nick Hill
|
d75c40734a
|
[Fix] Keep scheduler.running as deque (#2523)
|
2024-01-20 22:36:09 -08:00 |
|
Junda Chen
|
5b23c3f26f
|
Add group as an argument in broadcast ops (#2522)
|
2024-01-20 16:00:26 -08:00 |
|
Roy
|
91a61da9b1
|
[Bugfix] fix load local safetensors model (#2512)
|
2024-01-19 16:26:16 -08:00 |
|
Zhuohan Li
|
ef9b636e2d
|
Simplify broadcast logic for control messages (#2501)
|
2024-01-19 11:23:30 -08:00 |
|
Simon Mo
|
dd7e8f5f64
|
refactor complemention api for readability (#2499)
|
2024-01-18 16:45:14 -08:00 |
|
ljss
|
d2a68364c4
|
[BugFix] Fix abort_seq_group (#2463)
|
2024-01-18 15:10:42 -08:00 |
|
Nikola Borisov
|
7e1081139d
|
Don't download both safetensor and bin files. (#2480)
|
2024-01-18 11:05:53 -08:00 |
|
Liangfu Chen
|
18473cf498
|
[Neuron] Add an option to build with neuron (#2065)
|
2024-01-18 10:58:50 -08:00 |
|