Woosuk Kwon
|
4089985552
|
[V1] Integrate Piecewise CUDA graphs (#10058)
Signed-off-by: Woosuk Kwon <woosuk.kwon@berkeley.edu>
|
2024-11-05 22:16:04 -08:00 |
|
zifeitong
|
9d59b75593
|
[Bugfix] Remove CustomChatCompletionContentPartParam multimodal input type (#10054)
Signed-off-by: Zifei Tong <zifeitong@gmail.com>
|
2024-11-06 05:13:09 +00:00 |
|
arakowsk-amd
|
ea928f608c
|
[Bugfix] Gpt-j-6B patch kv_scale to k_scale path (#10063)
Signed-off-by: Alex Rakowski <alex.rakowski@amd.com>
Signed-off-by: Alex Rakowski <182798202+arakowsk-amd@users.noreply.github.com>
|
2024-11-06 05:10:40 +00:00 |
|
Travis Johnson
|
2bcbae704c
|
[Bugfix] Fix edge-case crash when using chat with the Mistral Tekken Tokenizer (#10051)
Signed-off-by: Travis Johnson <tsjohnso@us.ibm.com>
|
2024-11-06 04:28:29 +00:00 |
|
Peter Salas
|
ffc0f2b47a
|
[Model][OpenVINO] Fix regressions from #8346 (#10045)
Signed-off-by: Peter Salas <peter@fixie.ai>
|
2024-11-06 04:19:15 +00:00 |
|
Cyrus Leung
|
82bfc38d07
|
[Misc] Sort the list of embedding models (#10037)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2024-11-06 04:05:05 +00:00 |
|
youkaichao
|
c4cacbaa7f
|
[v1] reduce graph capture time for piecewise cudagraph (#10059)
Signed-off-by: youkaichao <youkaichao@gmail.com>
|
2024-11-05 18:19:50 -08:00 |
|
Sungjae Lee
|
0c63c34f72
|
[Bugfix][SpecDecode] kv corruption with bonus tokens in spec decode (#9730)
Co-authored-by: LiuXiaoxuanPKU <lilyliupku@gmail.com>
|
2024-11-06 01:45:45 +00:00 |
|
Wallas Henrique
|
966e31697b
|
[Bugfix] Fix pickle of input when async output processing is on (#9931)
Signed-off-by: Wallas Santos <wallashss@ibm.com>
|
2024-11-06 00:39:26 +00:00 |
|
zifeitong
|
43300bd98a
|
[Bugfix] Properly propagate trust_remote_code settings (#10047)
Signed-off-by: Zifei Tong <zifeitong@gmail.com>
|
2024-11-05 16:34:40 -08:00 |
|
youkaichao
|
ca9844b340
|
[bugfix] fix weak ref in piecewise cudagraph and tractable test (#10048)
Signed-off-by: youkaichao <youkaichao@gmail.com>
|
2024-11-05 14:49:20 -08:00 |
|
Michael Goin
|
235366fe2e
|
[CI] Prune back the number of tests in tests/kernels/* (#9932)
Signed-off-by: mgoin <michael@neuralmagic.com>
|
2024-11-05 16:02:32 -05:00 |
|
Michael Goin
|
02462465ea
|
[CI] Prune tests/models/decoder_only/language/* tests (#9940)
Signed-off-by: mgoin <michael@neuralmagic.com>
|
2024-11-05 16:02:23 -05:00 |
|
Jee Jee Li
|
b9c64c0ca7
|
[Misc] Modify BNB parameter name (#9997)
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>
|
2024-11-05 14:40:08 -05:00 |
|
lkchen
|
d2e80332a7
|
[Feature] Update benchmark_throughput.py to support image input (#9851)
Signed-off-by: Linkun Chen <github+anyscale@lkchen.net>
Co-authored-by: Linkun Chen <github+anyscale@lkchen.net>
|
2024-11-05 19:30:02 +00:00 |
|
Michael Goin
|
a53046b16f
|
[Model] Support quantization of PixtralHFTransformer for PixtralHF (#9921)
Signed-off-by: mgoin <michael@neuralmagic.com>
|
2024-11-05 10:42:20 -08:00 |
|
Russell Bryant
|
731aec5be7
|
[CI/Build] Limit github CI jobs based on files changed (#9928)
Signed-off-by: Russell Bryant <rbryant@redhat.com>
|
2024-11-05 10:30:42 -08:00 |
|
Chenghao (Alan) Yang
|
09d3550372
|
[Misc] Add logging for CUDA memory (#10027)
Signed-off-by: Chenghao Yang <yangalan1996@gmail.com>
Signed-off-by: youkaichao <youkaichao@gmail.com>
Co-authored-by: Chenghao Yang <yangalan1996@gmail.com>
Co-authored-by: youkaichao <youkaichao@gmail.com>
|
2024-11-05 09:50:50 -08:00 |
|
Richard Liu
|
cd34029e91
|
Refactor TPU requirements file and pin build dependencies (#10010)
Signed-off-by: Richard Liu <ricliu@google.com>
|
2024-11-05 16:48:44 +00:00 |
|
Russell Bryant
|
5952d81139
|
[Frontend] Fix tcp port reservation for api server (#10012)
Signed-off-by: Russell Bryant <rbryant@redhat.com>
|
2024-11-05 07:50:57 -08:00 |
|
Chauncey
|
93dee88f6b
|
[Misc] vllm CLI flags should be ordered for better user readability (#10017)
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>
|
2024-11-05 18:59:56 +08:00 |
|
Gene Der Su
|
7a83b1aec0
|
[BugFix] Lazy import ray (#10021)
|
2024-11-05 10:04:10 +00:00 |
|
Tyler Michael Smith
|
ad23318928
|
[Bugfix] Fixup Mamba (#10004)
Signed-off-by: Tyler Michael Smith <tyler@neuralmagic.com>
|
2024-11-05 03:46:38 +00:00 |
|
Cyrus Leung
|
bbc3619dc8
|
[Core] Make encoder-decoder inputs a nested structure to be more composable (#9604)
Signed-off-by: DarkLight1337 <tlleungac@connect.ust.hk>
|
2024-11-05 10:07:31 +08:00 |
|
Tyler Michael Smith
|
04bbf38e05
|
[Core] Use os.sched_yield in ShmRingBuffer instead of time.sleep (#9994)
Signed-off-by: Tyler Michael Smith <tyler@neuralmagic.com>
|
2024-11-05 01:08:21 +00:00 |
|
Michael Goin
|
8f0a9ca890
|
[Bugfix] Respect modules_to_not_convert within awq_marlin (#9895)
Signed-off-by: mgoin <michael@neuralmagic.com>
|
2024-11-04 16:57:44 -07:00 |
|
youkaichao
|
2094062b4e
|
[4.5/N] bugfix for quant config in speculative decode (#10007)
Signed-off-by: youkaichao <youkaichao@gmail.com>
|
2024-11-04 15:11:59 -08:00 |
|
bnellnm
|
d93478b399
|
[Bugfix] Upgrade to pytorch 2.5.1 (#10001)
Signed-off-by: Bill Nell <bill@neuralmagic.com>
|
2024-11-04 15:11:28 -08:00 |
|
tomeras91
|
ac04a97a9f
|
[Frontend] Add max_tokens prometheus metric (#9881)
Signed-off-by: Tomer Asida <tomera@ai21.com>
|
2024-11-04 22:53:24 +00:00 |
|
lkchen
|
9a5664d4a4
|
[Misc] Refactor benchmark_throughput.py (#9779)
Signed-off-by: Linkun Chen <github+anyscale@lkchen.net>
Co-authored-by: Linkun Chen <lkchen@github.com>
Co-authored-by: Linkun Chen <github+anyscale@lkchen.net>
|
2024-11-04 14:32:16 -08:00 |
|
Robert Shaw
|
04cef2c6ab
|
[Bugfix] Fix MQLLMEngine hanging (#9973)
Signed-off-by: rshaw@neuralmagic.com <rshaw@neuralmagic.com>
|
2024-11-04 16:01:43 -05:00 |
|
Roger Wang
|
6e056bcf04
|
[Doc] Update VLM doc about loading from local files (#9999)
Signed-off-by: Roger Wang <ywang@roblox.com>
|
2024-11-04 19:47:11 +00:00 |
|
hissu-hyvarinen
|
5208dc7a20
|
[Bugfix][CI/Build][Hardware][AMD] Shard ID parameters in AMD tests running parallel jobs (#9279)
Signed-off-by: Hissu Hyvarinen <hissu.hyvarinen@amd.com>
|
2024-11-04 11:37:46 -08:00 |
|
Robert Shaw
|
1c45f4c385
|
[CI] Basic Integration Test For TPU (#9968)
Signed-off-by: Robert Shaw <rshaw@neuralmagic.com>
|
2024-11-04 11:34:26 -08:00 |
|
Mor Zusman
|
603a661ae8
|
[Model] factoring out MambaMixer out of Jamba (#8993)
Signed-off-by: mzusman <mor.zusmann@gmail.com>
|
2024-11-04 18:00:00 +00:00 |
|
Jee Jee Li
|
fb2716d641
|
[Misc]Reduce BNB static variable (#9987)
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>
|
2024-11-04 17:04:40 +00:00 |
|
youkaichao
|
8d72bb20fa
|
[4/N] make quant config first-class citizen (#9978)
Signed-off-by: youkaichao <youkaichao@gmail.com>
|
2024-11-04 08:51:31 -08:00 |
|
Chauncey
|
ac6b8f19b9
|
[Frontend] Multi-Modality Support for Loading Local Image Files (#9915)
Signed-off-by: chaunceyjiang <chaunceyjiang@gmail.com>
|
2024-11-04 15:34:57 +00:00 |
|
Mengqing Cao
|
ccb5376a9a
|
[Bugfix][OpenVINO] Fix circular reference #9939 (#9974)
Signed-off-by: MengqingCao <cmq0113@163.com>
|
2024-11-04 18:14:13 +08:00 |
|
Tran Quang Dai
|
ea4adeddc1
|
[Bugfix] Fix E2EL mean and median stats (#9984)
Signed-off-by: daitran2k1 <tranquangdai7a@gmail.com>
|
2024-11-04 09:37:58 +00:00 |
|
Yang Zheng
|
4dbcbbeb09
|
[Misc] Compute query_start_loc/seq_start_loc on CPU (#9447)
Co-authored-by: Yang Zheng(SW)(Alex) <you@example.com>
|
2024-11-04 08:54:37 +00:00 |
|
Gregory Shtrasberg
|
b67feb1274
|
[Bugfix]Using the correct type hints (#9885)
Signed-off-by: Gregory Shtrasberg <Gregory.Shtrasberg@amd.com>
|
2024-11-04 06:19:51 +00:00 |
|
Jee Jee Li
|
c49f0407ba
|
[Bugfix] Fix MiniCPMV and Mllama BNB bug (#9917)
Signed-off-by: Jee Jee Li <pandaleefree@gmail.com>
|
2024-11-04 03:36:41 +00:00 |
|
Robert Shaw
|
91c9ebbb1b
|
[V1] Fix Configs (#9971)
|
2024-11-04 00:24:40 +00:00 |
|
shanshan wang
|
54597724f4
|
[Model] Add support for H2OVL-Mississippi models (#9747)
Signed-off-by: Shanshan Wang <shanshan.wang@h2o.ai>
Signed-off-by: Roger Wang <ywang@roblox.com>
Co-authored-by: Roger Wang <ywang@roblox.com>
|
2024-11-04 00:15:36 +00:00 |
|
Nick Hill
|
1f1b6d6eda
|
[V1] Support per-request seed (#9945)
Signed-off-by: Nick Hill <nickhill@us.ibm.com>
|
2024-11-03 09:14:17 -08:00 |
|
youkaichao
|
3bb4befea7
|
[bugfix] fix tsts (#9959)
Signed-off-by: youkaichao <youkaichao@gmail.com>
|
2024-11-02 15:54:05 -07:00 |
|
Yongzao
|
ae5279a163
|
[torch.compile] Adding torch compile to vision-language models (#9946)
|
2024-11-02 12:56:05 -07:00 |
|
Nikita Furin
|
1b73ab2a1f
|
[CI/Build] Quoting around > (#9956)
|
2024-11-02 12:50:28 -07:00 |
|
youkaichao
|
cea808f325
|
[3/N] model runner pass the whole config to model (#9958)
Signed-off-by: youkaichao <youkaichao@gmail.com>
|
2024-11-02 12:08:49 -07:00 |
|