vLLM releases
59 articles archived · Visit source ↗ · RSS
-
vLLM releases dev-tools 9d ago
v0.30.0rc2
[Bugfix][NIXL] Avoid receive reports for notification-only requests (…
22 -
vLLM releases dev-tools 12d ago
proto-v0.2.0: vllm-proto 0.2.0
Validated by PR #56538 CI at fa2a26f .
23 -
vLLM releases dev-tools 15d ago
v0.29.1rc0
[watermarking] Dual-key gumbel-max watermarking for speculative decod…
20 -
vLLM releases dev-tools 20d ago
v0.29.0rc6
[Bugfix][Core] Apply dense prefix cache default to hybrid models ( #55 …
36 -
vLLM releases dev-tools 20d ago
v0.29.0
[Bugfix][Core] Apply dense prefix cache default to hybrid models ( #55 …
37 -
vLLM releases dev-tools 20d ago
v0.29.0rc5
[Core] Default prefix_cache_retention_interval to dense for Mamba + E…
23 -
vLLM releases dev-tools 23d ago
v0.29.0rc4: [Bugfix] Avoid sync in TRT-LLM ragged prefill
Generated-by: Codex [email protected] Signed-off-by: Codex [email protected]
16 -
vLLM releases dev-tools 24d ago
v0.29.0rc3
[CI] Remove deleted nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF1…
33 -
vLLM releases dev-tools 25d ago
v0.29.0rc2
[Bugfix][Multimodal] Handle prefix-covered items in SHM worker cache …
9 -
vLLM releases dev-tools 1mo ago
v0.28.1rc0
[Tools][Recipes] Improve sweep recommendations and short-alias parsin…
7 -
vLLM releases dev-tools 1mo ago
v0.28.0rc1
[kv_offload] fix(metrics): rename kv_offload_tiering_block_{queries,h…
23 -
vLLM releases dev-tools 1mo ago
v0.27.2rc0: [Spec Decode] DSpark confidence-scheduled verification (#47808)
Signed-off-by: Lucas Wilkinson [email protected] Signed-off-by: Lucas Wilkinson [email protected] Signed-off-by: Benjamin Chislett [email protected] Signed-off-by: Lucas Wilkinson [email protected] Signed-off-by: Nick Hill…
26 -
vLLM releases dev-tools 1mo ago
v0.27.1: [CI] Limit Arctic import check to x86 test images
The arm64 test lockfile intentionally omits arctic-inference, so only validate its native extension on platforms where the package is installed. Co-authored-by: OpenAI Codex [email protected] Signed-off-by: khluu [email protected]
19 -
vLLM releases dev-tools 2mo ago
v0.26.1rc0
[CI][ROCm] Fix test_ocp_mx_wikitext_correctness reference value ( #4 …
7 -
vLLM releases dev-tools 2mo ago
v0.25.0: [CI] Fix cargo-deny config flag ordering (#48170)
Signed-off-by: Lucas Wilkinson [email protected]
35 -
vLLM releases dev-tools 2mo ago
v0.25.0rc3
[P/D][Bugfix] Fix PD async KV load lookahead handling for MTP spec de…
6 -
vLLM releases dev-tools 2mo ago
v0.25.0rc2
Fix embed scaling + CUDA graphs in Transformers modelling backend ( #4 …
12 -
vLLM releases dev-tools 2mo ago
v0.25.0rc1
[CPU][Bugfix] Fix flaky ShortConv prefill test on ARM (uninitialized …
5 -
vLLM releases dev-tools 3mo ago
v0.24.0
[CI] Raise gsm8k startup timeout for MoE Refactor Qwen3 NVFP4 configs…
23 -
vLLM releases dev-tools 3mo ago
v0.24.0rc2: Fix P/D with DP Supervisor (#46628)
Signed-off-by: Robert Shaw [email protected] (cherry picked from commit c5e3c40 )
7 -
vLLM releases dev-tools 3mo ago
v0.23.1rc0: [Bugfix][CI] Update Dockerfile dependency graph PNG (#45602)
Signed-off-by: sfeng33 [email protected]
37 -
vLLM releases dev-tools 3mo ago
v0.22.1rc2: fix: resolve CUTLASS fmin compatibility for DeepSeek-V4 init
Signed-off-by: khluu [email protected]
9 -
vLLM releases dev-tools 3mo ago
v0.22.1: fix: resolve CUTLASS fmin compatibility for DeepSeek-V4 init
Signed-off-by: khluu [email protected]
28 -
vLLM releases dev-tools 3mo ago
v0.22.1rc1: [docker] Stop using extra-index-url for flashinfer-jit-cache (#44366)
Signed-off-by: Kevin H. Luu [email protected]
34 -
vLLM releases dev-tools 4mo ago
v0.22.0rc2: Fix early CUDA init (#43791)
Signed-off-by: Harry Mellor [email protected] (cherry picked from commit 41688e2 )
11 -
vLLM releases dev-tools 4mo ago
v0.21.1rc0: [ROCm][CI] Stage B gating (#42025)
Signed-off-by: Andreas Karatzas [email protected]
17 -
vLLM releases dev-tools 4mo ago
v0.21.0
Highlights This release features 367 commits from 202 contributors (49 new)! Transformers v4 deprecated : This release formally deprecates transformers v4 support ( #40389 ). Users should migrate to transformers v5. C++20 build requirement : vLLM now requires a C++20-compatible…
23 -
vLLM releases dev-tools 4mo ago
v0.21.0rc3
[MLA Attention Backend] Add TOKENSPEED_MLA backend for DSR1/Kimi K25 …
28 -
vLLM releases dev-tools 4mo ago
v0.21.0rc2
[Bugfix] Install nvidia-cutlass-dsl[cu13] extra on CUDA 13 platforms …
16