Published event
ArtificialIntelligence ProductUpdate 1 source(s)

v0.20.1

Updated September 26, 2026 · 2:46 PM · source date May 4, 2026

Summary

v0.20.1 vllm-project / vllm Public Uh oh! There was an error while loading.

Why it matters

This ProductUpdate is relevant to the technology intelligence record because it involves DeepSeek, DeepSeek v4, gpt-oss, GPT. The source article should remain the factual reference for follow-up coverage.

Key facts
  • vllm-project / vllm Public Uh oh!
  • There was an error while loading.
  • Notifications You must be signed in to change notification settings Fork 22.7k Star 92.7k v0.20.1 khluu released this 04 May 10:36 · 5903 commits to main since this release v0.20.1 132765e vLLM v0.20.1 This is a patch release on top of v0.20.0 primarily focused on DeepSeek V4 stabilization and performance improvements , along with several important bug fixes.
  • DeepSeek V4 Base model support ( #41006 ).
  • Multi-stream pre-attention GEMM ( #41061 ), configurable pre-attn GEMM knob ( #41443 ), and tuned default VLLM_MULTI_STREAM_GEMM_TOKEN_THRESHOLD ( #41526 ).
  • BF16 and MXFP8 all-to-all support for FlashInfer one-sided communication ( #40960 ).
Entities in this story

Companies

DeepSeek→

Technologies

CUDA→
Related events