Engine and design goals
vLLM presents itself as the high-throughput and memory-efficient inference and serving engine for LLMs, aiming at easy, fast and cost-efficient LLM serving. Throughput is maximised through PagedAttention, with advanced scheduling and continuous batching intended to keep GPU utilisation high. The project frames cost efficiency as a result of maximising hardware efficiency rather than requiring additional hardware. A drop-in OpenAI-compatible API allows existing client integrations to be pointed at a vLLM server.
- PagedAttention for memory-efficient KV cache
- Continuous batching and advanced scheduling
- Drop-in OpenAI-compatible API
- Widest range of open-source models on any hardware
Installation and quick start
The vLLM site provides a quick-start selector covering stable and nightly builds, hardware backend, installation method and CUDA version. The project recommends uv for faster and more reliable installation, with the command uv pip install vllm --torch-backend auto for CUDA builds. Python 3.10 or later is required, and Python 3.12 or later is recommended. Installation options also include plain Python packaging and Docker images, with other platforms documented at docs.vllm.ai.
- uv pip install vllm --torch-backend auto
- Requires Python 3.10+, Python 3.12+ recommended
- Stable and nightly build channels
- Python (uv), Python and Docker install paths
- CUDA 13.0 and CUDA 12.9 selections
Hardware and model compatibility
vLLM describes a unified API across platforms, so that one engine runs models on different accelerators. The quick-start selector covers CUDA, ROCm, XPU and CPU backends, and the compatibility listing extends to further accelerators. The site also lists trending open-source model families that are optimised and production-ready on the engine. Project documentation on release quality cites support for more than 1,000 model architectures and 600+ accelerator types.
- NVIDIA CUDA GPU, AMD ROCm GPU, Gaudi XPU, CPU
- Huawei Ascend NPU, AWS Neuron, Google Cloud TPU, IBM Spyre
- Apple Silicon, Baidu Kunlun XPU, Cambricon MLU
- Model families including DeepSeek, Gemma, Llama, MiniMax, Mistral, Kimi, Nemotron, Qwen, Step and GLM
- 1000+ model architectures and 600+ accelerator types
Day-0 model support
vLLM publishes engineering posts describing support for newly released open-weight models at launch. Day-0 support has been announced for Kimi K3, a 2.8-trillion-parameter Mixture-of-Experts model from Moonshot AI with a 1M-token context window, and for the MiniMax M3 family in BF16 and MXFP8 checkpoints. DiffusionGemma, a 26B-parameter discrete diffusion language model from Google built on the Gemma4 backbone, is documented as the first diffusion LLM natively supported in vLLM. Each release is accompanied by serving commands and deployment recipes.
- Kimi K3 — day-0 serving with speculative decoding and prefill/decode disaggregation
- MiniMax M3 — 1M-token context and MiniMax Sparse Attention
- DiffusionGemma — first diffusion LLM supported natively
- GLM-5.2-NVFP4 — disaggregated serving on NVIDIA B300 GPUs
Release quality and engineering process
vLLM documents a three-layer path from pull request to release: continuous integration on every PR, nightly performance benchmarking and accuracy evaluation, and a release process that gates artefacts. Lightweight GitHub Actions checks run first, after which Buildkite assembles a testing pipeline dynamically from the diff; the suite comprises 37 test groups and 266 jobs. The project reports 86K+ GitHub stars, 5.6M+ monthly pip installs and 2.5M+ monthly image pulls. In June 2026 it merged 1,918 commits into main, with CI consuming 13 million job minutes and peaking at 1,400 concurrent runners.
- 37 CI test groups and 266 jobs
- Shared container image and pinned dependency graph
- Nightly benchmarking and accuracy evaluation
- 86K+ GitHub stars; 5.6M+ monthly pip installs
Recipes and documentation
vLLM Recipes is a catalogue of deployment guides under the strapline "deploy any model on any hardware with vLLM", published by the vLLM Project. Recipes cover individual model checkpoints with details such as parameter counts, quantisation format, context length and modality. Blog posts on model support link to the corresponding recipe for exact Docker and serve commands. The main site also signposts documentation, benchmarks, roadmap and example notebooks.
- 195+ recipes listed in the catalogue
- Per-model deployment instructions and Docker commands
- Documentation at docs.vllm.ai
Community, support and funding
The project directs questions to a Slack workspace, a forum serving as a searchable Q&A knowledge base, and GitHub Issues for bug reports and feature requests. Joining the Slack workspace is done through an invitation page and is subject to the project's terms of service and privacy statement. vLLM is a community project whose development and testing compute is supported by sponsoring organisations, including cloud providers, hardware vendors and universities, alongside cash donors. Donations are collected through GitHub and OpenCollective and are intended to support development, maintenance and adoption.
- Slack workspace for real-time help and discussion
- Forum for searchable Q&A
- GitHub Issues for bugs and feature requests
- Cash donations via GitHub and OpenCollective
- Compute sponsorship from cloud and hardware organisations
Sources
-
vLLM Recipes — Deploy any model on any hardware with vLLM https://recipes.vllm.ai/ Verified 02 Oct 2026
-
[Roadmap] vLLM Roadmap Q3 2026 · Issue #48168 · vllm-project/vllm · GitHub https://github.com/vllm-project/vllm/issues/48168 Verified 02 Oct 2026
-
Overview - vLLM Hardware Plugin for Intel® Gaudi® https://docs.vllm.ai/projects/gaudi/en/latest/ Verified 18 Aug 2026
-
Supported Models - vLLM https://docs.vllm.ai/en/latest/models/supported_models/ Verified 18 Aug 2026
-
vLLM http://www.vllm.ai/
-
Custom Logits Processors - vLLM https://docs.vllm.ai/en/latest/features/custom_logitsprocs/ Verified 02 Oct 2026
-
KV Offloading Usage Guide - vLLM https://docs.vllm.ai/en/latest/features/kv_offloading_usage/ Verified 17 Sep 2026
-
Distributed Layerwise Offload: Scaling Toward 200B+ DiT Models Efficiently in vLLM-Omni | vLLM Blog https://vllm.ai/blog/2026-08-17-distributed-layerwise-offload Verified 02 Oct 2026
-
Kimi K3 Is Here: Efficient Day-0 Support on vLLM | vLLM Blog https://vllm.ai/blog/2026-07-27-k3 Verified 02 Oct 2026
-
From Day 0 to Production SLAs: Serving GLM-5.2 on 24 NVIDIA B300 GPUs with vLLM | vLLM Blog https://vllm.ai/blog/2026-07-23-glm-5.2-nvfp4-b300-pd Verified 02 Oct 2026
-
Keeping vLLM Production Quality: A Look Inside CI, Benchmarking, and the Release Process | vLLM Blog https://vllm.ai/blog/2026-07-16-keeping-vllm-production-quality Verified 02 Oct 2026
-
vLLM × HPC-Ops: High-Performance Attention and MoE Backends from Tencent Hunyuan | vLLM Blog https://vllm.ai/blog/2026-07-06-vllm-hpc-ops Verified 02 Oct 2026
-
Experience and Lessons Learned from Serving Multi-Stage Qwen3-Omni in vLLM-Omni | vLLM Blog https://vllm.ai/blog/2026-07-01-qwen3-omni-optimization Verified 02 Oct 2026
-
MiniMax M3 in vLLM: Day-0 Serving for 1M-Token Multimodal Reasoning | vLLM Blog https://vllm.ai/blog/2026-06-12-minimax-m3-vllm Verified 02 Oct 2026
-
DiffusionGemma: The First Diffusion LLM (dLLM) Natively Supported in vLLM | vLLM Blog https://vllm.ai/blog/2026-06-10-diffusion-gemma Verified 02 Oct 2026
-
vLLM on the DGX Spark: Architecture, Configuration, and Local Evaluation | vLLM Blog https://vllm.ai/blog/2026-06-01-vllm-dgx-spark Verified 02 Oct 2026
Last verified 02 Oct 2026. This entry is compiled from the public web pages listed above. Nothing here is stated that those pages do not, and each of them was read on the date shown.