Engine and design goals

vLLM presents itself as the high-throughput and memory-efficient inference and serving engine for LLMs, aiming at easy, fast and cost-efficient LLM serving. Throughput is maximised through PagedAttention, with advanced scheduling and continuous batching intended to keep GPU utilisation high. The project frames cost efficiency as a result of maximising hardware efficiency rather than requiring additional hardware. A drop-in OpenAI-compatible API allows existing client integrations to be pointed at a vLLM server.

  • PagedAttention for memory-efficient KV cache
  • Continuous batching and advanced scheduling
  • Drop-in OpenAI-compatible API
  • Widest range of open-source models on any hardware

Source: vLLM

Installation and quick start

The vLLM site provides a quick-start selector covering stable and nightly builds, hardware backend, installation method and CUDA version. The project recommends uv for faster and more reliable installation, with the command uv pip install vllm --torch-backend auto for CUDA builds. Python 3.10 or later is required, and Python 3.12 or later is recommended. Installation options also include plain Python packaging and Docker images, with other platforms documented at docs.vllm.ai.

  • uv pip install vllm --torch-backend auto
  • Requires Python 3.10+, Python 3.12+ recommended
  • Stable and nightly build channels
  • Python (uv), Python and Docker install paths
  • CUDA 13.0 and CUDA 12.9 selections

Source: vLLM

Hardware and model compatibility

vLLM describes a unified API across platforms, so that one engine runs models on different accelerators. The quick-start selector covers CUDA, ROCm, XPU and CPU backends, and the compatibility listing extends to further accelerators. The site also lists trending open-source model families that are optimised and production-ready on the engine. Project documentation on release quality cites support for more than 1,000 model architectures and 600+ accelerator types.

  • NVIDIA CUDA GPU, AMD ROCm GPU, Gaudi XPU, CPU
  • Huawei Ascend NPU, AWS Neuron, Google Cloud TPU, IBM Spyre
  • Apple Silicon, Baidu Kunlun XPU, Cambricon MLU
  • Model families including DeepSeek, Gemma, Llama, MiniMax, Mistral, Kimi, Nemotron, Qwen, Step and GLM
  • 1000+ model architectures and 600+ accelerator types

Sources: vLLM, Keeping vLLM Production Quality: A Look Inside CI, Benchmarking, and the Release Process | vLLM Blog

Day-0 model support

vLLM publishes engineering posts describing support for newly released open-weight models at launch. Day-0 support has been announced for Kimi K3, a 2.8-trillion-parameter Mixture-of-Experts model from Moonshot AI with a 1M-token context window, and for the MiniMax M3 family in BF16 and MXFP8 checkpoints. DiffusionGemma, a 26B-parameter discrete diffusion language model from Google built on the Gemma4 backbone, is documented as the first diffusion LLM natively supported in vLLM. Each release is accompanied by serving commands and deployment recipes.

  • Kimi K3 — day-0 serving with speculative decoding and prefill/decode disaggregation
  • MiniMax M3 — 1M-token context and MiniMax Sparse Attention
  • DiffusionGemma — first diffusion LLM supported natively
  • GLM-5.2-NVFP4 — disaggregated serving on NVIDIA B300 GPUs

Sources: Kimi K3 Is Here: Efficient Day-0 Support on vLLM | vLLM Blog, MiniMax M3 in vLLM: Day-0 Serving for 1M-Token Multimodal Reasoning | vLLM Blog, DiffusionGemma: The First Diffusion LLM (dLLM) Natively Supported in vLLM | vLLM Blog, From Day 0 to Production SLAs: Serving GLM-5.2 on 24 NVIDIA B300 GPUs with vLLM | vLLM Blog

Related projects and integrations

Several sub-projects and third-party contributions extend the core engine. vLLM-Omni serves multi-stage multimodal models such as Qwen3-Omni as a staged Thinker–Talker–Code2Wav pipeline behind the /v1/chat/completions endpoint, and its Distributed Layerwise Offload feature shards and streams diffusion transformer weights across devices so that models larger than single-device HBM can run. HPC-Ops, the production operator library from the Tencent Hunyuan AI Infra team, contributes Attention and MoE kernels to the vLLM main branch as first-class backends. vLLM also documents local deployment on NVIDIA DGX Spark, pairing the OpenAI-compatible API with Prometheus telemetry and paged KV cache for unified-memory systems.

  • vLLM-Omni for staged multimodal and speech serving
  • Distributed Layerwise Offload for large DiT models
  • HPC-Ops attention and MoE backends
  • DGX Spark local serving with Prometheus metrics

Sources: Experience and Lessons Learned from Serving Multi-Stage Qwen3-Omni in vLLM-Omni | vLLM Blog, Distributed Layerwise Offload: Scaling Toward 200B+ DiT Models Efficiently in vLLM-Omni | vLLM Blog, vLLM × HPC-Ops: High-Performance Attention and MoE Backends from Tencent Hunyuan | vLLM Blog, vLLM on the DGX Spark: Architecture, Configuration, and Local Evaluation | vLLM Blog

Release quality and engineering process

vLLM documents a three-layer path from pull request to release: continuous integration on every PR, nightly performance benchmarking and accuracy evaluation, and a release process that gates artefacts. Lightweight GitHub Actions checks run first, after which Buildkite assembles a testing pipeline dynamically from the diff; the suite comprises 37 test groups and 266 jobs. The project reports 86K+ GitHub stars, 5.6M+ monthly pip installs and 2.5M+ monthly image pulls. In June 2026 it merged 1,918 commits into main, with CI consuming 13 million job minutes and peaking at 1,400 concurrent runners.

  • 37 CI test groups and 266 jobs
  • Shared container image and pinned dependency graph
  • Nightly benchmarking and accuracy evaluation
  • 86K+ GitHub stars; 5.6M+ monthly pip installs

Source: Keeping vLLM Production Quality: A Look Inside CI, Benchmarking, and the Release Process | vLLM Blog

Recipes and documentation

vLLM Recipes is a catalogue of deployment guides under the strapline "deploy any model on any hardware with vLLM", published by the vLLM Project. Recipes cover individual model checkpoints with details such as parameter counts, quantisation format, context length and modality. Blog posts on model support link to the corresponding recipe for exact Docker and serve commands. The main site also signposts documentation, benchmarks, roadmap and example notebooks.

  • 195+ recipes listed in the catalogue
  • Per-model deployment instructions and Docker commands
  • Documentation at docs.vllm.ai

Sources: vLLM Recipes — Deploy any model on any hardware with vLLM, vLLM, DiffusionGemma: The First Diffusion LLM (dLLM) Natively Supported in vLLM | vLLM Blog

Community, support and funding

The project directs questions to a Slack workspace, a forum serving as a searchable Q&A knowledge base, and GitHub Issues for bug reports and feature requests. Joining the Slack workspace is done through an invitation page and is subject to the project's terms of service and privacy statement. vLLM is a community project whose development and testing compute is supported by sponsoring organisations, including cloud providers, hardware vendors and universities, alongside cash donors. Donations are collected through GitHub and OpenCollective and are intended to support development, maintenance and adoption.

  • Slack workspace for real-time help and discussion
  • Forum for searchable Q&A
  • GitHub Issues for bugs and feature requests
  • Cash donations via GitHub and OpenCollective
  • Compute sponsorship from cloud and hardware organisations

Sources: vLLM, vLLM

Sources

  1. vLLM https://vllm.ai/ Verified 02 Oct 2026
  2. Contact | vLLM https://vllm.ai/contact Verified 02 Oct 2026
  3. vLLM https://docs.vllm.ai/en/latest/ Verified 02 Oct 2026
  4. vLLM https://inviter.co/vllm-slack Verified 02 Oct 2026
  5. vLLM Forums https://discuss.vllm.ai/ Verified 02 Oct 2026
  6. vLLM Recipes — Deploy any model on any hardware with vLLM https://recipes.vllm.ai/ Verified 02 Oct 2026
  7. [Roadmap] vLLM Roadmap Q3 2026 · Issue #48168 · vllm-project/vllm · GitHub https://github.com/vllm-project/vllm/issues/48168 Verified 02 Oct 2026
  8. Previous vLLM Releases | vLLM https://vllm.ai/releases Verified 02 Oct 2026
  9. Troubleshooting - vLLM https://docs.vllm.ai/en/latest/usage/troubleshooting/ Verified 18 Aug 2026
  10. Overview - vLLM Hardware Plugin for Intel® Gaudi® https://docs.vllm.ai/projects/gaudi/en/latest/ Verified 18 Aug 2026
  11. Supported Models - vLLM https://docs.vllm.ai/en/latest/models/supported_models/ Verified 18 Aug 2026
  12. Custom Logits Processors - vLLM https://docs.vllm.ai/en/latest/features/custom_logitsprocs/ Verified 02 Oct 2026
  13. KV Offloading Usage Guide - vLLM https://docs.vllm.ai/en/latest/features/kv_offloading_usage/ Verified 17 Sep 2026
  14. Distributed Layerwise Offload: Scaling Toward 200B+ DiT Models Efficiently in vLLM-Omni | vLLM Blog https://vllm.ai/blog/2026-08-17-distributed-layerwise-offload Verified 02 Oct 2026
  15. Kimi K3 Is Here: Efficient Day-0 Support on vLLM | vLLM Blog https://vllm.ai/blog/2026-07-27-k3 Verified 02 Oct 2026
  16. From Day 0 to Production SLAs: Serving GLM-5.2 on 24 NVIDIA B300 GPUs with vLLM | vLLM Blog https://vllm.ai/blog/2026-07-23-glm-5.2-nvfp4-b300-pd Verified 02 Oct 2026
  17. Keeping vLLM Production Quality: A Look Inside CI, Benchmarking, and the Release Process | vLLM Blog https://vllm.ai/blog/2026-07-16-keeping-vllm-production-quality Verified 02 Oct 2026
  18. vLLM × HPC-Ops: High-Performance Attention and MoE Backends from Tencent Hunyuan | vLLM Blog https://vllm.ai/blog/2026-07-06-vllm-hpc-ops Verified 02 Oct 2026
  19. Experience and Lessons Learned from Serving Multi-Stage Qwen3-Omni in vLLM-Omni | vLLM Blog https://vllm.ai/blog/2026-07-01-qwen3-omni-optimization Verified 02 Oct 2026
  20. MiniMax M3 in vLLM: Day-0 Serving for 1M-Token Multimodal Reasoning | vLLM Blog https://vllm.ai/blog/2026-06-12-minimax-m3-vllm Verified 02 Oct 2026
  21. DiffusionGemma: The First Diffusion LLM (dLLM) Natively Supported in vLLM | vLLM Blog https://vllm.ai/blog/2026-06-10-diffusion-gemma Verified 02 Oct 2026
  22. vLLM on the DGX Spark: Architecture, Configuration, and Local Evaluation | vLLM Blog https://vllm.ai/blog/2026-06-01-vllm-dgx-spark Verified 02 Oct 2026

Last verified 02 Oct 2026. This entry is compiled from the public web pages listed above. Nothing here is stated that those pages do not, and each of them was read on the date shown.