AI/TLDR

vLLM

Fast, memory-efficient inference and serving for large language models

Serving & DeploymentOpen source
Latest
v0.29.0
Updated
9 Sep 2026
Language
Python
License
Apache-2.0
Coverage
8 stories
$uv pip install vllm

What's new

v0.29.09 Sep 2026

Model Runner V2 becomes the default for all models, and ten deprecated architectures (Arctic, Chameleon, MPT and others) are removed. 594 commits from 277 contributors add Hy4-preview and Qwen3.8-Flash-Next support; `python -m vllm.entrypoints.openai.api_server` is deprecated in favour of `vllm serve`.

Latest news

all 8 ↓

Overview

vLLM is a library for running and serving large language models. It loads models from Hugging Face and exposes them either as a Python API for batch generation or as an OpenAI-compatible HTTP server. It was originally developed in the Sky Computing Lab at UC Berkeley and is now maintained by a large open-source community.

It is built for teams that need to serve many requests at the same time without running out of GPU memory. Its PagedAttention technique manages the attention key/value cache more efficiently, and continuous batching keeps the GPU busy by adding and removing requests on the fly. It supports 200+ model architectures, including decoder-only LLMs, mixture-of-expert models, and multimodal models.

Within the inference and serving category, vLLM sits at the high-throughput serving end. You point it at a model, and it handles batching, memory management, streaming, and an API surface so you can focus on your application instead of the serving plumbing.

What it does

  • PagedAttention manages attention key/value memory efficiently, reducing waste and fitting more concurrent requests in GPU memory
  • Continuous batching, chunked prefill, and prefix caching keep throughput high under load
  • OpenAI-compatible API server, plus Anthropic Messages API and gRPC support
  • Wide quantization support including FP8, INT8, INT4, GPTQ/AWQ, and GGUF
  • Tensor, pipeline, data, expert, and context parallelism for distributed inference
  • Runs on NVIDIA and AMD GPUs, x86/ARM/PowerPC CPUs, and other hardware via plugins

Getting started

Install vLLM, then either generate text from Python or start an OpenAI-compatible server.

Install vLLM

Install with uv (recommended) or pip. A GPU with a matching CUDA setup is the common target.

bashbash
uv pip install vllm

Run offline batched inference

Load a model and generate text for a list of prompts directly in Python.

pythonpython
from vllm import LLM, SamplingParams

prompts = [
    "Hello, my name is",
    "The capital of France is",
]
sampling_params = SamplingParams(temperature=0.8, top_p=0.95)

llm = LLM(model="facebook/opt-125m")
outputs = llm.generate(prompts, sampling_params)

for output in outputs:
    print(output.prompt, output.outputs[0].text)

Start an OpenAI-compatible server

Serve a model over HTTP at http://localhost:8000 with chat and completion endpoints that match the OpenAI API.

bashbash
vllm serve Qwen/Qwen2.5-1.5B-Instruct

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Serving a chat or completion model to many users at once behind an OpenAI-compatible API
  • Running high-volume offline batch generation over large prompt sets
  • Hosting open-weight Hugging Face models on your own GPUs instead of a paid API
  • Scaling a single large model across multiple GPUs with tensor or pipeline parallelism

Version history

Every verified update to vLLM that AI/TLDR tracked, newest first — each links to our coverage and the official changeset.

  1. 2026-09-09v0.29.0

    Model Runner V2 becomes the default for all models, and ten deprecated architectures (Arctic, Chameleon, MPT and others) are removed. 594 commits from 277 contributors add Hy4-preview and Qwen3.8-Flash-Next support; `python -m vllm.entrypoints.openai.api_server` is deprecated in favour of `vllm serve`.

  2. 2026-08-26v0.28.0

    vLLM v0.28.0 lands 584 commits from 270 contributors: a Kimi K3 speed push (decode context parallelism, fused kernels, expert sharding saving ~17 GiB per GPU) plus end-to-end sparse MLA for DeepSeek V4.

  3. 2026-04-27v0.20.0

    752 commits from 320 contributors: DeepSeek V4 support, FlashAttention 4 as the default MLA prefill backend, TurboQuant 2-bit KV cache for 4x capacity, and CUDA 13 plus PyTorch 2.11 as defaults.

vLLM in the news

  1. 2026-09-09MAJORvLLM v0.29.0 — Model Runner V2 becomes the default for every model
  2. 2026-09-08NOTABLECohere megakernel — one CUDA file serves North Mini Code faster than vLLM
  3. 2026-08-28MAJORGLM-5.3 weights go public — Z.ai's 753B coding model lands on Hugging Face
  4. 2026-08-26MAJORvLLM v0.28.0 — Kimi K3 gets a full-stack speed pass
  5. 2026-08-25MAJORPrime Intellect finds an offline sandbox escape — the inference API is the hole
  6. 2026-08-24NOTABLEBoyd Kane — a model could escape by attacking the engine that runs it
  7. 2026-08-18NOTABLESam Witteveen — 'Qwen3.8-27B & How to Serve it Fast'
  8. 2026-08-17MAJORAgent Lightning v1.0 — Microsoft's RL trainer plugs into real agent harnesses

From the AI/TLDR release feed — every item is source-verified when it ships.

How vLLM compares

vLLM alongside other open-source serving & deployment tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Transformers★ 166kHugging Face Transformers is a Python framework that defines and runs state-of-the-art pretrained models for text, vision, audio, and multimodal tasks, for both inference and training.
vLLM★ 92.1kFast, memory-efficient inference and serving for large language models
SGLang★ 36.1kA serving framework for LLMs and multimodal models that boosts throughput by reusing shared prompt prefixes across requests.
Modular Platform★ 29.8kModular's AI platform: the MAX inference framework with an OpenAI-compatible serving endpoint and GPU kernel library, plus the Mojo language and compiler.
TensorRT-LLM★ 14.7kNVIDIA's library that compiles LLMs into optimized engines for the fastest inference on its data-center GPUs.
OpenLLM★ 12.5kA tool to run any open-source LLM as an OpenAI-compatible API endpoint locally or in the cloud.
LMCache★ 11.9kA KV-cache layer that stores and shares cached attention state across engines and requests to cut repeated computation.
NVIDIA Triton Inference Server★ 11kA multi-framework model server that runs TensorRT, PyTorch, ONNX, and other models with dynamic batching and concurrent execution.