█

AI/TLDR

vllm-mlx

vLLM-style inference server for Apple Silicon with OpenAI and Anthropic APIs

Local RuntimesOpen source
Latest
v0.5.0
Updated
17 Sep 2026
Language
Python
License
Apache-2.0

What's new

v0.5.017 Sep 2026

Added opt-in prefix-trie conversation caching, DeepSeek-V4-Flash support with reasoning and tool calls, and assistant drafting for multimodal models with continuous batching. Structured output now rejects requests whose constraints cannot be enforced.

Overview

vllm-mlx is a vLLM-style inference server for Apple Silicon Macs. It runs language, vision, audio and embedding models on Metal through Apple's MLX framework, using unified memory with no model conversion step, and exposes them from a single process through both OpenAI-compatible `/v1/*` endpoints (chat completions, completions, embeddings, rerank, responses) and the Anthropic-compatible `/v1/messages` endpoint. Any client that accepts a custom base URL, including the OpenAI and Anthropic SDKs and Claude Code, can point at it.

What sets it apart from running `mlx-lm` directly or using Ollama, in the project's own words, is the serving layer borrowed from vLLM's design: continuous batching for concurrent requests, a paged KV cache with prefix sharing, a trie-based prefix cache, and an SSD-tiered KV cache that spills prefixes to disk for long-context agents. Underneath, it routes requests to mlx-lm for LLMs, mlx-vlm for vision models, mlx-audio for speech and mlx-embeddings for embeddings.

The server also handles tool calling (with 19 parsers covering formats such as OpenAI, Anthropic, Gemini, Qwen, DeepSeek and Gemma), JSON Schema structured output, reasoning extraction for Qwen3 and DeepSeek-R1, speculative decoding for supported models, and Prometheus metrics. The README reports that OpenCode, pi, Codex, Claude Code, GitHub Copilot CLI, Cline CLI and OpenClaw's embedded agent each completed a streamed tool interaction and a file edit against it in a local test run. The project is Apache-2.0 licensed and written in Python; Rapid-MLX began as a community fork of it.

What it does

  • OpenAI-compatible `/v1/chat/completions`, `/v1/completions`, `/v1/embeddings`, `/v1/rerank` and `/v1/responses`, plus Anthropic-compatible `/v1/messages` from one process
  • Continuous batching, paged KV cache, trie-based prefix cache and an SSD-tiered KV cache (`--ssd-cache-dir`)
  • Text, image, video and audio input, native text-to-speech and Whisper-family speech-to-text
  • Tool calling with 19 parsers, JSON Schema structured output and reasoning extraction (`--reasoning-parser`)
  • Built-in `vllm-mlx bench-serve` benchmarker and a Prometheus `/metrics` endpoint
  • `vllm-mlx model inspect / acquire / convert` commands for fetching and converting Hugging Face weights

Getting started

vllm-mlx requires a Mac with Apple Silicon (M1 to M5) and an arm64-native Python 3.10 or later. The project recommends installing a pinned release into a fresh virtual environment; the README pins 0.4.1 as its reference release.

Install into a fresh environment

Check that Python runs natively (not under Rosetta), then install a pinned release. If you manage CLI tools with uv, `uv tool install 'vllm-mlx==0.4.1'` is the isolated equivalent.

bashbash
python3.12 -c 'import platform; print(platform.system(), platform.machine())'
# Expected: Darwin arm64
python3.12 -m venv .venv
source .venv/bin/activate
python -m pip install 'vllm-mlx==0.4.1'

Serve a model

Models load straight from Hugging Face's mlx-community repos. Once the first response works, restart with `--continuous-batching` to enable batching and its cache options.

bashbash
vllm-mlx serve mlx-community/Llama-3.2-3B-Instruct-4bit --host 127.0.0.1 --port 8000

Call it with the OpenAI SDK

Install `openai` in your client environment; the server package does not include it.

pythonpython
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")
r = client.chat.completions.create(model="mlx-community/Llama-3.2-3B-Instruct-4bit", messages=[{"role": "user", "content": "Hi!"}])
print(r.choices[0].message.content)

Point Claude Code at it

The Anthropic-compatible endpoint lives at the server root, so Anthropic clients use the base URL without `/v1`.

bashbash
vllm-mlx serve mlx-community/Qwen3-8B-4bit --port 8000
export ANTHROPIC_BASE_URL=http://localhost:8000
export ANTHROPIC_API_KEY=not-needed
claude

Benchmark and monitor

`bench-serve` sweeps prompts at a given concurrency and writes CSV or JSON; `--enable-metrics` exposes Prometheus metrics.

bashbash
vllm-mlx bench-serve --url http://localhost:8000 --concurrency 5 --prompts prompts.txt --output results.csv

vllm-mlx serve <model> --enable-metrics
curl http://localhost:8000/metrics

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Run Claude Code, Codex, OpenCode or another coding agent against a local model on a Mac
  • Serve several concurrent users or agent threads from one Mac with continuous batching and shared prefix caching
  • Give existing OpenAI- or Anthropic-SDK code a private local backend by changing only the base URL
  • Run vision, speech-to-text, text-to-speech, embeddings and reranking locally through one API server

How vllm-mlx compares

vllm-mlx alongside other open-source local runtimes tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Ollama★ 182kA developer-friendly tool that downloads and runs local LLMs from the terminal with a built-in OpenAI-compatible API.
llama.cpp★ 130kA C/C++ inference engine that runs LLMs in the GGUF format on CPUs, Apple Silicon, and GPUs with low memory use.
GPT4All★ 77.4kGPT4All is a free desktop app and Python client that runs large language models locally on your own computer, with no API calls or GPU required.
LocalAI★ 49.4kA self-hosted server that exposes an OpenAI-compatible API for running text, vision, voice, and image models on local hardware.
Jan★ 44.7kAn open-source desktop app that runs LLMs fully offline as a ChatGPT-style assistant on your own computer.
Colibrì★ 38.9kA pure-C inference engine that keeps a Mixture-of-Experts model's dense trunk resident in RAM and streams its routed experts from disk, so 744B-2.8T models run on consumer hardware.
llmfit★ 37.4kA Rust terminal tool that inspects your CPU, RAM, GPUs and VRAM and scores which open-weight models and quantizations will actually run well on that machine, with a TUI, CLI, REST API and local-runtime integrations.
vllm-mlx★ 1.6kvLLM-style inference server for Apple Silicon with OpenAI and Anthropic APIs