Overview
Parlor is a real-time multimodal assistant that runs entirely on your own machine. You open a page in the browser, grant microphone and camera access, and talk: Gemma 4 running through llama.cpp hears your speech and sees the camera frame, streams a reply, and Kokoro TTS speaks it back sentence by sentence while it is still generating. The author built it to match the feel of cloud voice assistants on a MacBook M3 Pro, and it is labelled a research preview — an early experiment with rough edges.
It is a classic cascade rather than a single full-duplex model. The browser runs Silero VAD with a roughly 200ms silence cutoff, so there is no push-to-talk; Pipecat's smart-turn-v3 classifier then decides whether you actually finished your thought, holding mid-sentence pauses instead of answering them. Audio and the camera frame are pushed through llama.cpp's prompt cache while you are still talking, so long questions start answering quickly, and you can speak over the assistant to interrupt it — generation is aborted server-side.
Actions are kept out of the spoken reply: timers, mode switches and research requests are decided by a separate grammar-forced JSON request over the same prompt cache. That powers server-owned timers with a countdown chip, a live translation mode that interprets each utterance after a short pause, and a just-listen mode that only transcribes. An optional background reasoner hands tasks such as web research to a frontier model on any OpenAI-compatible endpoint while the conversation continues; it stays off unless `REASONER_API_KEY` is set, so by default nothing leaves the device. The project is Python, licensed Apache-2.0, and its README states that it was developed with strong assistance from Claude.
What it does
- Hears and sees in one model: Gemma 4 (E2B, E4B or 12B) via llama.cpp takes speech and the browser camera frame
- Hands-free turn-taking with browser-side Silero VAD plus the smart-turn-v3 end-of-turn classifier
- Streaming replies spoken sentence by sentence through Kokoro TTS (MLX on Mac, ONNX on Linux), with barge-in to interrupt
- Grammar-forced JSON action head for timers, live translation mode and just-listen transcription mode
- Optional background research on any OpenAI-compatible endpoint, off unless an API key is configured
- End-to-end test suite and latency, turn-detection and architecture benchmarks in the repository
Getting started
Parlor needs Python 3.12+, uv, and llama.cpp build b9503 or newer (b9512 for `MODEL=12b`), on macOS with Apple Silicon or Linux with a supported GPU. The default E4B model needs about 6 GB of free RAM; `MODEL=e2b` fits in about 4 GB.
Clone and install the prerequisites
On macOS, llama.cpp comes from Homebrew; other platforms follow llama.cpp's own install guide.
git clone https://github.com/fikrikarim/parlor.git
cd parlor
curl -LsSf https://astral.sh/uv/install.sh | sh
brew install llama.cppSync and run
Models download automatically on first run — about 5.7 GB for Gemma 4 E4B QAT and its multimodal projector, plus the TTS models.
uv sync
uv run parlorStart talking
Open http://localhost:8000, grant camera and microphone access, and speak. Headphones avoid echo, though Parlor handles it without the browser's echo canceller.
Configure the model and optional research
Set variables in your shell or a .env at the repo root. `MODEL` picks the Gemma 4 size and `PORT` the server port; setting `REASONER_API_KEY` (with `REASONER_BASE_URL`, default OpenRouter, and `REASONER_MODEL`) turns on background research. The full list is in docs/configuration.md.
MODEL=e2b
PORT=8000
# REASONER_API_KEY=...Test and benchmark
The test suite spawns the real server and drives it over WebSocket with synthesized speech; the benchmark measures end-of-utterance to first audio against a running server.
uv run pytest
uv run python benchmarks/bench.py --label before --out benchmarks/results/before.jsonCommands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Hold a private spoken conversation with an assistant that can see your camera, with no cloud service involved
- Practise speaking a language, or use live translation mode as a consecutive interpreter
- Think out loud in just-listen mode and get an on-screen transcript without replies
- Study or extend a local cascade voice pipeline with its turn detection, streaming TTS and benchmarks already in place
Version history
Every verified update to Parlor that AI/TLDR tracked, newest first — each links to our coverage and the official changeset.
- 2026-08-02v2.0.0
Near-complete rebuild: inference moved from LiteRT-LM to llama.cpp, turn-taking to the smart-turn-v3 classifier and actions to a grammar-forced JSON head, adding background research, timers, live translation and just-listen modes; the default model became Gemma 4 E4B.
- 2026-04-07v1.0.0
Initial release, published retroactively on GitHub on 2026-07-29.
How Parlor compares
Parlor alongside other open-source assistants & chatbots tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| OpenClaw | ★ 391k | OpenClaw is a self-hosted personal AI assistant that answers you on WhatsApp, Telegram, Slack, Discord, and many other channels, with voice and a live visual canvas. |
| Hermes Agent | ★ 251k | A self-improving personal AI agent from Nous Research that builds skills from experience, remembers across sessions, and reaches you on Telegram, Discord, Slack, and more. |
| Odysseus | ★ 88.9k | A self-hosted AI workspace that puts chat, agents, deep research, documents, email, notes, tasks and calendar behind one Docker Compose stack, over local or API models. |
| CowAgent | ★ 47.2k | A self-hosted assistant that plans and executes tasks with built-in file, terminal, browser and search tools, and answers across a web console plus a dozen messaging platforms. |
| AstrBot | ★ 41.3k | An all-in-one agent chatbot platform that puts LLM conversations, tools, knowledge bases and a plugin marketplace inside messaging apps like Telegram, Slack, Discord, QQ and Feishu. |
| OpenHuman | ★ 40.5k | A local-first desktop personal AI for macOS, Windows and Linux that keeps a compressed memory tree on your machine and orchestrates checkpointed research and automation workflows. |
| MindsHub | ★ 39.8k | An agent workspace for knowledge work and software development that runs swappable open-source agent harnesses against your choice of frontier or open models. |
| Parlor | ★ 2.1k | A fully on-device, real-time voice and vision assistant built on Gemma 4 and llama.cpp |