█

AI/TLDR

Parlor

A fully on-device, real-time voice and vision assistant built on Gemma 4 and llama.cpp

Assistants & ChatbotsOpen source
Latest
v2.0.0
Updated
2 Aug 2026
Language
Python
License
Apache-2.0
$git clone https://github.com/fikrikarim/parlor.git

What's new

v2.0.02 Aug 2026

Near-complete rebuild: inference moved from LiteRT-LM to llama.cpp, turn-taking to the smart-turn-v3 classifier and actions to a grammar-forced JSON head, adding background research, timers, live translation and just-listen modes; the default model became Gemma 4 E4B.

Overview

Parlor is a real-time multimodal assistant that runs entirely on your own machine. You open a page in the browser, grant microphone and camera access, and talk: Gemma 4 running through llama.cpp hears your speech and sees the camera frame, streams a reply, and Kokoro TTS speaks it back sentence by sentence while it is still generating. The author built it to match the feel of cloud voice assistants on a MacBook M3 Pro, and it is labelled a research preview — an early experiment with rough edges.

It is a classic cascade rather than a single full-duplex model. The browser runs Silero VAD with a roughly 200ms silence cutoff, so there is no push-to-talk; Pipecat's smart-turn-v3 classifier then decides whether you actually finished your thought, holding mid-sentence pauses instead of answering them. Audio and the camera frame are pushed through llama.cpp's prompt cache while you are still talking, so long questions start answering quickly, and you can speak over the assistant to interrupt it — generation is aborted server-side.

Actions are kept out of the spoken reply: timers, mode switches and research requests are decided by a separate grammar-forced JSON request over the same prompt cache. That powers server-owned timers with a countdown chip, a live translation mode that interprets each utterance after a short pause, and a just-listen mode that only transcribes. An optional background reasoner hands tasks such as web research to a frontier model on any OpenAI-compatible endpoint while the conversation continues; it stays off unless `REASONER_API_KEY` is set, so by default nothing leaves the device. The project is Python, licensed Apache-2.0, and its README states that it was developed with strong assistance from Claude.

What it does

  • Hears and sees in one model: Gemma 4 (E2B, E4B or 12B) via llama.cpp takes speech and the browser camera frame
  • Hands-free turn-taking with browser-side Silero VAD plus the smart-turn-v3 end-of-turn classifier
  • Streaming replies spoken sentence by sentence through Kokoro TTS (MLX on Mac, ONNX on Linux), with barge-in to interrupt
  • Grammar-forced JSON action head for timers, live translation mode and just-listen transcription mode
  • Optional background research on any OpenAI-compatible endpoint, off unless an API key is configured
  • End-to-end test suite and latency, turn-detection and architecture benchmarks in the repository

Getting started

Parlor needs Python 3.12+, uv, and llama.cpp build b9503 or newer (b9512 for `MODEL=12b`), on macOS with Apple Silicon or Linux with a supported GPU. The default E4B model needs about 6 GB of free RAM; `MODEL=e2b` fits in about 4 GB.

Clone and install the prerequisites

On macOS, llama.cpp comes from Homebrew; other platforms follow llama.cpp's own install guide.

bashbash
git clone https://github.com/fikrikarim/parlor.git
cd parlor

curl -LsSf https://astral.sh/uv/install.sh | sh
brew install llama.cpp

Sync and run

Models download automatically on first run — about 5.7 GB for Gemma 4 E4B QAT and its multimodal projector, plus the TTS models.

bashbash
uv sync
uv run parlor

Start talking

Open http://localhost:8000, grant camera and microphone access, and speak. Headphones avoid echo, though Parlor handles it without the browser's echo canceller.

Configure the model and optional research

Set variables in your shell or a .env at the repo root. `MODEL` picks the Gemma 4 size and `PORT` the server port; setting `REASONER_API_KEY` (with `REASONER_BASE_URL`, default OpenRouter, and `REASONER_MODEL`) turns on background research. The full list is in docs/configuration.md.

bashbash
MODEL=e2b
PORT=8000
# REASONER_API_KEY=...

Test and benchmark

The test suite spawns the real server and drives it over WebSocket with synthesized speech; the benchmark measures end-of-utterance to first audio against a running server.

bashbash
uv run pytest
uv run python benchmarks/bench.py --label before --out benchmarks/results/before.json

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Hold a private spoken conversation with an assistant that can see your camera, with no cloud service involved
  • Practise speaking a language, or use live translation mode as a consecutive interpreter
  • Think out loud in just-listen mode and get an on-screen transcript without replies
  • Study or extend a local cascade voice pipeline with its turn detection, streaming TTS and benchmarks already in place

Version history

Every verified update to Parlor that AI/TLDR tracked, newest first — each links to our coverage and the official changeset.

  1. 2026-08-02v2.0.0

    Near-complete rebuild: inference moved from LiteRT-LM to llama.cpp, turn-taking to the smart-turn-v3 classifier and actions to a grammar-forced JSON head, adding background research, timers, live translation and just-listen modes; the default model became Gemma 4 E4B.

  2. 2026-04-07v1.0.0

    Initial release, published retroactively on GitHub on 2026-07-29.

How Parlor compares

Parlor alongside other open-source assistants & chatbots tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
OpenClaw★ 391kOpenClaw is a self-hosted personal AI assistant that answers you on WhatsApp, Telegram, Slack, Discord, and many other channels, with voice and a live visual canvas.
Hermes Agent★ 251kA self-improving personal AI agent from Nous Research that builds skills from experience, remembers across sessions, and reaches you on Telegram, Discord, Slack, and more.
Odysseus★ 88.9kA self-hosted AI workspace that puts chat, agents, deep research, documents, email, notes, tasks and calendar behind one Docker Compose stack, over local or API models.
CowAgent★ 47.2kA self-hosted assistant that plans and executes tasks with built-in file, terminal, browser and search tools, and answers across a web console plus a dozen messaging platforms.
AstrBot★ 41.3kAn all-in-one agent chatbot platform that puts LLM conversations, tools, knowledge bases and a plugin marketplace inside messaging apps like Telegram, Slack, Discord, QQ and Feishu.
OpenHuman★ 40.5kA local-first desktop personal AI for macOS, Windows and Linux that keeps a compressed memory tree on your machine and orchestrates checkpointed research and automation workflows.
MindsHub★ 39.8kAn agent workspace for knowledge work and software development that runs swappable open-source agent harnesses against your choice of frontier or open models.
Parlor★ 2.1kA fully on-device, real-time voice and vision assistant built on Gemma 4 and llama.cpp