█

AI/TLDR

Strata

A one-click local inference engine that runs the 125B Qwen3.8-Flash-Next on a gaming PC with a 12 GB graphics card

Strata's browser monitor on an RTX 5070 showing 50.4 tokens per second, 11.2 of 12 GB VRAM used with 1,644 experts cached, beside a terminal coding agent writing a voxel pagoda garden
Strata's monitor page while a local coding agent works: speed, VRAM, cached experts, PCIe, CPU and disk reads, plus recent requests.Strata README ↗
Local RuntimesOpen source
Latest
v0.1.39
Updated
4 Oct 2026
Language
C++
License
MIT
Coverage
1 story

What's new

v0.1.394 Oct 2026

Strata v0.1.39 adds parallel requests, the OpenAI Responses API for Codex CLI, experimental support for older NVIDIA, AMD and Intel Arc cards and AVX-only CPUs, and faster decoding on long prompts and multi-GPU setups.

Latest news

Overview

Strata is an open-source inference engine built for one model family: Qwen3.8-Flash-Next, the Qwen team's 125-billion-parameter Mixture-of-Experts model. Its goal is to run that model on an ordinary gaming PC — a 12 GB or larger NVIDIA or AMD graphics card plus 32 GB of RAM — with nothing leaving the machine. A one-click installer for Windows and Linux downloads the model, picks settings for the hardware and starts a local server with a chat page in the browser.

The trick is to split the model across the whole computer. Qwen3.8-Flash-Next is made of 24,576 small experts and asks only 10 of them for each token, so Strata keeps the most-used experts on the graphics card, holds all of them in RAM, lets the CPU run the ones the card does not have, and reads a large lookup table from the SSD a few rows at a time. On top of that it uses speculative decoding: a small helper guesses the next few tokens and the big model checks them all in one pass, keeping the ones it agrees with.

Diagram of how Strata splits a model of 24,576 experts across the graphics card, RAM plus processor, and a 29 GB SSD lookup table
Only 10 of 24,576 experts work on each word, so the busy ones live on the GPU, the rest in RAM, and a lookup table stays on the SSD.Strata README ↗

The README reports 94 tokens per second of output on an RTX 5070 (12 GB) with the 2-bit Q2_0 build and 60 tokens per second on an AMD RX 9070 XT. Strata exposes an OpenAI-compatible API and an Anthropic Messages API on localhost, so coding agents and other apps can point at it, and it accepts image input. It is built on llama.cpp/ggml, uses compressed model files from ISTA-DASLab, UkisAI and Unsloth, and is released under the MIT licence.

What it does

  • Runs the 125B Qwen3.8-Flash-Next on a 12 GB+ NVIDIA (RTX 20–50 series) or AMD (RX 7000/9000) graphics card with 32 GB+ RAM
  • Places hot experts on the GPU, all experts in RAM, CPU fallback for misses, and a lookup table read from the SSD
  • Speculative decoding: a small helper guesses ahead and the big model verifies several tokens per step
Diagram of Strata's guess-then-check decoding: a small helper guesses four words and the big model accepts three of them in one step
Guess, then check: the big model still decides every word, but several arrive per step.Strata README ↗
  • OpenAI-compatible /v1 API, OpenAI Responses API and Anthropic /v1/messages on localhost
  • Optional image input, parallel requests and a browser chat and monitor page
  • Model choices include Q2_0 and IQ2/IQ3 compressions, a Coder variant, Swift 1.5 and Unsloth 4-bit builds

Getting started

Strata installs from a cloned repository with one script per platform. The installer asks which model build, context window and image support you want, then downloads about 70 GB (the download resumes if interrupted). You need roughly 80 GB of free disk space, ideally on an SSD.

Run the installer on Windows

From the repository folder, start the setup script. It checks the hardware and recommends a model build.

bashbash
START-HERE.bat

Or on Linux

The Linux script does the same checks and download.

bashbash
./setup.sh

Open the chat or point an app at the API

Once the engine is running, the browser interface and the OpenAI-compatible API are on localhost; Anthropic-style clients use /v1/messages.

texttext
http://127.0.0.1:8080        # chat + monitor
http://127.0.0.1:8080/v1     # OpenAI-compatible API
http://127.0.0.1:8080/v1/messages  # Anthropic API

Serve other machines on your network

Re-run setup with a host and an API key to expose the server beyond localhost.

bashbash
START-HERE.bat --setup --host 0.0.0.0 --api-key <secret>

Keep it updated

Releases land often; the update script pulls the newest engine.

bashbash
UPDATE.bat      # Windows
./update.sh     # Linux

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Run a 125B open-weight model on a single gaming PC instead of paying for API tokens
  • Back a local coding agent such as Codex CLI with a private model through the OpenAI or Anthropic API
  • Keep sensitive code and documents on your own machine while still using a large model
  • Share one home GPU box as a private model server for other machines on the network

How Strata compares

Strata alongside other open-source local runtimes tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Ollama★ 182kA developer-friendly tool that downloads and runs local LLMs from the terminal with a built-in OpenAI-compatible API.
llama.cpp★ 130kA C/C++ inference engine that runs LLMs in the GGUF format on CPUs, Apple Silicon, and GPUs with low memory use.
GPT4All★ 77.4kGPT4All is a free desktop app and Python client that runs large language models locally on your own computer, with no API calls or GPU required.
LocalAI★ 49.4kA self-hosted server that exposes an OpenAI-compatible API for running text, vision, voice, and image models on local hardware.
Jan★ 44.8kAn open-source desktop app that runs LLMs fully offline as a ChatGPT-style assistant on your own computer.
Colibrì★ 39.5kA pure-C inference engine that keeps a Mixture-of-Experts model's dense trunk resident in RAM and streams its routed experts from disk, so 744B-2.8T models run on consumer hardware.
llmfit★ 37.6kA Rust terminal tool that inspects your CPU, RAM, GPUs and VRAM and scores which open-weight models and quantizations will actually run well on that machine, with a TUI, CLI, REST API and local-runtime integrations.
Strata—A one-click local inference engine that runs the 125B Qwen3.8-Flash-Next on a gaming PC with a 12 GB graphics card