Overview
Strata is an open-source inference engine built for one model family: Qwen3.8-Flash-Next, the Qwen team's 125-billion-parameter Mixture-of-Experts model. Its goal is to run that model on an ordinary gaming PC — a 12 GB or larger NVIDIA or AMD graphics card plus 32 GB of RAM — with nothing leaving the machine. A one-click installer for Windows and Linux downloads the model, picks settings for the hardware and starts a local server with a chat page in the browser.
The trick is to split the model across the whole computer. Qwen3.8-Flash-Next is made of 24,576 small experts and asks only 10 of them for each token, so Strata keeps the most-used experts on the graphics card, holds all of them in RAM, lets the CPU run the ones the card does not have, and reads a large lookup table from the SSD a few rows at a time. On top of that it uses speculative decoding: a small helper guesses the next few tokens and the big model checks them all in one pass, keeping the ones it agrees with.

The README reports 94 tokens per second of output on an RTX 5070 (12 GB) with the 2-bit Q2_0 build and 60 tokens per second on an AMD RX 9070 XT. Strata exposes an OpenAI-compatible API and an Anthropic Messages API on localhost, so coding agents and other apps can point at it, and it accepts image input. It is built on llama.cpp/ggml, uses compressed model files from ISTA-DASLab, UkisAI and Unsloth, and is released under the MIT licence.
What it does
- Runs the 125B Qwen3.8-Flash-Next on a 12 GB+ NVIDIA (RTX 20–50 series) or AMD (RX 7000/9000) graphics card with 32 GB+ RAM
- Places hot experts on the GPU, all experts in RAM, CPU fallback for misses, and a lookup table read from the SSD
- Speculative decoding: a small helper guesses ahead and the big model verifies several tokens per step

- OpenAI-compatible /v1 API, OpenAI Responses API and Anthropic /v1/messages on localhost
- Optional image input, parallel requests and a browser chat and monitor page
- Model choices include Q2_0 and IQ2/IQ3 compressions, a Coder variant, Swift 1.5 and Unsloth 4-bit builds
Getting started
Strata installs from a cloned repository with one script per platform. The installer asks which model build, context window and image support you want, then downloads about 70 GB (the download resumes if interrupted). You need roughly 80 GB of free disk space, ideally on an SSD.
Run the installer on Windows
From the repository folder, start the setup script. It checks the hardware and recommends a model build.
START-HERE.batOr on Linux
The Linux script does the same checks and download.
./setup.shOpen the chat or point an app at the API
Once the engine is running, the browser interface and the OpenAI-compatible API are on localhost; Anthropic-style clients use /v1/messages.
http://127.0.0.1:8080 # chat + monitor
http://127.0.0.1:8080/v1 # OpenAI-compatible API
http://127.0.0.1:8080/v1/messages # Anthropic APIServe other machines on your network
Re-run setup with a host and an API key to expose the server beyond localhost.
START-HERE.bat --setup --host 0.0.0.0 --api-key <secret>Keep it updated
Releases land often; the update script pulls the newest engine.
UPDATE.bat # Windows
./update.sh # LinuxCommands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Run a 125B open-weight model on a single gaming PC instead of paying for API tokens
- Back a local coding agent such as Codex CLI with a private model through the OpenAI or Anthropic API
- Keep sensitive code and documents on your own machine while still using a large model
- Share one home GPU box as a private model server for other machines on the network
How Strata compares
Strata alongside other open-source local runtimes tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| Ollama | ★ 182k | A developer-friendly tool that downloads and runs local LLMs from the terminal with a built-in OpenAI-compatible API. |
| llama.cpp | ★ 130k | A C/C++ inference engine that runs LLMs in the GGUF format on CPUs, Apple Silicon, and GPUs with low memory use. |
| GPT4All | ★ 77.4k | GPT4All is a free desktop app and Python client that runs large language models locally on your own computer, with no API calls or GPU required. |
| LocalAI | ★ 49.4k | A self-hosted server that exposes an OpenAI-compatible API for running text, vision, voice, and image models on local hardware. |
| Jan | ★ 44.8k | An open-source desktop app that runs LLMs fully offline as a ChatGPT-style assistant on your own computer. |
| Colibrì | ★ 39.5k | A pure-C inference engine that keeps a Mixture-of-Experts model's dense trunk resident in RAM and streams its routed experts from disk, so 744B-2.8T models run on consumer hardware. |
| llmfit | ★ 37.6k | A Rust terminal tool that inspects your CPU, RAM, GPUs and VRAM and scores which open-weight models and quantizations will actually run well on that machine, with a TUI, CLI, REST API and local-runtime integrations. |
| Strata | — | A one-click local inference engine that runs the 125B Qwen3.8-Flash-Next on a gaming PC with a 12 GB graphics card |
