█

AI/TLDR

Strata · 2026-10-04 · major

Strata v0.1.39 — Qwen3.8-Flash-Next 125B on a 12 GB gaming GPU

Strata is an MIT-licensed engine that runs the 125B Qwen3.8-Flash-Next on a gaming PC with a 12 GB+ card. v0.1.39 adds parallel requests, the OpenAI Responses API for Codex CLI, and support for older GPUs.

GitHub social card for the Niko1221/Strata repository

A one-click engine that splits a 125B open model across your GPU, RAM, CPU and SSD so it runs on a normal gaming PC.

Key specs

GitHub stars10,228
Rtx 5070 output (q2 0)94 tok/s

Quick facts

Versionv0.1.39 (4 Oct 2026)
ModelQwen3.8-Flash-Next 125B MoE
Minimum hardware12 GB GPU, 32 GB RAM, ~80 GB disk
PlatformsWindows 10/11, Linux
APIsOpenAI, OpenAI Responses, Anthropic Messages
LicenseMIT

What is it?

Strata v0.1.39 is the newest release of an open-source engine that runs Qwen3.8-Flash-Next, a 125B Mixture-of-Experts model, on consumer hardware. This release adds several requests at once, a POST /v1/responses endpoint so Codex CLI can use it, and experimental support for older NVIDIA, AMD and Intel Arc cards and AVX-only CPUs.

How does it work?

Qwen3.8-Flash-Next has 24,576 experts and uses only 10 per token, so Strata keeps the most-used experts on the graphics card, all experts in RAM with the CPU covering misses, and a lookup table on the SSD. A small helper guesses the next few tokens and the big model checks them in one pass. The engine is built on llama.cpp/ggml.

Why does it matter?

A 125B model at 60–94 tokens per second on a 12–16 GB card means a private, local coding model without a server GPU or API bill. The Responses API and parallel requests in this release make it a direct backend for agents like Codex CLI.

Who is it for?

local-LLM users and developers running coding agents on their own PC

Frequently asked questions

What hardware do I need to run Qwen3.8-Flash-Next with Strata?
Strata needs an NVIDIA RTX 20–50 series or AMD RX 7000/9000 graphics card with at least 12 GB of VRAM, 32 GB of system RAM (64 GB recommended) and about 80 GB of free disk space, ideally on an SSD. Strata v0.1.39 also adds experimental builds for older Pascal and Volta cards, Intel Arc on Linux and CPUs without AVX2.
How fast is Strata on a consumer graphics card?
The Strata README reports 94 tokens per second of output and 2,650 tokens per second of prompt reading on an RTX 5070 with 12 GB using the Q2_0 build, and 60 tokens per second on an AMD RX 9070 XT. Strata v0.1.39 adds about 6% faster decoding on Q2_0 and up to 132% more speed for code on two RTX 4090s.
Can I use Strata with Codex CLI or other coding agents?
Yes. Strata serves an OpenAI-compatible API at 127.0.0.1:8080/v1 and an Anthropic Messages API at /v1/messages. Strata v0.1.39 adds the OpenAI Responses API (POST /v1/responses), which the release notes say was tested with Codex CLI 0.160.0 at about 96% prompt-cache reuse.
What does parallel request support change in Strata v0.1.39?
Strata v0.1.39 can serve several requests at once, set with "parallel": N in the config or setup --parallel N. On an RTX 5070 with Q2_0, four requests at once decode 11% slower in total, but the wait before an answer starts drops from 11.2 seconds to 1.8 seconds.

Try it

git clone https://github.com/Niko1221/Strata && cd Strata && ./setup.sh

Sources · 2 outlets

Tags

  • strata
  • qwen3-8-flash-next
  • qwen
  • local-inference
  • llama-cpp
  • mixture-of-experts
  • speculative-decoding
  • consumer-gpu
  • open-source
  • codex-cli

← All releases · Learn AI