125 Billion Parameters, One Gaming PC, No Server in Sight

Strata tiers a huge mixture-of-experts model across GPU, RAM, and SSD so a server-class AI runs — briskly — on hardware you already own.
The local-LLM scene has spent years reconciling itself to a modest ceiling: if you want to run models at home, you run the small ones. The honest self-hosters’ takes all converge on the same trade-off — privacy and predictable cost in exchange for a model that’s noticeably dumber than what the cloud offers. A widely shared guide to local LLMs in 2026 frames the choice as picking your compromise: which tool, which hardware, which quantization of a 7B-to-70B model you can tolerate. Meanwhile, forum threads like this hardware-buying discussion on Hugging Face show people agonizing over unified-memory machines — a DGX Spark, a Mac Studio — precisely because big models don’t fit the consumer-GPU-plus-system-RAM arrangement most of us already have.

Strata, a small open-source project by a single developer, walks into that conversation and makes an impolite claim: a 125-billion-parameter model — Qwen3.8-Flash-Next, the kind of thing that normally lives on a server with hundreds of gigabytes of graphics memory — can run on a gaming PC with one 12 GB NVIDIA card and 64 GB of system RAM. Not crawl. Run, at 60–95 tokens per second, which is faster than most people read.
If that claim survives scrutiny, it doesn’t just add another option to the local-LLM menu. It moves the ceiling.
Layers, as advertised
The name is not decoration. A stratum, per the geologists, is a layer of material with internally consistent characteristics, stacked in parallel layers one upon another. (The internet, unhelpfully, will also tell you a strata is a breakfast casserole of bread and custard — several food sites are eager to explain this — but the geological reading is the one that matters here.) Strata the project is literally about layers: a memory hierarchy in which the model’s weights are distributed across three tiers of your machine, each doing the job it’s suited for.
The README’s own analogy is a kitchen. The things you use constantly stay on the counter; the rest waits in the pantry. It’s a good analogy, and it’s also where the actual engineering lives — the boring, load-bearing part that most hype writeups skip.
The insight: not all 125 billion parameters show up to work
The whole scheme rests on a property of the model itself. Qwen3.8-Flash-Next is a mixture-of-experts architecture: 24,576 small specialist sub-networks, of which only 10 are consulted for any given token. That’s the crucial asymmetry. A dense 125B model needs all 125 billion parameters in fast memory for every single word it writes. A sparse one needs a rotating cast of roughly ten experts out of twenty-five thousand.
Strata exploits that rotation. The graphics card keeps the few thousand experts that get asked for most often — and, notably, keeps learning which ones those are while you use it, so the residency set adapts to your workload. System RAM holds every expert, so nothing is ever more than main memory away. When a token needs an expert the card doesn’t have, the CPU computes it, in parallel with the GPU’s work, so neither sits idle waiting for the other. And the SSD holds a large lookup table — 29 GB of it — from which only a few small rows are read per token.
This is the same family of tricks that offloading engines in the llama.cpp ecosystem have been refining for a while, and Strata is candid about its lineage: it’s built on parts of llama.cpp/ggml, and the README credits ideas from three other projects — Splash, ninfer, and HyperQwen. The contribution here is integration and tuning, not a new theory of inference. That matters for calibrating the hype, and we’ll come back to it.
The second trick is speculative decoding, described in the README as “guess, then check.” A small, fast helper model built into the package guesses the next few words; the big model verifies all the guesses in one pass, keeps the ones it agrees with, and writes the next word itself. Because the big model still decides every word, the claim is that output quality is unchanged — you just get each word 1.6–1.8 times sooner. Prompt ingestion gets its own optimization: long inputs are read in chunks of up to 8,192 tokens at a time, which is why a 32,000-token document is processed at over 1,000 tokens per second rather than the trickle that makes local-model users contemplate their life choices while a summary finishes.
The numbers, and the asterisks
The published measurements come from one machine: an RTX 5070 (12 GB), a Ryzen 5 7600, 64 GB of RAM. On it, the recommended IQ2_XS build writes at 74 tokens per second in short chat and 60 at a 128K context; the most aggressive compression, Q2_0, hits 90. A 24 GB card is estimated — estimated, not measured — at 100–140 tokens per second, on the reasoning that more of the model on the GPU means fewer trips to RAM.
The compression ladder is where users will make their real trade. Four sizes of the same model, from Q2_0 (37.6 GB of combined memory, fastest, “good”) up to IQ3_S (54.8 GB, slowest, and — per the README — matching the full model on the published tests, though that claim is qualified as applying to the original model only). The fit question is governed by RAM, not VRAM: shard one’s experts load into system memory at startup, and a bigger graphics card makes things faster without lowering the RAM floor. Sixty-four gigabytes fits everything; forty-eight fits the two smaller sizes.
There are also two derived variants worth knowing about, both from third parties. A “Coder” build from ISTA-DASLab deletes half the experts, keeping the ones that coding, tool use, and image understanding lean on — its authors report 91% of the full model’s SWE-bench Verified score and 99% of LiveCodeBench, at a memory footprint that fits a 32 GB machine. And a fine-tune called Swift 1.5, from UkisAI, shortens the model’s chain-of-thought so answers arrive sooner at roughly the same quality. These numbers are the variant authors’, not Strata’s, and the README is appropriately careful about attributing them.
What Strata actually is
Strip away the packaging and Strata is orchestration: a curated bundle of someone else’s model (Qwen’s), someone else’s quantizations (ISTA-DASLab’s), and an inference engine descended from llama.cpp, wrapped in an installer, a browser app, and an OpenAI-compatible API endpoint so coding agents and ordinary applications can treat it like any other provider. It is, in the pejorative sense, glue code.
But that’s the wrong sense. The hard problem in local inference at this scale has never been a missing algorithm — offloading and speculative decoding are established techniques. The hard problem is that assembling them correctly, calibrating them to specific hardware, and packaging the result so a non-expert can actually use it is genuinely difficult work that almost nobody does well. Strata’s value is that it makes a configuration that was previously a weekend of forum archaeology into something that behaves like a product. The README even includes a calibration mode that benchmarks engine settings on your specific machine and keeps the fastest — on the reference PC it found a 7% gain. That’s the unglamorous 7% that separates a demo from a tool.
The honesty of the project’s credits section is itself a signal: a repo that names its influences and licenses, separates its own MIT code from the model files it merely distributes pointers to, and flags an experimental feature (a “speed projection” control vector, off by default, derived from the model’s activations under a separate Qwen license) is a repo that expects to be read by people who check.
The rough edges, visible in plain text
To its credit, the README doesn’t hide the costs. First startup loads 35–55 GB into RAM and locks part of it for the GPU, during which the PC “can be slow or stop responding for 1–3 minutes” — the mouse may freeze. The model download is around 70 GB, with roughly 80 GB of disk expected. The server answers one request at a time, which is fine for a personal endpoint and disqualifying for a team. Out-of-RAM conditions are a recurring theme in the troubleshooting section, including an engine that dies mid-answer and needs the message re-sent. These are the honest frictions of running a server-class model on hardware that was never promised to it.
The deeper caveat is provenance. Every performance and quality number in the README comes from the project’s own single-machine measurements or from the variant authors’ claims; there’s no independent benchmarking here, and the broader community’s verdict — the Reddit local-LLM threads where this kind of thing gets stress-tested by thousands of opinionated hobbyists — wasn’t accessible for this piece, so the crowd’s reaction remains an open question. The 125B-on-a-gaming-PC claim deserves independent replication before it’s treated as settled fact.
Why the timing matters
The local-inference community has been circling this exact problem from the other direction. The self-hoster’s honest assessment of the scene is blunt that local models are smaller models, and that the largest thing you can reasonably run — a 70B — is still smaller than what commercial APIs serve. Enthusiast writeups of daily local-LLM workflows describe real, useful projects built on 8B-class models, with RAM and speed as the acknowledged taxes. And the hardware-advice ecosystem — buying guides, forum threads, an entire genre of what-to-buy-for-local-LLMs content — has been steering people toward expensive unified-memory machines as the only path to big models.
Strata’s argument, whether or not it fully holds, is that this framing is stale. If a sparse 125B model can be tiered across a 12 GB card, 64 GB of RAM, and an SSD, then the machine that runs frontier-adjacent open weights isn’t a $4,000 appliance — it’s the gaming PC a lot of people already have, plus a RAM upgrade. That reframing ripples outward: it changes what hardware advice should say, what the minimum viable local setup looks like, and how big the gap between local and cloud really is in 2026.
The open questions
Three things will decide whether Strata is a moment or a fixture. First, quality at the low end: the whole pitch rests on aggressive quantization, and “good” at Q2_0 is doing a lot of work in that table — the README’s own recommendation hedges toward IQ2_XS, and the strongest quality claim is reserved for the size that needs 55 GB. Second, the project’s dependence on its model: Strata is a delivery mechanism, and its fortunes ride on Qwen’s release cadence and on third parties continuing to produce these cleverly pruned and quantized variants. Third, single-user throughput caps how far this scales beyond the enthusiast desk.
None of those are fatal. They’re the normal shape of a young project whose core bet — that memory hierarchy, not raw VRAM, is the real constraint on local AI — is looking increasingly like the right bet. The geologists would appreciate the naming: the value is all in the layers.
Sources
- Breakfast Strata Recipe
- What to Buy for Local LLMs (April 2026) - Julien Simon - Medium
- What is your actual daily use case for local LLMs?
- Stratum
- Do you think dedicated hardware for running local LLMs will become ...
- Running local LLM: Best use cases and tools
- Build-Your-Own Strata
- BUYING ADVICE for local LLM machine - Beginners
- Local LLMs for self-hosters: what's worth running at home
- What The Heck Is Strata???
- Guide to Local LLMs in 2026: Privacy, Tools & Hardware
- How I Use Local LLMs for Real Projects (and Why You ...