← all repositories
gufo-org/gufo

One chip, 128 GiB, and a very short model list

Gufo ditches portability on purpose: a C++ inference engine built solely for AMD's Strix Halo APU, where 128 GiB of unified memory makes a dedicated engine worth writing.

★533 stars C++ Inference · Serving
gufo
Collecting fresh signals — velocity needs a few days of history.
collecting data…
star history

What it does Gufo is a C++ local inference server for exactly one machine: AMD Strix Halo systems (Ryzen AI MAX+ 395 with the Radeon 8060S, gfx1151) and their up to 128 GiB of unified memory. It serves a deliberately curated model list — Qwen3.8 27B, Qwen3.8 Flash-Next, DeepSeek V4 Flash, Qwen3-ASR, Qwen3-TTS — through an OpenAI-compatible API, with Qwen-Image-2.1 and MiniMax H3 (image and video) still in progress. It ships as a container image and builds from source via CMake or Nix; Linux on Strix Halo is the only supported target.

The interesting bit The whole bet is verticality: every kernel is tuned for this one piece of silicon, and kernels are deliberately not shared between models, so a change to one model can’t quietly break another. Each model ships a quality report with independent numerical checks, so the speed claims come with receipts. Single-user decode rates lean on speculative modes — DFlash2, MTP, DSpark — which is where the fast decode numbers come from.

Key highlights

  • Qwen3.8 Flash-Next Q4: 1,628.52 tok/s prompt processing; 59.41 tok/s single-user decode; 157.22 tok/s aggregated across 8 concurrent requests (MTP).
  • Qwen3.8 27B Q4: 656.33 tok/s pp; 70.56 tok/s single-user decode with DFlash2; 123.00 tok/s aggregated on 8 requests.
  • Speech is first-class: ASR at 15.27× realtime, TTS up to 2.54× realtime with 201 ms to first audio and voice cloning support.
  • Concurrency, cancellation, and conversation caching are treated as first-class workloads, not afterthoughts.
  • Attention and audio convolution kernels compile straight from HIP — no Python, Triton, or MIOpen in the production runtime.

Caveats

  • The published figures are peaks, and the README is upfront that peaks include repetitive output; the 27B single-user number also uses a short-prompt workload. Treat them as ceilings, not typical chat throughput.
  • Image and video generation are listed as in progress, so the text-and-audio story is what’s real today.
  • Windows and a dual-Strix-Halo RDMA variant exist only as unmerged forks.

Verdict If you own a Strix Halo box with the big memory configuration, this is the tuned local stack aimed squarely at you. If you’re on anything else, keep walking — the single-chip tunnel vision is the feature, not a flaw.

Frequently asked

What is gufo-org/gufo?
Gufo ditches portability on purpose: a C++ inference engine built solely for AMD's Strix Halo APU, where 128 GiB of unified memory makes a dedicated engine worth writing.
Is gufo open source?
Yes — gufo-org/gufo is open source, released under the MIT license.
What language is gufo written in?
gufo-org/gufo is primarily written in C++.
How popular is gufo?
gufo-org/gufo has 533 stars on GitHub.
Where can I find gufo?
gufo-org/gufo is on GitHub at https://github.com/gufo-org/gufo.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.