← all repositories
warpfront/hipfire

A Rust LLM engine that actually prefers your AMD gaming GPU

Most LLM inference stacks treat consumer AMD GPUs as second-class citizens; hipfire targets the full RDNA family with a single Rust binary that keeps Python out of the hot path.

hipfire
Not currently ranked — collecting fresh signals.
star history

What it does

hipfire is a single-binary LLM inference engine written in Rust for AMD’s RDNA GPUs, covering RDNA1 through RDNA4 including consumer cards, pro variants, and APUs. It avoids the standard ROCm userspace stack at runtime by shipping pre-compiled kernel blobs where possible and JIT-compiling the rest through HIP. The tool handles model pulling, prompt execution, and an OpenAI-compatible HTTP API without Python or PyTorch touching the hot path.

The interesting bit

Instead of treating consumer graphics as a compatibility afterthought, the project makes them the primary target, borrowing architecture-specific optimizations from community Vega 20 forks and speculative decode techniques inspired by the Lucebox project. It also implements CASK-based KV cache eviction via optional sidecars to keep long-context work from OOM-ing on limited VRAM.

Key highlights

  • On a 7900 XTX, decode throughput reaches up to 2.10× Ollama’s Q4_K_M speed for a 0.8B Qwen 3.5 model, and 1.71× for the 9B variant; DFlash speculative decoding peaks at 372 tok/s on that same 9B model.
  • Ships with first-class NixOS support via a flake and system module, alongside pipeline-parallel multi-GPU deployment and asymmetric KV cache quantization.
  • The codebase is dual-licensed under MIT or Apache-2.0, with an explicit prior-art catalogue and documented governance around its 2026 license correction.

Caveats

  • DFlash speculative decode gains are genre-conditional; the README warns that speedups vary significantly by prompt type.
  • Performance numbers are self-reported on a single gfx1100 card (7900 XTX) against Ollama, so broader hardware comparisons remain thin.

Verdict

Worth a look if you run local models on consumer AMD hardware and want to drop the ROCm Python stack. If your workflow is CUDA-only or cloud-datacenter bound, this is not your scene.

Frequently asked

What is warpfront/hipfire?
Most LLM inference stacks treat consumer AMD GPUs as second-class citizens; hipfire targets the full RDNA family with a single Rust binary that keeps Python out of the hot path.
Is hipfire open source?
Yes — warpfront/hipfire is an open-source project tracked on heatdrop.
What language is hipfire written in?
warpfront/hipfire is primarily written in Rust.
How popular is hipfire?
warpfront/hipfire has 653 stars on GitHub.
Where can I find hipfire?
warpfront/hipfire is on GitHub at https://github.com/warpfront/hipfire.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.