← all repositories

incoai/splash

A local LLM inference engine for Apple silicon that runs coding agents and OpenAI/Anthropic-compatible applications on a single Mac.

splash
Collecting fresh signals — velocity needs a few days of history.
collecting data…
star history

Splash is a local inference engine targeting Apple silicon Macs. It combines DFlash 2 speculative decoding, specialized Metal kernels, and automatic memory planning to serve LLMs with vision, tool calling, and a built-in chat page. It exposes OpenAI Chat Completions, Responses, and Completions endpoints as well as Anthropic Messages, with streaming, JSON Schema output, image and PDF inputs, prefix caching, and automatic request batching. The project ships CLI helpers to launch coding agents such as opencode, claude, codex, hermes, and pi against the local server.

Frequently asked

What is incoai/splash?
A local LLM inference engine for Apple silicon that runs coding agents and OpenAI/Anthropic-compatible applications on a single Mac.
Is splash open source?
Yes — incoai/splash is open source, released under the Apache-2.0 license.
What language is splash written in?
incoai/splash is primarily written in C++.
How popular is splash?
incoai/splash has 1k stars on GitHub.
Where can I find splash?
incoai/splash is on GitHub at https://github.com/incoai/splash.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.