Your junk-drawer Android is now a 5-watt LLM server
OlliteRT turns a spare Android phone into a fully local, OpenAI- and Anthropic-compatible LLM server — no cloud, no API keys, just your LAN.

What it does It’s an Android app that runs LLMs on your phone’s GPU or CPU via Google’s LiteRT-LM runtime and serves them over the local network as a standard OpenAI-compatible HTTP API — the README calls it “Ollama for Android,” which is about right. Any client that speaks OpenAI (Open WebUI, Home Assistant, Python, curl) points at the phone and gets chat completions, streaming, and audio transcription. It also implements the Anthropic Messages API, so Claude-shaped clients work too. Everything runs on-device, and the privacy policy is short because there’s nothing to collect.
The interesting bit The economics are the pitch: roughly 5–10W versus 300W+ for a GPU server, per the README, which reframes the phone in your drawer as an always-on home inference box — one the author explicitly begs you not to run under your pillow. It ships more ops tooling than some production servers: bearer token auth, client IP access rules, Prometheus metrics, and Home Assistant remote control. Under the hood it’s honest glue — model execution belongs to Google’s LiteRT-LM, and the app builds on Google’s AI Edge Gallery — with the value sitting in the Ktor server, model management, and monitoring wrapped around it.
Key highlights
- Speaks OpenAI (
/v1/chat/completions,/v1/completions,/v1/responses) and Anthropic (/v1/messages), streaming included, plus/v1/audio/transcriptions - Model roster from Gemma 4 E2B (2.4 GB, 32K context, vision/audio/thinking/tools) down to Gemma 3 1B at 0.5 GB; one-tap HuggingFace downloads or
.litertlmimports - Server-grade touches: bearer token auth, IP rules, 29 Prometheus metrics, Home Assistant REST control, auto-start on boot, idle model unload
- Built-in benchmark to compare models on your actual hardware before committing the storage
- Runs fully offline; tool calling is experimental and model-dependent, which the README flags rather than hides
Caveats
- One model loaded at a time and requests queue sequentially — a LiteRT SDK limitation, so forget concurrency
- Token counts are estimated as characters ÷ 4 because LiteRT exposes no tokenizer API; fine for English, shakier for code
- arm64-v8a only, Android 12+, 6 GB RAM minimum, and no GGUF —
.litertlmmodels only, so your existing Ollama library doesn’t transfer
Verdict Worth a look if you have a spare arm64 phone and want a private LAN LLM for Home Assistant or Open WebUI without feeding a GPU. Skip it if you need concurrent requests, GGUF support, or logprobs — the LiteRT runtime’s constraints are real, and the README deserves credit for listing every one of them.
Frequently asked
- What is NightMean/OlliteRT?
- OlliteRT turns a spare Android phone into a fully local, OpenAI- and Anthropic-compatible LLM server — no cloud, no API keys, just your LAN.
- Is OlliteRT open source?
- Yes — NightMean/OlliteRT is open source, released under the Apache-2.0 license.
- What language is OlliteRT written in?
- NightMean/OlliteRT is primarily written in Kotlin.
- How popular is OlliteRT?
- NightMean/OlliteRT has 508 stars on GitHub.
- Where can I find OlliteRT?
- NightMean/OlliteRT is on GitHub at https://github.com/NightMean/OlliteRT.