← all repositories

toverainc/willow-inference-server

An open-source inference server for running Whisper-based ASR, TTS, and LLM models locally with WebRTC support.

willow-inference-server
Not currently ranked — collecting fresh signals.
star history

Willow Inference Server is a self-hosted language inference system that serves Whisper for speech recognition, TTS for speech synthesis, and LLM models (llama, vicuna). It uses CTranslate2 for optimized Whisper inference and supports multiple transports including WebRTC for real-time streaming, REST, and WebSockets. The server is memory-optimized to load multiple Whisper models and TTS simultaneously within 6GB VRAM and targets CUDA GPUs ranging from consumer to datacenter cards.

Frequently asked

What is toverainc/willow-inference-server?
An open-source inference server for running Whisper-based ASR, TTS, and LLM models locally with WebRTC support.
Is willow-inference-server open source?
Yes — toverainc/willow-inference-server is open source, released under the Apache-2.0 license.
What language is willow-inference-server written in?
toverainc/willow-inference-server is primarily written in Python.
How popular is willow-inference-server?
toverainc/willow-inference-server has 508 stars on GitHub.
Where can I find willow-inference-server?
toverainc/willow-inference-server is on GitHub at https://github.com/toverainc/willow-inference-server.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.