vectorch-ai/ScaleLLM
A high-performance C++ inference runtime for large language models with GPU acceleration and speculative decoding.

Not currently ranked — collecting fresh signals.
star history
ScaleLLM is a production-grade LLM inference system written in C++. It provides GPU acceleration via CUDA for efficient serving of large language models and supports popular open-source models including Llama3.1, Gemma2, and Phi. The system targets production environments with optimizations like speculative decoding for improved throughput.
Frequently asked
- What is vectorch-ai/ScaleLLM?
- A high-performance C++ inference runtime for large language models with GPU acceleration and speculative decoding.
- Is ScaleLLM open source?
- Yes — vectorch-ai/ScaleLLM is open source, released under the Apache-2.0 license.
- What language is ScaleLLM written in?
- vectorch-ai/ScaleLLM is primarily written in C++.
- How popular is ScaleLLM?
- vectorch-ai/ScaleLLM has 499 stars on GitHub.
- Where can I find ScaleLLM?
- vectorch-ai/ScaleLLM is on GitHub at https://github.com/vectorch-ai/ScaleLLM.