← all repositories

vectorch-ai/ScaleLLM

A high-performance C++ inference runtime for large language models with GPU acceleration and speculative decoding.

499 stars C++ Inference · Serving
ScaleLLM
Not currently ranked — collecting fresh signals.
star history

ScaleLLM is a production-grade LLM inference system written in C++. It provides GPU acceleration via CUDA for efficient serving of large language models and supports popular open-source models including Llama3.1, Gemma2, and Phi. The system targets production environments with optimizations like speculative decoding for improved throughput.

Frequently asked

What is vectorch-ai/ScaleLLM?
A high-performance C++ inference runtime for large language models with GPU acceleration and speculative decoding.
Is ScaleLLM open source?
Yes — vectorch-ai/ScaleLLM is open source, released under the Apache-2.0 license.
What language is ScaleLLM written in?
vectorch-ai/ScaleLLM is primarily written in C++.
How popular is ScaleLLM?
vectorch-ai/ScaleLLM has 499 stars on GitHub.
Where can I find ScaleLLM?
vectorch-ai/ScaleLLM is on GitHub at https://github.com/vectorch-ai/ScaleLLM.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.