← all repositories

microsoft/sarathi-serve

Microsoft's Sarathi-Serve is a research LLM serving framework optimized for throughput-latency tradeoff, originally forked from vLLM.

512 stars Python Inference · Serving
sarathi-serve
Not currently ranked — collecting fresh signals.
star history

Sarathi-Serve is a high-throughput and low-latency serving engine for large language models. The project is a research prototype that originated as a fork of vLLM, adapted specifically to tame the throughput-latency tradeoff in LLM inference. It targets H100 and A100 GPUs and was developed as part of an academic paper published at the 2024 USENIX OSDI conference.

Frequently asked

What is microsoft/sarathi-serve?
Microsoft's Sarathi-Serve is a research LLM serving framework optimized for throughput-latency tradeoff, originally forked from vLLM.
Is sarathi-serve open source?
Yes — microsoft/sarathi-serve is open source, released under the Apache-2.0 license.
What language is sarathi-serve written in?
microsoft/sarathi-serve is primarily written in Python.
How popular is sarathi-serve?
microsoft/sarathi-serve has 512 stars on GitHub.
Where can I find sarathi-serve?
microsoft/sarathi-serve is on GitHub at https://github.com/microsoft/sarathi-serve.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.