microsoft/sarathi-serve
Microsoft's Sarathi-Serve is a research LLM serving framework optimized for throughput-latency tradeoff, originally forked from vLLM.

Not currently ranked — collecting fresh signals.
star history
Sarathi-Serve is a high-throughput and low-latency serving engine for large language models. The project is a research prototype that originated as a fork of vLLM, adapted specifically to tame the throughput-latency tradeoff in LLM inference. It targets H100 and A100 GPUs and was developed as part of an academic paper published at the 2024 USENIX OSDI conference.
Frequently asked
- What is microsoft/sarathi-serve?
- Microsoft's Sarathi-Serve is a research LLM serving framework optimized for throughput-latency tradeoff, originally forked from vLLM.
- Is sarathi-serve open source?
- Yes — microsoft/sarathi-serve is open source, released under the Apache-2.0 license.
- What language is sarathi-serve written in?
- microsoft/sarathi-serve is primarily written in Python.
- How popular is sarathi-serve?
- microsoft/sarathi-serve has 512 stars on GitHub.
- Where can I find sarathi-serve?
- microsoft/sarathi-serve is on GitHub at https://github.com/microsoft/sarathi-serve.