← all repositories

huggingface/text-generation-inference

A Rust and Python gRPC inference server for serving large language models with optimization for text generation workloads.

text-generation-inference
Not currently ranked — collecting fresh signals.
star history

Text Generation Inference (TGI) is a production inference server designed for serving large language models. Built with Rust for performance-critical components and Python for higher-level logic, it provides gRPC APIs for text generation. The system includes optimizations for quantization and supports major model architectures including transformers, Falcon, BLOOM, StarCoder, and GPT variants. It is the backbone infrastructure powering Hugging Face’s production services including Hugging Chat and the Inference API.

Frequently asked

What is huggingface/text-generation-inference?
A Rust and Python gRPC inference server for serving large language models with optimization for text generation workloads.
Is text-generation-inference open source?
Yes — huggingface/text-generation-inference is open source, released under the Apache-2.0 license.
What language is text-generation-inference written in?
huggingface/text-generation-inference is primarily written in Python.
How popular is text-generation-inference?
huggingface/text-generation-inference has 10.9k stars on GitHub.
Where can I find text-generation-inference?
huggingface/text-generation-inference is on GitHub at https://github.com/huggingface/text-generation-inference.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.