← all repositories

NVIDIA/FasterTransformer

NVIDIA's optimized transformer inference library for BERT and GPT models on Volta, Turing, and Ampere GPUs.

FasterTransformer
Not currently ranked — collecting fresh signals.
star history

FasterTransformer provides highly optimized encoder and decoder transformer components for inference on NVIDIA GPUs. It supports BERT and GPT model families and integrates with PyTorch and TensorFlow. The library has transitioned development to TensorRT-LLM but remains available for existing use cases.

Frequently asked

What is NVIDIA/FasterTransformer?
NVIDIA's optimized transformer inference library for BERT and GPT models on Volta, Turing, and Ampere GPUs.
Is FasterTransformer open source?
Yes — NVIDIA/FasterTransformer is open source, released under the Apache-2.0 license.
What language is FasterTransformer written in?
NVIDIA/FasterTransformer is primarily written in C++.
How popular is FasterTransformer?
NVIDIA/FasterTransformer has 6.4k stars on GitHub.
Where can I find FasterTransformer?
NVIDIA/FasterTransformer is on GitHub at https://github.com/NVIDIA/FasterTransformer.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.