← all repositories

intel/xFasterTransformer

An optimized inference solution for running large language models on Intel X86/Xeon platforms.

xFasterTransformer
Not currently ranked — collecting fresh signals.
star history

xFasterTransformer provides a GPU FasterTransformer-equivalent solution for CPU-based LLM inference on Intel Xeon processors. It supports distributed inference across multiple sockets and nodes for running larger models, and offers both C++ and Python APIs spanning from high-level to low-level interfaces. The project supports popular model architectures including Qwen, DeepSeek-R1, ChatGLM, and LLaMA, and includes integration with vLLM for OpenAI-compatible serving.

Frequently asked

What is intel/xFasterTransformer?
An optimized inference solution for running large language models on Intel X86/Xeon platforms.
Is xFasterTransformer open source?
Yes — intel/xFasterTransformer is open source, released under the Apache-2.0 license.
What language is xFasterTransformer written in?
intel/xFasterTransformer is primarily written in C++.
How popular is xFasterTransformer?
intel/xFasterTransformer has 436 stars on GitHub.
Where can I find xFasterTransformer?
intel/xFasterTransformer is on GitHub at https://github.com/intel/xFasterTransformer.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.