PaddlePaddle/FastDeploy
High-performance inference and deployment toolkit for LLMs and VLMs built on the PaddlePaddle framework.

Not currently ranked — collecting fresh signals.
star history
FastDeploy is a production-focused deployment suite for large language models and vision-language models. It provides quantized model serving, OpenAI-compatible API endpoints, and support for popular models including ERNIE, Qwen3-VL, and their variants. The toolkit targets GPU-based inference optimization with techniques like W4AFP8 quantization and integrates with vLLM for serving workflows.
Frequently asked
- What is PaddlePaddle/FastDeploy?
- High-performance inference and deployment toolkit for LLMs and VLMs built on the PaddlePaddle framework.
- Is FastDeploy open source?
- Yes — PaddlePaddle/FastDeploy is open source, released under the Apache-2.0 license.
- What language is FastDeploy written in?
- PaddlePaddle/FastDeploy is primarily written in Python.
- How popular is FastDeploy?
- PaddlePaddle/FastDeploy has 3.7k stars on GitHub.
- Where can I find FastDeploy?
- PaddlePaddle/FastDeploy is on GitHub at https://github.com/PaddlePaddle/FastDeploy.