← all repositories

PaddlePaddle/FastDeploy

High-performance inference and deployment toolkit for LLMs and VLMs built on the PaddlePaddle framework.

FastDeploy
Not currently ranked — collecting fresh signals.
star history

FastDeploy is a production-focused deployment suite for large language models and vision-language models. It provides quantized model serving, OpenAI-compatible API endpoints, and support for popular models including ERNIE, Qwen3-VL, and their variants. The toolkit targets GPU-based inference optimization with techniques like W4AFP8 quantization and integrates with vLLM for serving workflows.

Frequently asked

What is PaddlePaddle/FastDeploy?
High-performance inference and deployment toolkit for LLMs and VLMs built on the PaddlePaddle framework.
Is FastDeploy open source?
Yes — PaddlePaddle/FastDeploy is open source, released under the Apache-2.0 license.
What language is FastDeploy written in?
PaddlePaddle/FastDeploy is primarily written in Python.
How popular is FastDeploy?
PaddlePaddle/FastDeploy has 3.7k stars on GitHub.
Where can I find FastDeploy?
PaddlePaddle/FastDeploy is on GitHub at https://github.com/PaddlePaddle/FastDeploy.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.