horseee/Awesome-Efficient-LLM
A curated collection of papers and projects on efficient LLM techniques including quantization, pruning, knowledge distillation, and inference acceleration.

This repository maintains an organized list of research papers and open-source projects focused on making large language models more efficient. It covers areas such as network pruning, model quantization, knowledge distillation, inference acceleration, efficient architectures, and KV cache compression. The list is structured by sub-topic with separate markdown files and includes a project directory for implementations.
Frequently asked
- What is horseee/Awesome-Efficient-LLM?
- A curated collection of papers and projects on efficient LLM techniques including quantization, pruning, knowledge distillation, and inference acceleration.
- Is Awesome-Efficient-LLM open source?
- Yes — horseee/Awesome-Efficient-LLM is an open-source project tracked on heatdrop.
- What language is Awesome-Efficient-LLM written in?
- horseee/Awesome-Efficient-LLM is primarily written in Python.
- How popular is Awesome-Efficient-LLM?
- horseee/Awesome-Efficient-LLM has 2k stars on GitHub.
- Where can I find Awesome-Efficient-LLM?
- horseee/Awesome-Efficient-LLM is on GitHub at https://github.com/horseee/Awesome-Efficient-LLM.