← all repositories

Zefan-Cai/R-KV

A NeurIPS 2025 paper presenting redundancy-aware KV cache compression to serve reasoning models with reduced memory while preserving full accuracy.

R-KV
Not currently ranked — collecting fresh signals.
star history

R-KV is a technique that compresses key-value cache entries by discarding repetitive tokens during LLM decoding, targeting memory reduction for reasoning models. It integrates with popular inference frameworks including vLLM, SGLang, and flash attention. The method focuses on math and reasoning benchmarks such as AIME24, supporting models like DeepSeek-R1-Distill-Llama-8B.

Frequently asked

What is Zefan-Cai/R-KV?
A NeurIPS 2025 paper presenting redundancy-aware KV cache compression to serve reasoning models with reduced memory while preserving full accuracy.
Is R-KV open source?
Yes — Zefan-Cai/R-KV is an open-source project tracked on heatdrop.
What language is R-KV written in?
Zefan-Cai/R-KV is primarily written in Python.
How popular is R-KV?
Zefan-Cai/R-KV has 1.2k stars on GitHub.
Where can I find R-KV?
Zefan-Cai/R-KV is on GitHub at https://github.com/Zefan-Cai/R-KV.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.