← all repositories

FMInference/H2O

A KV cache eviction policy for LLM inference that dynamically retains heavy-hitter tokens to improve throughput while maintaining accuracy.

H2O
Not currently ranked — collecting fresh signals.
star history

H2O is a research implementation from NeurIPS 2023 that optimizes LLM inference by dynamically managing the KV cache. It identifies that a small subset of tokens contribute most to attention scores and proposes an eviction policy that balances recent tokens with these high-value heavy-hitter tokens. The approach reduces memory footprint and improves throughput significantly while maintaining model accuracy across OPT, LLaMA, and GPT-NeoX architectures.

Frequently asked

What is FMInference/H2O?
A KV cache eviction policy for LLM inference that dynamically retains heavy-hitter tokens to improve throughput while maintaining accuracy.
Is H2O open source?
Yes — FMInference/H2O is an open-source project tracked on heatdrop.
What language is H2O written in?
FMInference/H2O is primarily written in Python.
How popular is H2O?
FMInference/H2O has 530 stars on GitHub.
Where can I find H2O?
FMInference/H2O is on GitHub at https://github.com/FMInference/H2O.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.