← all repositories

dleemiller/WordLlama

A lightweight NLP toolkit for similarity, ranking, deduplication, and clustering using LLM token embeddings.

1.5k stars Python RAG · SearchData Tooling
WordLlama
Not currently ranked — collecting fresh signals.
star history

WordLlama is a fast, lightweight NLP toolkit that operates on token embeddings from LLMs. It provides CPU-optimized functionality for fuzzy deduplication, semantic similarity computation, document ranking, clustering, and text splitting. The toolkit supports model2vec static embeddings and achieves competitive results on the MTEB benchmark.

Frequently asked

What is dleemiller/WordLlama?
A lightweight NLP toolkit for similarity, ranking, deduplication, and clustering using LLM token embeddings.
Is WordLlama open source?
Yes — dleemiller/WordLlama is open source, released under the MIT license.
What language is WordLlama written in?
dleemiller/WordLlama is primarily written in Python.
How popular is WordLlama?
dleemiller/WordLlama has 1.5k stars on GitHub.
Where can I find WordLlama?
dleemiller/WordLlama is on GitHub at https://github.com/dleemiller/WordLlama.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.