NVIDIA/NeMo-Retriever
NVIDIA's document extraction and embedding pipeline for retrieval-augmented generation and generative AI applications.

Not currently ranked — collecting fresh signals.
star history
NeMo Retriever Library extracts text, tables, charts, and infographics from documents using OCR and classification, then computes vector embeddings for the extracted content and stores them in LanceDB for downstream generative AI and RAG applications. It leverages NVIDIA NIM microservices for scalable, production-grade document processing and can be deployed on Kubernetes using Helm charts.
Frequently asked
- What is NVIDIA/NeMo-Retriever?
- NVIDIA's document extraction and embedding pipeline for retrieval-augmented generation and generative AI applications.
- Is NeMo-Retriever open source?
- Yes — NVIDIA/NeMo-Retriever is open source, released under the Apache-2.0 license.
- What language is NeMo-Retriever written in?
- NVIDIA/NeMo-Retriever is primarily written in Python.
- How popular is NeMo-Retriever?
- NVIDIA/NeMo-Retriever has 3k stars on GitHub.
- Where can I find NeMo-Retriever?
- NVIDIA/NeMo-Retriever is on GitHub at https://github.com/NVIDIA/NeMo-Retriever.