← all repositories

vorojar/Folio-OCR

Batch OCR workbench powered by local GLM-OCR and Ollama models for document digitization.

448 stars Python Computer VisionData Tooling
Folio-OCR
Not currently ranked — collecting fresh signals.
star history

Folio-OCR is a three-panel document OCR application that uses local GLM-OCR and Ollama to recognize text from PDFs and images in batch. It performs layout detection to automatically segment pages into text regions, merges adjacent areas to reduce API calls, and exports results as Markdown, Word, EPUB, or plain text. The tool runs entirely offline with SQLite for session persistence.

Frequently asked

What is vorojar/Folio-OCR?
Batch OCR workbench powered by local GLM-OCR and Ollama models for document digitization.
Is Folio-OCR open source?
Yes — vorojar/Folio-OCR is an open-source project tracked on heatdrop.
What language is Folio-OCR written in?
vorojar/Folio-OCR is primarily written in Python.
How popular is Folio-OCR?
vorojar/Folio-OCR has 448 stars on GitHub.
Where can I find Folio-OCR?
vorojar/Folio-OCR is on GitHub at https://github.com/vorojar/Folio-OCR.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.