lxe/llavavision
A web accessibility tool that uses a local LLaVA vision model to narrate visual content for users.

Not currently ranked — collecting fresh signals.
star history
LLaVaVision is a browser-based application inspired by Be My Eyes that captures video frames, sends them to a locally-running multimodal LLaVA model via llama.cpp, and uses the Web Speech API to narrate the model’s descriptions of what it sees. The app uses the BakLLaVA-1 model quantized to q4_k for reasonable performance on consumer hardware with roughly 5GB RAM.
Frequently asked
- What is lxe/llavavision?
- A web accessibility tool that uses a local LLaVA vision model to narrate visual content for users.
- Is llavavision open source?
- Yes — lxe/llavavision is an open-source project tracked on heatdrop.
- What language is llavavision written in?
- lxe/llavavision is primarily written in JavaScript.
- How popular is llavavision?
- lxe/llavavision has 495 stars on GitHub.
- Where can I find llavavision?
- lxe/llavavision is on GitHub at https://github.com/lxe/llavavision.