← all repositories

airaria/Visual-Chinese-LLaMA-Alpaca

A multimodal Chinese LLaMA model extended with visual encoding to process and understand image inputs alongside text.

Visual-Chinese-LLaMA-Alpaca
Not currently ranked — collecting fresh signals.
star history

VisualCLA extends the Chinese LLaMA/Alpaca foundation model with image encoding modules, enabling it to process visual information. It uses Chinese image-text pairs for multimodal pretraining to align visual and textual representations, followed by instruction tuning on multimodal datasets to improve instruction following and conversational abilities. The project provides inference code and deployment scripts via Gradio and Text-Generation-WebUI.

Frequently asked

What is airaria/Visual-Chinese-LLaMA-Alpaca?
A multimodal Chinese LLaMA model extended with visual encoding to process and understand image inputs alongside text.
Is Visual-Chinese-LLaMA-Alpaca open source?
Yes — airaria/Visual-Chinese-LLaMA-Alpaca is open source, released under the Apache-2.0 license.
What language is Visual-Chinese-LLaMA-Alpaca written in?
airaria/Visual-Chinese-LLaMA-Alpaca is primarily written in Python.
How popular is Visual-Chinese-LLaMA-Alpaca?
airaria/Visual-Chinese-LLaMA-Alpaca has 461 stars on GitHub.
Where can I find Visual-Chinese-LLaMA-Alpaca?
airaria/Visual-Chinese-LLaMA-Alpaca is on GitHub at https://github.com/airaria/Visual-Chinese-LLaMA-Alpaca.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.