← all repositories

RLHF-V/RLAIF-V

Open-source framework for aligning multimodal large language models using AI feedback, achieving GPT-4V-level trustworthiness.

RLAIF-V
Not currently ranked — collecting fresh signals.
star history

RLAIF-V introduces a novel paradigm for training and aligning multimodal large language models using open-source AI feedback. The project provides a full pipeline including high-quality feedback data, online feedback learning algorithms, and pre-trained model weights (7B and 12B variants). The resulting models and training data are used by projects like MiniCPM-Llora3-V 2.5 for building competitive vision-language models.

Frequently asked

What is RLHF-V/RLAIF-V?
Open-source framework for aligning multimodal large language models using AI feedback, achieving GPT-4V-level trustworthiness.
Is RLAIF-V open source?
Yes — RLHF-V/RLAIF-V is an open-source project tracked on heatdrop.
What language is RLAIF-V written in?
RLHF-V/RLAIF-V is primarily written in Python.
How popular is RLAIF-V?
RLHF-V/RLAIF-V has 457 stars on GitHub.
Where can I find RLAIF-V?
RLHF-V/RLAIF-V is on GitHub at https://github.com/RLHF-V/RLAIF-V.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.