RLHF-V/RLAIF-V
Open-source framework for aligning multimodal large language models using AI feedback, achieving GPT-4V-level trustworthiness.

Not currently ranked — collecting fresh signals.
star history
RLAIF-V introduces a novel paradigm for training and aligning multimodal large language models using open-source AI feedback. The project provides a full pipeline including high-quality feedback data, online feedback learning algorithms, and pre-trained model weights (7B and 12B variants). The resulting models and training data are used by projects like MiniCPM-Llora3-V 2.5 for building competitive vision-language models.
Frequently asked
- What is RLHF-V/RLAIF-V?
- Open-source framework for aligning multimodal large language models using AI feedback, achieving GPT-4V-level trustworthiness.
- Is RLAIF-V open source?
- Yes — RLHF-V/RLAIF-V is an open-source project tracked on heatdrop.
- What language is RLAIF-V written in?
- RLHF-V/RLAIF-V is primarily written in Python.
- How popular is RLAIF-V?
- RLHF-V/RLAIF-V has 457 stars on GitHub.
- Where can I find RLAIF-V?
- RLHF-V/RLAIF-V is on GitHub at https://github.com/RLHF-V/RLAIF-V.