HIT-SCIR/Chinese-Mixtral-8x7B
A Chinese-extended vocabulary large language model based on Mixtral-8x7B with vocabulary expansion and incremental pretraining.

Not currently ranked — collecting fresh signals.
star history
This project extends the Mixtral-8x7B model with a Chinese-optimized vocabulary to improve encoding and decoding efficiency for Chinese text. It provides incremental pretraining code on large-scale open-source corpora and releases both LoRA adapter weights and fully merged model weights for download. The extended vocabulary significantly boosts the model’s Chinese language generation and comprehension capabilities.
Frequently asked
- What is HIT-SCIR/Chinese-Mixtral-8x7B?
- A Chinese-extended vocabulary large language model based on Mixtral-8x7B with vocabulary expansion and incremental pretraining.
- Is Chinese-Mixtral-8x7B open source?
- Yes — HIT-SCIR/Chinese-Mixtral-8x7B is open source, released under the Apache-2.0 license.
- What language is Chinese-Mixtral-8x7B written in?
- HIT-SCIR/Chinese-Mixtral-8x7B is primarily written in Python.
- How popular is Chinese-Mixtral-8x7B?
- HIT-SCIR/Chinese-Mixtral-8x7B has 651 stars on GitHub.
- Where can I find Chinese-Mixtral-8x7B?
- HIT-SCIR/Chinese-Mixtral-8x7B is on GitHub at https://github.com/HIT-SCIR/Chinese-Mixtral-8x7B.