deepseek-ai/DeepSeek-MoE
A Mixture-of-Experts language model architecture achieving expert specialization in large language models.

Not currently ranked — collecting fresh signals.
star history
DeepSeek-MoE implements a novel Mixture-of-Experts architecture for language models, focusing on achieving ultimate expert specialization. The architecture uses sparse activation to efficiently route tokens through specialized expert networks. It provides model weights, training code, and evaluation results as a foundation model research project.
Frequently asked
- What is deepseek-ai/DeepSeek-MoE?
- A Mixture-of-Experts language model architecture achieving expert specialization in large language models.
- Is DeepSeek-MoE open source?
- Yes — deepseek-ai/DeepSeek-MoE is open source, released under the MIT license.
- What language is DeepSeek-MoE written in?
- deepseek-ai/DeepSeek-MoE is primarily written in Python.
- How popular is DeepSeek-MoE?
- deepseek-ai/DeepSeek-MoE has 1.9k stars on GitHub.
- Where can I find DeepSeek-MoE?
- deepseek-ai/DeepSeek-MoE is on GitHub at https://github.com/deepseek-ai/DeepSeek-MoE.