imcaspar/gpt2-ml
A multilingual GPT-2 implementation providing 1.5B parameter Chinese pretrained models with TPU training support.

Not currently ranked — collecting fresh signals.
star history
This repository provides GPT-2 adapted for multiple languages, with a focus on Chinese. It includes 1.5 billion parameter pretrained Chinese language models trained on large corpora (~15G and ~30G text), along with training scripts supporting TPU execution based on Grover, and ported BERT tokenizers for multilingual compatibility. The project offers ready-to-use Colab demos for model inference.
Frequently asked
- What is imcaspar/gpt2-ml?
- A multilingual GPT-2 implementation providing 1.5B parameter Chinese pretrained models with TPU training support.
- Is gpt2-ml open source?
- Yes — imcaspar/gpt2-ml is open source, released under the Apache-2.0 license.
- What language is gpt2-ml written in?
- imcaspar/gpt2-ml is primarily written in Python.
- How popular is gpt2-ml?
- imcaspar/gpt2-ml has 1.7k stars on GitHub.
- Where can I find gpt2-ml?
- imcaspar/gpt2-ml is on GitHub at https://github.com/imcaspar/gpt2-ml.