Beomi/KcBERT
Pretrained BERT model and WordPiece tokenizer trained on Korean comments, with associated datasets for fine-tuning Korean NLP tasks.
★492 stars Language Models

Not currently ranked — collecting fresh signals.
star history
KcBERT provides a Korean-specific pretrained BERT model and WordPiece tokenizer trained on Korean comment data. The repository includes released datasets (v2022.3Q with 45GB and 340M records) and supports fine-tuning for downstream Korean NLP tasks via Colab notebooks. The project also led to the development of KcELECTRA, a newer model with improved performance across most tasks.
Frequently asked
- What is Beomi/KcBERT?
- Pretrained BERT model and WordPiece tokenizer trained on Korean comments, with associated datasets for fine-tuning Korean NLP tasks.
- Is KcBERT open source?
- Yes — Beomi/KcBERT is open source, released under the MIT license.
- How popular is KcBERT?
- Beomi/KcBERT has 492 stars on GitHub.
- Where can I find KcBERT?
- Beomi/KcBERT is on GitHub at https://github.com/Beomi/KcBERT.