← all repositories

Beomi/KcBERT

Pretrained BERT model and WordPiece tokenizer trained on Korean comments, with associated datasets for fine-tuning Korean NLP tasks.

492 stars Language Models
KcBERT
Not currently ranked — collecting fresh signals.
star history

KcBERT provides a Korean-specific pretrained BERT model and WordPiece tokenizer trained on Korean comment data. The repository includes released datasets (v2022.3Q with 45GB and 340M records) and supports fine-tuning for downstream Korean NLP tasks via Colab notebooks. The project also led to the development of KcELECTRA, a newer model with improved performance across most tasks.

Frequently asked

What is Beomi/KcBERT?
Pretrained BERT model and WordPiece tokenizer trained on Korean comments, with associated datasets for fine-tuning Korean NLP tasks.
Is KcBERT open source?
Yes — Beomi/KcBERT is open source, released under the MIT license.
How popular is KcBERT?
Beomi/KcBERT has 492 stars on GitHub.
Where can I find KcBERT?
Beomi/KcBERT is on GitHub at https://github.com/Beomi/KcBERT.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.