CLUEbenchmark/SuperCLUE
A benchmark for evaluating Chinese foundation models and large language models across multiple capability dimensions including agent performance.

SuperCLUE is a comprehensive benchmark for evaluating Chinese foundation models and LLMs. It assesses models across four primary capability quadrants: language understanding and generation, professional skills and knowledge, AI agents, and safety. The benchmark includes specific sub-evaluations such as SuperCLUE-Agent for agent task performance and SuperCLUE-Safety for adversarial safety testing. It provides monthly leaderboards and annual reports tracking the progress of Chinese AI models.
Frequently asked
- What is CLUEbenchmark/SuperCLUE?
- A benchmark for evaluating Chinese foundation models and large language models across multiple capability dimensions including agent performance.
- Is SuperCLUE open source?
- Yes — CLUEbenchmark/SuperCLUE is an open-source project tracked on heatdrop.
- How popular is SuperCLUE?
- CLUEbenchmark/SuperCLUE has 3.3k stars on GitHub.
- Where can I find SuperCLUE?
- CLUEbenchmark/SuperCLUE is on GitHub at https://github.com/CLUEbenchmark/SuperCLUE.