← all repositories
SeekingDream/Static-to-Dynamic-LLMEval

Benchmarks rot when models train on the test set

A curated, actively maintained survey tracking how LLM benchmarking is evolving from fixed datasets to dynamic, contamination-resistant evaluation methods.

Static-to-Dynamic-LLMEval
Velocity · 7d
+0.0
★ / day
star history

What it does

This repository hosts the living companion to a survey on data contamination in LLM evaluation. It catalogs the field’s migration from static benchmarks—fixed, human-curated datasets that inevitably leak into training corpora—toward dynamic benchmarks that generate fresh tasks on demand. The authors also propose design principles for these dynamic evaluations, noting that the community still lacks standardized criteria to judge them.

The interesting bit

Instead of treating publication as the finish line, the authors openly state the EMNLP camera-ready is frozen while this repo stays current, and they actively solicit pull requests for missed papers or corrected venues. It is essentially a collaboratively edited literature index disguised as a GitHub repository.

Key highlights

  • Extensive taxonomy of static benchmarks spanning math, coding, reasoning, safety, language, and reading comprehension, plus mitigation tactics like canary strings, encryption, and post-hoc detection.
  • Breakdown of dynamic methods into temporal cutoffs, rule-based generation (templates, tables, graphs), LLM-based generation (rewriting, interactive, multi-agent), and hybrid approaches.
  • Explicit proposal of optimal design principles for dynamic benchmarks to fill the gap left by missing standardized evaluation criteria.
  • Community maintenance workflow: authors invite emails and pull requests to keep the paper list and taxonomy updated as the field evolves.

Caveats

  • This is a survey companion, not a framework; you will find paper citations and taxonomies rather than installable code or unified benchmarking scripts.
  • The README is a single, dense scroll of bibliographic entries; navigating it means relying on markdown anchor links.

Verdict

Bookmark this if you are building benchmarks or writing a literature review on LLM evaluation. Look elsewhere if you need a ready-to-run dynamic evaluation toolkit.

Frequently asked

What is SeekingDream/Static-to-Dynamic-LLMEval?
A curated, actively maintained survey tracking how LLM benchmarking is evolving from fixed datasets to dynamic, contamination-resistant evaluation methods.
Is Static-to-Dynamic-LLMEval open source?
Yes — SeekingDream/Static-to-Dynamic-LLMEval is an open-source project tracked on heatdrop.
How popular is Static-to-Dynamic-LLMEval?
SeekingDream/Static-to-Dynamic-LLMEval has 498 stars on GitHub.
Where can I find Static-to-Dynamic-LLMEval?
SeekingDream/Static-to-Dynamic-LLMEval is on GitHub at https://github.com/SeekingDream/Static-to-Dynamic-LLMEval.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.