← all repositories
xinxuxin/keystone-bench

Delete the one load-bearing fact, see who notices

HealthBench grades a reply against a rubric written for the exact message the physician saw; Keystone edits the decisive fact underneath it and checks whether the action — and the rubric — still hold.

★510 stars Python LLMOps · EvalDomain Apps
keystone-bench
Collecting fresh signals — velocity needs a few days of history.
collecting data…
star history

What it does Every HealthBench source conversation gets a decision frame: what is being decided, the action the stated evidence supports, and the facts it rests on. Keystone then builds twins that change exactly one thing — remove a decisive fact, add a credible contradiction, swap a demographic, or mention a red flag in passing. Each twin carries an annotated evidence state with the actions a clinician would accept or forbid, and a judge from a different vendor maps every reply onto it, so the primary outcomes are about actions rather than prose-matching.

The interesting bit It grades the grader. Delete the fact a rubric criterion depends on and the rubric may stop measuring anything — Keystone scores that applicability alongside the model, which is rare honesty in benchmark design. The name is the method: pull the keystone and see who notices the arch is gone.

Key highlights

  • 7,318 twins over 1,236 HealthBench conversations across eight edit families; 2,815 are multi-turn with byte-identical earlier turns, so only the decisive edit differs.
  • Two of the eight families are negative controls where correct behaviour is no change at all — a trap for models that fidget when any sentence moves.
  • Paraphrase-only controls for 1,119 sources (reworded, nothing removed) separate models reacting to evidence from models reacting to wording.
  • Reference results for five assistants ship with every reply and verdict; on the quick set, asking the decisive question after a fact is removed ranges from 0.37 to 0.74 across models.
  • The twins aren’t distributed as prose — the release rebuilds them locally from OpenAI’s public HealthBench copy and verifies every file against published SHA-256 hashes; 81 tests run without an API key.

Caveats

  • Labels are tier silver: model-drafted and reviewed by a second vendor, with clinician-confirmed rows promoted to gold — the accept/forbid annotations aren’t all physician-written yet.
  • Scoring is LLM-judged (GPT-4.1 in the reference run), validated at 0.90–0.99 on 166 authored items with intended labels; solid, but still a model grading models.
  • Pooled numbers mislead — the two negative-control families are 51% of the primary layer, so the README itself insists you report per family.

Verdict Worth a close look if you evaluate clinical assistants or benchmark design — the decision-frame-plus-twins method ports to any domain where one fact should flip an action. Skip it if you want a single clinical-quality score; this is a narrow, deliberate instrument that will tell on your rubric as readily as on your model.

Frequently asked

What is xinxuxin/keystone-bench?
HealthBench grades a reply against a rubric written for the exact message the physician saw; Keystone edits the decisive fact underneath it and checks whether the action — and the rubric — still hold.
Is keystone-bench open source?
Yes — xinxuxin/keystone-bench is an open-source project tracked on heatdrop.
What language is keystone-bench written in?
xinxuxin/keystone-bench is primarily written in Python.
How popular is keystone-bench?
xinxuxin/keystone-bench has 510 stars on GitHub.
Where can I find keystone-bench?
xinxuxin/keystone-bench is on GitHub at https://github.com/xinxuxin/keystone-bench.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.