Delete the one load-bearing fact, see who notices
HealthBench grades a reply against a rubric written for the exact message the physician saw; Keystone edits the decisive fact underneath it and checks whether the action — and the rubric — still hold.
What it does Every HealthBench source conversation gets a decision frame: what is being decided, the action the stated evidence supports, and the facts it rests on. Keystone then builds twins that change exactly one thing — remove a decisive fact, add a credible contradiction, swap a demographic, or mention a red flag in passing. Each twin carries an annotated evidence state with the actions a clinician would accept or forbid, and a judge from a different vendor maps every reply onto it, so the primary outcomes are about actions rather than prose-matching.
The interesting bit It grades the grader. Delete the fact a rubric criterion depends on and the rubric may stop measuring anything — Keystone scores that applicability alongside the model, which is rare honesty in benchmark design. The name is the method: pull the keystone and see who notices the arch is gone.
Key highlights
- 7,318 twins over 1,236 HealthBench conversations across eight edit families; 2,815 are multi-turn with byte-identical earlier turns, so only the decisive edit differs.
- Two of the eight families are negative controls where correct behaviour is no change at all — a trap for models that fidget when any sentence moves.
- Paraphrase-only controls for 1,119 sources (reworded, nothing removed) separate models reacting to evidence from models reacting to wording.
- Reference results for five assistants ship with every reply and verdict; on the quick set, asking the decisive question after a fact is removed ranges from 0.37 to 0.74 across models.
- The twins aren’t distributed as prose — the release rebuilds them locally from OpenAI’s public HealthBench copy and verifies every file against published SHA-256 hashes; 81 tests run without an API key.
Caveats
- Labels are tier
silver: model-drafted and reviewed by a second vendor, with clinician-confirmed rows promoted togold— the accept/forbid annotations aren’t all physician-written yet. - Scoring is LLM-judged (GPT-4.1 in the reference run), validated at 0.90–0.99 on 166 authored items with intended labels; solid, but still a model grading models.
- Pooled numbers mislead — the two negative-control families are 51% of the
primarylayer, so the README itself insists you report per family.
Verdict Worth a close look if you evaluate clinical assistants or benchmark design — the decision-frame-plus-twins method ports to any domain where one fact should flip an action. Skip it if you want a single clinical-quality score; this is a narrow, deliberate instrument that will tell on your rubric as readily as on your model.
Frequently asked
- What is xinxuxin/keystone-bench?
- HealthBench grades a reply against a rubric written for the exact message the physician saw; Keystone edits the decisive fact underneath it and checks whether the action — and the rubric — still hold.
- Is keystone-bench open source?
- Yes — xinxuxin/keystone-bench is an open-source project tracked on heatdrop.
- What language is keystone-bench written in?
- xinxuxin/keystone-bench is primarily written in Python.
- How popular is keystone-bench?
- xinxuxin/keystone-bench has 510 stars on GitHub.
- Where can I find keystone-bench?
- xinxuxin/keystone-bench is on GitHub at https://github.com/xinxuxin/keystone-bench.