← all repositories
pariskang/CMLM-ZhongJing

135k instructions to teach LLMs tongue diagnosis and herbal formulas

It exists because off-the-shelf LLMs confuse herbal formulas with ibuprofen and botch tongue diagnosis.

512 stars Jupyter Notebook Language ModelsDomain Apps
CMLM-ZhongJing
Not currently ranked — collecting fresh signals.
star history

What it does

CMLM-ZhongJing is a family of fine-tuned models—built on Baichuan2-13B-Chat and Qwen1.5-1.8B—that specializes in Traditional Chinese Medicine dialogue. The team constructed roughly 135,000 instruction pairs drawn from classical medical texts, symptom synonyms, and fifteen distinct clinical scenarios ranging from tongue-and-pulse reading to prescription dosage. A 1.8-billion-parameter variant is designed to run on a single Tesla T4, while a 13B version is available for users with more GPU memory.

The interesting bit

The authors treat medical hallucinations as a safety hazard, not a novelty, so they decompose TCM reasoning into fifteen granular tasks—such as diagnostic_analysis, herb_dosage, and narrative_medicine—to force structured clinical logic instead of surface-level pattern matching. In their own physician-led evaluations, they claim the resulting model outperforms GPT-4 on specific TCM prescription tasks, though the benchmarks are internal and the outputs carry an explicit “academic use only” disclaimer.

Key highlights

  • Fine-tuned weights for Baichuan2-13B-Chat and Qwen1.5-1.8B-Chat; not a ground-up pretrain.
  • Instruction corpus of ~135k entries spanning 15 clinical task types, from tongue_palse to critical_thinking_data.
  • 1.8B model targets single-GPU (Tesla T4) deployment; 13B variant for high-performance setups.
  • Paper accepted in Tsinghua Science & Technology.
  • Explicit disclaimer that all output is for academic research, not clinical diagnosis.

Caveats

  • The claim of surpassing GPT-4 rests on the project’s own small-scale physician tests; no independent benchmark or open evaluation suite is visible.
  • The README is almost entirely in Chinese, with only a partial English translation, so non-Chinese speakers will need translation tools to dig into dataset details.
  • Because these are fine-tunes atop existing chat models, inherited base-model biases and safety guardrails remain, rather than TCM-specific ones.

Verdict

Worth a look if you study medical NLP, low-resource domain adaptation, or TCM informatics. Skip it if you need a clinically validated diagnostic tool—the repository itself warns that it is not one.

Frequently asked

What is pariskang/CMLM-ZhongJing?
It exists because off-the-shelf LLMs confuse herbal formulas with ibuprofen and botch tongue diagnosis.
Is CMLM-ZhongJing open source?
Yes — pariskang/CMLM-ZhongJing is open source, released under the MIT license.
What language is CMLM-ZhongJing written in?
pariskang/CMLM-ZhongJing is primarily written in Jupyter Notebook.
How popular is CMLM-ZhongJing?
pariskang/CMLM-ZhongJing has 512 stars on GitHub.
Where can I find CMLM-ZhongJing?
pariskang/CMLM-ZhongJing is on GitHub at https://github.com/pariskang/CMLM-ZhongJing.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.