← all repositories
SemiAnalysisAI/InferenceX

The benchmark that chases daily kernel commits

InferenceX continuously re-benchmarks vLLM, SGLang, and TensorRT-LLM so static numbers don't go stale before your next deploy.

InferenceX
Not currently ranked — collecting fresh signals.
star history

What it does

InferenceX is an automated, continuously-running benchmark platform that tracks real-world LLM inference performance across major open-source frameworks and hardware stacks. It publishes live results to a public dashboard, aiming to reflect software improvements as they land rather than capturing a single frozen moment.

The interesting bit

The project treats “speed is the moat” as infrastructure: because kernel optimizations and scheduling improvements ship every few days, a benchmark from last quarter is already historical fiction. InferenceX attempts to close that gap by re-running against the latest vLLM, SGLang, TensorRT-LLM, CUDA, and ROCm builds. The backing hardware is notably serious—GB200 NVL72 racks, B200s, MI355X/CDNA3 GPUs—with TPUv6e/v7, Trainium2/3, and GB300 NVL72 listed as coming soon.

Key highlights

  • Live public dashboard at inferencex.com with open-sourced frontend
  • Hardware coverage spans NVIDIA Blackwell, AMD CDNA3, and planned Google/Amazon silicon
  • Endorsements from OpenAI Stargate infrastructure, Tri Dao, and vLLM project leads
  • Apache 2.0 licensed; SemiAnalysisAI/InferenceX repo claims exclusive “official” result status
  • Explicitly warns that forks and unofficial replicas may use different machine configs, producing subpar or misleading numbers

Caveats

  • The README is heavy on vision and light on methodology: no detail on workload shapes, prompt distributions, or how “continuous” the cadence actually is
  • “Soon™” hardware list (TPUv6e/v7, Trainium2/3, GB300 NVL72) is aspirational; no timeline given
  • The “official results only in this repo” disclaimer, while understandable for quality control, creates a centralized gatekeeper for what is otherwise pitched as open infrastructure

Verdict

Worth watching if you buy inference hardware or tune serving stacks and need defensible, current numbers rather than vendor datasheets. Less useful if you want to audit the benchmark methodology yourself—the project trusts its own process more than it explains it.

Frequently asked

What is SemiAnalysisAI/InferenceX?
InferenceX continuously re-benchmarks vLLM, SGLang, and TensorRT-LLM so static numbers don't go stale before your next deploy.
Is InferenceX open source?
Yes — SemiAnalysisAI/InferenceX is open source, released under the Apache-2.0 license.
What language is InferenceX written in?
SemiAnalysisAI/InferenceX is primarily written in Python.
How popular is InferenceX?
SemiAnalysisAI/InferenceX has 1.3k stars on GitHub.
Where can I find InferenceX?
SemiAnalysisAI/InferenceX is on GitHub at https://github.com/SemiAnalysisAI/InferenceX.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.