█

AI/TLDR

InferenceX

SemiAnalysis's open harness for continuously benchmarking LLM inference across GPUs and serving engines

Benchmark HarnessesOpen source
Language
Python
License
Apache-2.0
Coverage
1 story
$git clone https://github.com/SemiAnalysisAI/InferenceX.git

Overview

InferenceX, formerly called InferenceMAX, is SemiAnalysis's open-source inference performance research platform. Its premise is that LLM serving speed depends on two things moving at different rates: hardware improves in yearly steps as new GPUs and systems ship, while serving software such as SGLang, vLLM, TensorRT-LLM, CUDA and ROCm improves with every release through kernel optimisations, distributed inference strategies and scheduling changes. A benchmark taken once goes stale within days, so InferenceX re-runs its benchmarks automatically as those software stacks change and publishes the results to a free public dashboard at inferencex.com.

InferenceX chart of token throughput per GPU against per-user interactivity for DeepSeek V4 Pro in FP4 at 8K input and 1K output, with one curve per hardware and serving-engine pairing across GB300 NVL72, GB200 NVL72, B300, B200 and MI355X.
A typical InferenceX result: throughput per GPU against interactivity, one curve per hardware and serving-engine pairing.InferenceX README ↗

The repository is the harness behind that dashboard. Benchmarks are declared in master YAML configs, one entry per model, container image, precision, serving framework and runner, each with a search space of tensor, pipeline and expert parallelism and a range of concurrencies. A Python package (`infx`) validates those configs and expands them into a job matrix, GitHub Actions workflows fan the matrix out to self-hosted GPU runners, launcher scripts start the serving engine and drive the client workload, and the resulting JSON artifacts are collected and handed to the separate, also open-source InferenceX-app for ingestion and display.

Alongside the end-to-end serving benchmarks (InferenceX-e2e), the repository holds two experimental beta projects. CollectiveX measures dispatch, combine and roundtrip latency of mixture-of-experts expert-parallel communication libraries across accelerator systems, and OperatorX times one operator at a time (GEMM, attention, MoE, collectives and more) on NVIDIA, AMD, TPU and Trainium machines, writing one JSON file per run. The project states that only the SemiAnalysisAI/InferenceX repository produces official InferenceX results and that runs from forks must be labelled unofficial.

What it does

  • Declarative master configs describe each benchmark's model, container image, precision, framework, runner and parallelism/concurrency search space, validated with Pydantic before anything runs
  • Covers vLLM, SGLang and TensorRT-LLM, including NVIDIA Dynamo disaggregated variants, plus ATOM on AMD, across NVIDIA, AMD and TPU hardware such as GB300 NVL72, B200, MI355X, H100 and TPUv7x Ironwood
  • Fixed input/output sequence-length scenarios plus agentic-coding trace replay (AgentX) for long-context, multi-turn workloads
  • GitHub Actions pipeline that expands configs into a job matrix, runs them on self-hosted runners and collects benchmark, eval and trace artifacts
  • CollectiveX (beta) benchmarks MoE expert-parallel dispatch and combine latency across EP libraries such as DeepEP, MoRI, UCCL-EP and NCCL EP
  • OperatorX (beta) times individual inference operators per backend and emits one JSON result per run, submittable as SLURM jobs or run directly on TPU and Trainium VMs

Getting started

The end-to-end tooling lives in the inferencex-e2e/ project and is managed with uv. Validating configs and generating matrices runs on any machine; actually executing a benchmark needs GPU runners, which the project drives through its GitHub Actions workflows.

Clone and install the tooling

The Python manifest and lockfile belong to inferencex-e2e/, which also pins its own Python version.

bashbash
git clone https://github.com/SemiAnalysisAI/InferenceX.git
cd InferenceX/inferencex-e2e
uv sync --locked

Generate the job matrix for a config key

Each top-level key in configs/nvidia-master.yaml or configs/amd-master.yaml is one benchmark definition, named for its model, precision, hardware and framework. The generator validates it and prints the concrete jobs it expands into.

bashbash
uv run --locked python -m infx.matrix.generate test-config \
  --config-files configs/nvidia-master.yaml --config-keys dsr1-fp4-b200-sglang

Run benchmarks on GPU runners

Full runs go through the repository's GitHub Actions workflows: run-sweep.yml is the PR and push gate that fans a matrix out to self-hosted runners, and e2e-tests.yml is the manually dispatched end-to-end path. configs/runners.yaml maps runner labels to concrete machines, and the runners/ launchers handle model paths, containers and Slurm allocation. The docs index at inferencex-e2e/docs/index.md routes to the configuration, CI and troubleshooting guides.

Time individual operators with OperatorX

On a TPU VM or Trainium instance (no SLURM), run the operator suite directly and pick the cluster id; on a SLURM cluster, scripts/submit_run.py <platform> submits one job per container image and world size instead.

bashbash
cd InferenceX/operatorx
OPERATORX_CLUSTER=v6e_4x python -m operatorx

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Compare how the same model performs across GPU generations and serving engines using a reproducible, published methodology
  • Track how serving throughput and interactivity change as vLLM, SGLang or TensorRT-LLM releases land, instead of relying on a one-off benchmark
  • Contribute a tuned recipe for a new model, framework or hardware SKU and have it benchmarked alongside the others
  • Measure MoE expert-parallel communication or single-operator performance on your own cluster with CollectiveX or OperatorX

How InferenceX compares

InferenceX alongside other open-source benchmark harnesses tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
LM Evaluation Harness★ 14.2kEleutherAI's framework for few-shot evaluation of language models across 60+ academic benchmarks, used as the backend for many leaderboards.
OpenCompass★ 7.5kAn LLM evaluation platform that runs models against 100+ datasets covering reasoning, knowledge, coding, and domain tasks, with leaderboards and multi-model support.
SWE-bench★ 6kA benchmark and containerized harness that tests whether language models can resolve real GitHub issues by generating patches that pass a repository's tests.
simple-evals★ 4.7kOpenAI's lightweight library for running standard zero-shot, chain-of-thought benchmarks like MMLU, MATH, and GPQA to measure model accuracy.
lmms-eval★ 4.4kAn evaluation suite for large multimodal models that runs image, video, and audio benchmarks across many tasks with a unified, reproducible interface.
AgentBench★ 3.8kA benchmark that evaluates LLMs as agents across diverse interactive environments such as operating systems, databases, web browsing, and games.
EvalScope★ 3.5kModelScope evaluation framework that runs standard benchmarks against local or OpenAI-compatible models, with agent-loop evaluation, inference stress testing and a comparison dashboard.
InferenceX★ 1.8kSemiAnalysis's open harness for continuously benchmarking LLM inference across GPUs and serving engines