Overview
InferenceX, formerly called InferenceMAX, is SemiAnalysis's open-source inference performance research platform. Its premise is that LLM serving speed depends on two things moving at different rates: hardware improves in yearly steps as new GPUs and systems ship, while serving software such as SGLang, vLLM, TensorRT-LLM, CUDA and ROCm improves with every release through kernel optimisations, distributed inference strategies and scheduling changes. A benchmark taken once goes stale within days, so InferenceX re-runs its benchmarks automatically as those software stacks change and publishes the results to a free public dashboard at inferencex.com.

The repository is the harness behind that dashboard. Benchmarks are declared in master YAML configs, one entry per model, container image, precision, serving framework and runner, each with a search space of tensor, pipeline and expert parallelism and a range of concurrencies. A Python package (`infx`) validates those configs and expands them into a job matrix, GitHub Actions workflows fan the matrix out to self-hosted GPU runners, launcher scripts start the serving engine and drive the client workload, and the resulting JSON artifacts are collected and handed to the separate, also open-source InferenceX-app for ingestion and display.
Alongside the end-to-end serving benchmarks (InferenceX-e2e), the repository holds two experimental beta projects. CollectiveX measures dispatch, combine and roundtrip latency of mixture-of-experts expert-parallel communication libraries across accelerator systems, and OperatorX times one operator at a time (GEMM, attention, MoE, collectives and more) on NVIDIA, AMD, TPU and Trainium machines, writing one JSON file per run. The project states that only the SemiAnalysisAI/InferenceX repository produces official InferenceX results and that runs from forks must be labelled unofficial.
What it does
- Declarative master configs describe each benchmark's model, container image, precision, framework, runner and parallelism/concurrency search space, validated with Pydantic before anything runs
- Covers vLLM, SGLang and TensorRT-LLM, including NVIDIA Dynamo disaggregated variants, plus ATOM on AMD, across NVIDIA, AMD and TPU hardware such as GB300 NVL72, B200, MI355X, H100 and TPUv7x Ironwood
- Fixed input/output sequence-length scenarios plus agentic-coding trace replay (AgentX) for long-context, multi-turn workloads
- GitHub Actions pipeline that expands configs into a job matrix, runs them on self-hosted runners and collects benchmark, eval and trace artifacts
- CollectiveX (beta) benchmarks MoE expert-parallel dispatch and combine latency across EP libraries such as DeepEP, MoRI, UCCL-EP and NCCL EP
- OperatorX (beta) times individual inference operators per backend and emits one JSON result per run, submittable as SLURM jobs or run directly on TPU and Trainium VMs
Getting started
The end-to-end tooling lives in the inferencex-e2e/ project and is managed with uv. Validating configs and generating matrices runs on any machine; actually executing a benchmark needs GPU runners, which the project drives through its GitHub Actions workflows.
Clone and install the tooling
The Python manifest and lockfile belong to inferencex-e2e/, which also pins its own Python version.
git clone https://github.com/SemiAnalysisAI/InferenceX.git
cd InferenceX/inferencex-e2e
uv sync --lockedGenerate the job matrix for a config key
Each top-level key in configs/nvidia-master.yaml or configs/amd-master.yaml is one benchmark definition, named for its model, precision, hardware and framework. The generator validates it and prints the concrete jobs it expands into.
uv run --locked python -m infx.matrix.generate test-config \
--config-files configs/nvidia-master.yaml --config-keys dsr1-fp4-b200-sglangRun benchmarks on GPU runners
Full runs go through the repository's GitHub Actions workflows: run-sweep.yml is the PR and push gate that fans a matrix out to self-hosted runners, and e2e-tests.yml is the manually dispatched end-to-end path. configs/runners.yaml maps runner labels to concrete machines, and the runners/ launchers handle model paths, containers and Slurm allocation. The docs index at inferencex-e2e/docs/index.md routes to the configuration, CI and troubleshooting guides.
Time individual operators with OperatorX
On a TPU VM or Trainium instance (no SLURM), run the operator suite directly and pick the cluster id; on a SLURM cluster, scripts/submit_run.py <platform> submits one job per container image and world size instead.
cd InferenceX/operatorx
OPERATORX_CLUSTER=v6e_4x python -m operatorxCommands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Compare how the same model performs across GPU generations and serving engines using a reproducible, published methodology
- Track how serving throughput and interactivity change as vLLM, SGLang or TensorRT-LLM releases land, instead of relying on a one-off benchmark
- Contribute a tuned recipe for a new model, framework or hardware SKU and have it benchmarked alongside the others
- Measure MoE expert-parallel communication or single-operator performance on your own cluster with CollectiveX or OperatorX
How InferenceX compares
InferenceX alongside other open-source benchmark harnesses tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| LM Evaluation Harness | ★ 14.2k | EleutherAI's framework for few-shot evaluation of language models across 60+ academic benchmarks, used as the backend for many leaderboards. |
| OpenCompass | ★ 7.5k | An LLM evaluation platform that runs models against 100+ datasets covering reasoning, knowledge, coding, and domain tasks, with leaderboards and multi-model support. |
| SWE-bench | ★ 6k | A benchmark and containerized harness that tests whether language models can resolve real GitHub issues by generating patches that pass a repository's tests. |
| simple-evals | ★ 4.7k | OpenAI's lightweight library for running standard zero-shot, chain-of-thought benchmarks like MMLU, MATH, and GPQA to measure model accuracy. |
| lmms-eval | ★ 4.4k | An evaluation suite for large multimodal models that runs image, video, and audio benchmarks across many tasks with a unified, reproducible interface. |
| AgentBench | ★ 3.8k | A benchmark that evaluates LLMs as agents across diverse interactive environments such as operating systems, databases, web browsing, and games. |
| EvalScope | ★ 3.5k | ModelScope evaluation framework that runs standard benchmarks against local or OpenAI-compatible models, with agent-loop evaluation, inference stress testing and a comparison dashboard. |
| InferenceX | ★ 1.8k | SemiAnalysis's open harness for continuously benchmarking LLM inference across GPUs and serving engines |