█

AI/TLDR

Rerun

Log, query and visualise multimodal robotics data, then stream it into training

Data WranglingOpen source
Language
Rust
License
MIT OR Apache-2.0
$pip install rerun-sdk

Overview

Rerun is a data layer for physical AI: you log multimodal, multi-rate data from your code, and a viewer renders all of it in sync while the same data stays queryable. It ingests images, point clouds, transforms, time series, tensors, joint states and video from robot logs, egocentric and UMI rigs, simulation and web video, reading formats including MCAP, its own rrd recordings and LeRobot datasets.

The viewer is the part people meet first — scrub through an episode, compare sensors side by side, watch a CV pipeline run live — but the storage underneath is the point. Rerun is built in Rust on column-chunk storage designed for multi-rate physical data, so the same recording can be queried with dataframes or SQL and streamed straight into training without an export step or a second stale copy of the dataset.

SDKs exist for Python, Rust and C++, and the viewer also runs in a browser. The Python package bundles the viewer binary; C++ and Rust users install the `rerun` CLI separately. The project is dual-licensed under MIT and Apache 2.0.

What it does

  • Log multi-rate multimodal data — images, point clouds, transforms, time series, tensors, joint states, video — from one API
  • Built-in viewer that renders every stream in sync and in realtime, locally or in the browser
  • Query the same recordings with dataframes or SQL, and stream dataset mixes directly into training
  • Reads robot logs, human-data rigs, simulation and web video; MCAP, rrd and LeRobot formats
  • SDKs for Python, Rust and C++, over column-chunk storage written in Rust
  • Stream over the network to a remote viewer, or save recordings to `.rrd` files on disk

Getting started

The Python path is the fastest — the SDK ships the viewer with it. All commands and code below come from the project README.

Install the Python SDK

This also installs the `rerun` viewer binary, so `rerun --help` works afterwards.

bashbash
pip install rerun-sdk

Log your first recording

Initialise a recording, spawn the viewer, put your data on a timeline and log it against an entity path.

pythonpython
import rerun as rr  # pip install rerun-sdk

rr.init("rerun_example_app")

rr.spawn()  # Spawn a child process with a viewer and connect
# rr.save("recording.rrd")  # Stream all logs to disk
# rr.connect_grpc()  # Connect to a remote viewer

# Associate subsequent data with 42 on the "frame" timeline
rr.set_time("frame", sequence=42)

# Log colored 3D points to the entity at `path/to/points`
rr.log("path/to/points", rr.Points3D(positions, colors=colors))

Or use the Rust or C++ SDK

Rust and C++ do not bundle the viewer, so install the CLI alongside the SDK.

bashbash
cargo add rerun
cargo install rerun-cli --locked --features nasm

Check the viewer binary is on your path

The `rerun` binary is what loads `.rrd` files and receives log data streamed over the network.

bashbash
rerun --help

Get data back out

The dataframe and SQL query APIs turn a recording into rows you can analyse or feed to a training job — see the how-to on querying and transforming.

texttext
https://rerun.io/docs/howto/query-and-transform/get-data-out

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Reach for it to debug a robotics or computer-vision pipeline by watching every sensor stream in sync
  • Reach for it when episodes from robots, rigs, simulation and video have to live in one queryable substrate
  • Reach for it to inspect intermediate and derived data from a CV pipeline, not just the final output
  • Reach for it when you want to stream dataset mixes into training without export jobs or stale copies

How Rerun compares

Rerun alongside other open-source data wrangling tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Hugging Face Datasets★ 22kAn Apache Arrow-backed Python library that loads datasets from the Hugging Face Hub or local files in one line and maps, filters and streams them without holding them in RAM.
Rerun★ 11.5kLog, query and visualise multimodal robotics data, then stream it into training
Datasette★ 11.5kAn open source multi-tool that points at a SQLite file and serves it as a browsable website with a JSON API, plus commands for publishing the result online.
ChatLab★ 7.5kA local-first desktop app that imports chat exports from eight messaging platforms into one normalized store, then queries them with SQL and tool-calling AI agents.
Lance★ 7.1kAn open lakehouse format for multimodal AI: one dataset holding images, video, audio, text and embeddings, with 100x faster random access than Parquet, vector and full-text indices, and zero-copy versioning.
Daft★ 5.8kHigh-performance data engine for AI: process images, audio, video and embeddings alongside structured data in one Python dataframe, with a Rust core that scales from a laptop to a Ray or Kubernetes cluster.
sqlite-utils★ 2.2kA Python CLI and library that turns JSON, CSV and TSV into SQLite databases, creating schemas automatically and adding full-text search, table transforms and migrations.