Overview
Lance is an open lakehouse format for multimodal AI. It is not a database you run but a file format, table format and catalog spec you build on top of object storage, so images, video, audio, text and embeddings all live in one dataset that your training and search code reads directly.
Its defining trade-off is random access. Columnar formats like Parquet are built for full scans, which makes fetching individual rows — exactly what shuffled ML training and point lookups do — expensive. The project reports Lance is 100x faster than Parquet for random access without giving up scan performance, which is what lets the same copy of the data serve both analytics and training.
Around that core sit the features a retrieval stack usually bolts on separately: accelerated secondary indices for vector similarity and BM25 full-text search alongside SQL analytics, blob encoding with lazy loading for large media, data evolution that adds backfilled columns without rewriting the table, and automatic zero-copy versioning with ACID transactions, time travel, tags and branches. It reads through Apache Arrow, Pandas, Polars, DuckDB, Spark, Ray and Trino, and is the storage layer beneath LanceDB.
What it does
- Hybrid search on one dataset: vector similarity, BM25 full-text and SQL analytics backed by accelerated secondary indices
- Random access the project benchmarks at 100x faster than Parquet or Iceberg, without sacrificing scan throughput
- Native multimodal storage for images, video, audio, text and embeddings with blob encoding and lazy loading
- Data evolution: add columns with backfilled values without a full table rewrite
- Zero-copy versioning with ACID transactions, time travel, tags and branches — no extra infrastructure
- Reads from Apache Arrow, Pandas, Polars, DuckDB, Spark, Ray, Trino and open catalogs such as Apache Polaris and Unity Catalog
Getting started
Lance ships as the `pylance` Python package. Converting an existing Parquet dataset is the fastest way to see the format, and the result is readable as a normal PyArrow dataset.
Install pylance
The Python bindings to the Rust core.
pip install pylanceConvert an existing dataset to Lance
`lance.write_dataset` accepts a PyArrow dataset, so any Parquet table converts in a couple of lines.
import lance
import pandas as pd
import pyarrow as pa
import pyarrow.dataset
df = pd.DataFrame({"a": [5], "b": [10]})
uri = "/tmp/test.parquet"
tbl = pa.Table.from_pandas(df)
pa.dataset.write_dataset(tbl, uri, format='parquet')
parquet = pa.dataset.dataset(uri, format='parquet')
lance.write_dataset(parquet, "/tmp/test.lance")Read it back
A Lance dataset is a PyArrow dataset, so the rest of your stack does not need to change.
dataset = lance.dataset("/tmp/test.lance")
assert isinstance(dataset, pa.dataset.Dataset)
df = dataset.to_table().to_pandas()Query it from DuckDB
DuckDB v0.7+ reads the dataset in place, with no export step.
import duckdb
duckdb.query("SELECT * FROM dataset LIMIT 10").to_df()Pin a storage version for production
Lance releases often, but the file format is versioned separately by the `data_storage_version` stored in each dataset. Write production data with a stable version — the `next` alias is unstable and meant only for experimentation. See the format versioning guide for the compatibility matrix.
Commands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Store a training corpus of images or video alongside its embeddings so shuffled random-access reads stay fast
- Build a search engine or feature store that needs vector similarity, full-text and SQL over one copy of the data
- Add engineered feature columns to an existing table without rewriting terabytes of data
- Use time travel and branches to reproduce exactly which rows a model was trained on
How Lance compares
Lance alongside other open-source data wrangling tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| Hugging Face Datasets | ★ 22k | An Apache Arrow-backed Python library that loads datasets from the Hugging Face Hub or local files in one line and maps, filters and streams them without holding them in RAM. |
| Rerun | ★ 11.5k | A data layer for physical AI that logs, visualises and queries multi-rate multimodal streams — images, point clouds, transforms, joint states, video — and streams them straight into training. |
| Datasette | ★ 11.5k | An open source multi-tool that points at a SQLite file and serves it as a browsable website with a JSON API, plus commands for publishing the result online. |
| ChatLab | ★ 7.5k | A local-first desktop app that imports chat exports from eight messaging platforms into one normalized store, then queries them with SQL and tool-calling AI agents. |
| Lance | ★ 7.1k | An open lakehouse format built for multimodal AI data |
| Daft | ★ 5.8k | High-performance data engine for AI: process images, audio, video and embeddings alongside structured data in one Python dataframe, with a Rust core that scales from a laptop to a Ray or Kubernetes cluster. |
| sqlite-utils | ★ 2.2k | A Python CLI and library that turns JSON, CSV and TSV into SQLite databases, creating schemas automatically and adding full-text search, table transforms and migrations. |