█

AI/TLDR

Daft

A Python-native, Rust-powered data engine for multimodal AI workloads

Data WranglingOpen source
Language
Rust
License
Apache-2.0
$pip install daft

Overview

Daft is a high-performance data engine built for AI and multimodal workloads. Where most dataframe libraries treat an image or an audio clip as an opaque Python object, Daft handles those columns natively, so you can decode images, transcribe audio and score embeddings alongside ordinary structured columns inside a single pipeline.

The project's design choice is Python at the surface and Rust underneath. You write ordinary Python and never touch the JVM, while the query optimizer, vectorized execution engine and Arrow-backed memory layout live in Rust. That pairing is what lets the same code run on a laptop and then scale out unchanged to a distributed Ray or Kubernetes cluster.

Daft also ships AI operations as first-class dataframe expressions: running LLM prompts, generating embeddings and classifying rows at scale using OpenAI, Transformers or your own models. It reads from S3, GCS, Iceberg, Delta Lake, Hugging Face and Unity Catalog, which makes it a practical bridge between an existing lakehouse and a model-training or inference pipeline.

What it does

  • Native multimodal columns — process images, audio, video and embeddings next to structured data in one framework
  • Built-in AI operations: run LLM prompts, generate embeddings and classify data at scale with OpenAI, Transformers or custom models
  • Python-native with a Rust core: no JVM, with a query optimizer and vectorized, Arrow-backed execution engine
  • Start local, then scale the same code to distributed clusters on Ray or Kubernetes
  • Universal connectivity to S3, GCS, Iceberg, Delta Lake, Hugging Face and Unity Catalog
  • Out-of-core execution with intelligent memory management, so datasets larger than RAM still run

Getting started

Daft installs as a single Python package and requires Python 3.10 or higher. The project's quickstart walks through loading a real e-commerce dataset, processing its product images and running AI inference over them.

Install Daft

One pip install pulls in the Rust engine as a prebuilt wheel.

bashbash
pip install daft

Add distributed and cloud extras when you need them

Ray and AWS support are optional extras; see the installation guide for building from source.

bashbash
pip install "daft[ray]"

Work through the quickstart

The docs quickstart loads a real-world e-commerce dataset, processes product images and runs AI inference at scale — the fastest way to see multimodal columns and AI expressions together. Full API reference and per-topic user guides are at docs.daft.ai.

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Build a training-data pipeline that decodes and transforms millions of images or audio files without leaving Python
  • Run LLM prompts or embedding generation as a dataframe column over a whole table, batched and parallelized for you
  • Query multimodal data already sitting in Iceberg, Delta Lake or Hugging Face without copying it into a new store
  • Prototype a data pipeline locally, then run the identical code on a Ray or Kubernetes cluster when the dataset outgrows one machine

How Daft compares

Daft alongside other open-source data wrangling tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Hugging Face Datasets★ 22kAn Apache Arrow-backed Python library that loads datasets from the Hugging Face Hub or local files in one line and maps, filters and streams them without holding them in RAM.
Rerun★ 11.5kA data layer for physical AI that logs, visualises and queries multi-rate multimodal streams — images, point clouds, transforms, joint states, video — and streams them straight into training.
Datasette★ 11.5kAn open source multi-tool that points at a SQLite file and serves it as a browsable website with a JSON API, plus commands for publishing the result online.
ChatLab★ 7.5kA local-first desktop app that imports chat exports from eight messaging platforms into one normalized store, then queries them with SQL and tool-calling AI agents.
Lance★ 7.1kAn open lakehouse format for multimodal AI: one dataset holding images, video, audio, text and embeddings, with 100x faster random access than Parquet, vector and full-text indices, and zero-copy versioning.
Daft★ 5.8kA Python-native, Rust-powered data engine for multimodal AI workloads
sqlite-utils★ 2.2kA Python CLI and library that turns JSON, CSV and TSV into SQLite databases, creating schemas automatically and adding full-text search, table transforms and migrations.