█

AI/TLDR

React Native ExecuTorch

Software Mansion's library for running LLMs, vision, speech and embedding models on the phone inside a React Native app, built on Meta's ExecuTorch runtime

Local RuntimesOpen source
Latest
v0.10.3
Updated
25 Sep 2026
Language
TypeScript
License
MIT (bundled third-party components such as ExecuTorch, OpenCV and MLX keep their own BSD-3-Clause, Apache-2.0 and MIT licences)

What's new

v0.10.325 Sep 2026

Patch release moving to the v0.10.4-libs native build, whose XNNPACK backend fixes PReLU models returning NaN and Kokoro TTS being killed on iPhone, plus LLM and install fixes. `LLMRunner.generate` no longer takes `echo`.

Overview

Phone screen recording of an LLM chat screen running LFM 2.5 1.2B on the device's CPU, answering a prompt about on-device AI
LLM chat with LFM 2.5 (1.2B) running on the phoneReact Native ExecuTorch Gallery ↗

React Native ExecuTorch is an on-device AI inference library for React Native, built by Software Mansion on top of ExecuTorch, Meta's on-device inference runtime for PyTorch models. It runs machine learning models directly on the user's phone: after a model has been downloaded there are no network calls, the app works offline, and the user's data never leaves the device.

The library ships with a curated catalogue of pre-exported models covering large and vision-language models, computer vision (object detection, segmentation, classification, OCR), speech-to-text with Whisper, text-to-speech with Kokoro, and image and text embeddings. They are published in Software Mansion's Hugging Face collections and addressed from code through a typed `models` registry, so picking a model is a single import. You can also bring your own `.pte` models exported with ExecuTorch and plug them into the existing pipelines, or build new ones.

The API has two layers. On top sit ready-made task hooks such as `useLLMChatSession`, which handle downloading, caching and the model's lifecycle and hand back a ready flag and progress. Underneath is a lower-level core that exposes the ExecuTorch runtime, tensors and fast native operators, so custom multi-stage pipelines can be written entirely in TypeScript, with work moved off the JS thread via `react-native-worklets`. Execution is accelerated per platform through XNNPACK on the CPU, Core ML and MLX on Apple silicon, and Vulkan on Android GPUs.

The library itself is MIT-licensed. It bundles or links third-party components under their own permissive licences, among them ExecuTorch (BSD 3-Clause), OpenCV (Apache-2.0) and MLX (MIT), as listed in the repository's LICENSE file. Software Mansion also uses it in its own Private Mind app on the App Store and Google Play.

What it does

  • Task hooks (`use<Task>`) for LLM chat, computer vision, speech and embeddings, with automatic model download, caching and lifecycle management
  • A typed `models` registry over Software Mansion's Hugging Face collections of pre-exported models, plus support for your own `.pte` exports
  • Hardware acceleration through XNNPACK (CPU), Core ML and MLX (Apple silicon) and Vulkan (Android GPU)
  • A lower-level core with tensors, native operators, schema validation and worklet threading for building custom pipelines in TypeScript
  • Model downloads go to a persistent cache keyed by URL, with progress reporting, cancellation and deduplication of concurrent downloads
  • A `features` list in `package.json` limits which native backends and libraries are pulled in, to keep app size and install time down

Getting started

React Native ExecuTorch needs the New Architecture, React Native 0.83+ or Expo SDK 55+ with a development build (Expo Go is not supported because of its custom C++ native libraries), `react-native-worklets` 0.10 or newer, and iOS 17.0+ or Android 13+ (minSdkVersion 26). The steps below follow the project's README and Getting Started guide.

VideoSoftware Mansion's original announcement of the library, which predates the v0.10 API rewriteSoftware Mansion ↗

Install the package and its peer dependencies

Add the library together with `react-native-worklets` and `react-native-blob-util`. On Expo SDK 55 or 56 the docs pin the worklets package explicitly.

bashbash
npm install react-native-executorch react-native-worklets react-native-blob-util
# or
yarn add react-native-executorch react-native-worklets react-native-blob-util
# or
pnpm add react-native-executorch react-native-worklets react-native-blob-util

# Expo SDK 55/56
npm install react-native-worklets@^0.10.0

Optionally declare only the features you use

A `react-native-executorch` block in `package.json` takes high-level task names; each expands to the backends and native libraries it needs, so the install only downloads those.

jsonjson
{
  "react-native-executorch": {
    "features": ["classification", "styleTransfer"]
  }
}

Run an LLM with a task hook

`useLLMChatSession` downloads the chosen model from the registry, reports progress while it loads, and streams tokens back through a callback once `isReady` is true.

tsxtsx
import { Button, View } from 'react-native';
import { models, useLLMChatSession } from 'react-native-executorch';

export function App() {
  const session = useLLMChatSession(models.llm.LFM2_5_1_2B.DEFAULT);

  const handleGenerate = async () => {
    if (!session.isReady || !session.sendMessage) return;

    const turn = await session.sendMessage(
      'Explain on-device AI in one sentence.',
      (token) => console.log(token)
    );

    console.log('Result messages:', turn.messages);
  };

  return (
    <View style={{ flex: 1, justifyContent: 'center', alignItems: 'center' }}>
      <Button
        title={session.isReady ? 'Generate' : `Loading (${session.downloadProgress.toFixed(0)}%)`}
        onPress={handleGenerate}
        disabled={!session.isReady}
      />
    </View>
  );
}

Download models ahead of time

Outside a task hook, `useResourceDownload` or the imperative `download` function fetch and cache a model's files and return the same config with every URL replaced by a local path. Local paths are passed through untouched; only http(s) URLs are downloaded.

tsts
import { download, models } from 'react-native-executorch';

const model = await download(models.classification.EFFICIENTNET_V2_S.XNNPACK_FP32, {
  onProgress: (p) => console.log(`${Math.round(p * 100)}%`),
});

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Add a private, offline chat assistant to a React Native or Expo app without paying for or depending on a cloud LLM API
  • Run camera-driven computer vision on the phone, such as object detection, segmentation or OCR on a photographed document
  • Transcribe speech with Whisper or synthesise it with Kokoro on-device, so audio never leaves the user's phone
  • Build local semantic search or RAG over on-device content using the library's image and text embedding models
Phone screen recording of a text-to-speech screen using the Kokoro 82M model to synthesise a typed sentence
On-device text-to-speech with Kokoro 82M
Phone screen recording of an OCR screen running PaddleOCR PP-OCRv6 on-device over a photographed newspaper stock table
OCR on a photographed page with PaddleOCR PP-OCRv6

Version history

Every verified update to React Native ExecuTorch that AI/TLDR tracked, newest first — each links to our coverage and the official changeset.

  1. 2026-09-25v0.10.3

    Patch release moving to the v0.10.4-libs native build, whose XNNPACK backend fixes PReLU models returning NaN and Kokoro TTS being killed on iPhone, plus LLM and install fixes. `LLMRunner.generate` no longer takes `echo`.

  2. 2026-09-11v0.10.2

    Patch release fixing a build failure when opting out of a native library and an LLM error that was masked by its own rollback.

  3. 2026-09-08v0.10.0

    A ground-up rewrite: task pipelines move into inspectable TypeScript over a lower-level native core, and Core ML, MLX and Vulkan acceleration join the XNNPACK CPU path across more than 130 pre-exported model variants, with a legacy module for gradual migration.

How React Native ExecuTorch compares

React Native ExecuTorch alongside other open-source local runtimes tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Ollama★ 182kA developer-friendly tool that downloads and runs local LLMs from the terminal with a built-in OpenAI-compatible API.
llama.cpp★ 130kA C/C++ inference engine that runs LLMs in the GGUF format on CPUs, Apple Silicon, and GPUs with low memory use.
GPT4All★ 77.4kGPT4All is a free desktop app and Python client that runs large language models locally on your own computer, with no API calls or GPU required.
LocalAI★ 49.3kA self-hosted server that exposes an OpenAI-compatible API for running text, vision, voice, and image models on local hardware.
Jan★ 44.7kAn open-source desktop app that runs LLMs fully offline as a ChatGPT-style assistant on your own computer.
Colibrì★ 38kA pure-C inference engine that keeps a Mixture-of-Experts model's dense trunk resident in RAM and streams its routed experts from disk, so 744B-2.8T models run on consumer hardware.
llmfit★ 37.2kA Rust terminal tool that inspects your CPU, RAM, GPUs and VRAM and scores which open-weight models and quantizations will actually run well on that machine, with a TUI, CLI, REST API and local-runtime integrations.
React Native ExecuTorch★ 1.7kSoftware Mansion's library for running LLMs, vision, speech and embedding models on the phone inside a React Native app, built on Meta's ExecuTorch runtime