Overview

React Native ExecuTorch is an on-device AI inference library for React Native, built by Software Mansion on top of ExecuTorch, Meta's on-device inference runtime for PyTorch models. It runs machine learning models directly on the user's phone: after a model has been downloaded there are no network calls, the app works offline, and the user's data never leaves the device.
The library ships with a curated catalogue of pre-exported models covering large and vision-language models, computer vision (object detection, segmentation, classification, OCR), speech-to-text with Whisper, text-to-speech with Kokoro, and image and text embeddings. They are published in Software Mansion's Hugging Face collections and addressed from code through a typed `models` registry, so picking a model is a single import. You can also bring your own `.pte` models exported with ExecuTorch and plug them into the existing pipelines, or build new ones.
The API has two layers. On top sit ready-made task hooks such as `useLLMChatSession`, which handle downloading, caching and the model's lifecycle and hand back a ready flag and progress. Underneath is a lower-level core that exposes the ExecuTorch runtime, tensors and fast native operators, so custom multi-stage pipelines can be written entirely in TypeScript, with work moved off the JS thread via `react-native-worklets`. Execution is accelerated per platform through XNNPACK on the CPU, Core ML and MLX on Apple silicon, and Vulkan on Android GPUs.
The library itself is MIT-licensed. It bundles or links third-party components under their own permissive licences, among them ExecuTorch (BSD 3-Clause), OpenCV (Apache-2.0) and MLX (MIT), as listed in the repository's LICENSE file. Software Mansion also uses it in its own Private Mind app on the App Store and Google Play.
What it does
- Task hooks (`use<Task>`) for LLM chat, computer vision, speech and embeddings, with automatic model download, caching and lifecycle management
- A typed `models` registry over Software Mansion's Hugging Face collections of pre-exported models, plus support for your own `.pte` exports
- Hardware acceleration through XNNPACK (CPU), Core ML and MLX (Apple silicon) and Vulkan (Android GPU)
- A lower-level core with tensors, native operators, schema validation and worklet threading for building custom pipelines in TypeScript
- Model downloads go to a persistent cache keyed by URL, with progress reporting, cancellation and deduplication of concurrent downloads
- A `features` list in `package.json` limits which native backends and libraries are pulled in, to keep app size and install time down
Getting started
React Native ExecuTorch needs the New Architecture, React Native 0.83+ or Expo SDK 55+ with a development build (Expo Go is not supported because of its custom C++ native libraries), `react-native-worklets` 0.10 or newer, and iOS 17.0+ or Android 13+ (minSdkVersion 26). The steps below follow the project's README and Getting Started guide.

Install the package and its peer dependencies
Add the library together with `react-native-worklets` and `react-native-blob-util`. On Expo SDK 55 or 56 the docs pin the worklets package explicitly.
npm install react-native-executorch react-native-worklets react-native-blob-util
# or
yarn add react-native-executorch react-native-worklets react-native-blob-util
# or
pnpm add react-native-executorch react-native-worklets react-native-blob-util
# Expo SDK 55/56
npm install react-native-worklets@^0.10.0Optionally declare only the features you use
A `react-native-executorch` block in `package.json` takes high-level task names; each expands to the backends and native libraries it needs, so the install only downloads those.
{
"react-native-executorch": {
"features": ["classification", "styleTransfer"]
}
}Run an LLM with a task hook
`useLLMChatSession` downloads the chosen model from the registry, reports progress while it loads, and streams tokens back through a callback once `isReady` is true.
import { Button, View } from 'react-native';
import { models, useLLMChatSession } from 'react-native-executorch';
export function App() {
const session = useLLMChatSession(models.llm.LFM2_5_1_2B.DEFAULT);
const handleGenerate = async () => {
if (!session.isReady || !session.sendMessage) return;
const turn = await session.sendMessage(
'Explain on-device AI in one sentence.',
(token) => console.log(token)
);
console.log('Result messages:', turn.messages);
};
return (
<View style={{ flex: 1, justifyContent: 'center', alignItems: 'center' }}>
<Button
title={session.isReady ? 'Generate' : `Loading (${session.downloadProgress.toFixed(0)}%)`}
onPress={handleGenerate}
disabled={!session.isReady}
/>
</View>
);
}Download models ahead of time
Outside a task hook, `useResourceDownload` or the imperative `download` function fetch and cache a model's files and return the same config with every URL replaced by a local path. Local paths are passed through untouched; only http(s) URLs are downloaded.
import { download, models } from 'react-native-executorch';
const model = await download(models.classification.EFFICIENTNET_V2_S.XNNPACK_FP32, {
onProgress: (p) => console.log(`${Math.round(p * 100)}%`),
});Commands and code are distilled from the project's own documentation — always check the official repo for the latest.
When to use it
- Add a private, offline chat assistant to a React Native or Expo app without paying for or depending on a cloud LLM API
- Run camera-driven computer vision on the phone, such as object detection, segmentation or OCR on a photographed document
- Transcribe speech with Whisper or synthesise it with Kokoro on-device, so audio never leaves the user's phone
- Build local semantic search or RAG over on-device content using the library's image and text embedding models
Version history
Every verified update to React Native ExecuTorch that AI/TLDR tracked, newest first — each links to our coverage and the official changeset.
- 2026-09-25v0.10.3
Patch release moving to the v0.10.4-libs native build, whose XNNPACK backend fixes PReLU models returning NaN and Kokoro TTS being killed on iPhone, plus LLM and install fixes. `LLMRunner.generate` no longer takes `echo`.
- 2026-09-11v0.10.2
Patch release fixing a build failure when opting out of a native library and an LLM error that was masked by its own rollback.
- 2026-09-08v0.10.0
A ground-up rewrite: task pipelines move into inspectable TypeScript over a lower-level native core, and Core ML, MLX and Vulkan acceleration join the XNNPACK CPU path across more than 130 pre-exported model variants, with a legacy module for gradual migration.
How React Native ExecuTorch compares
React Native ExecuTorch alongside other open-source local runtimes tools AI/TLDR tracks, ranked by GitHub stars.
| Tool | Stars | What it does |
|---|---|---|
| Ollama | ★ 182k | A developer-friendly tool that downloads and runs local LLMs from the terminal with a built-in OpenAI-compatible API. |
| llama.cpp | ★ 130k | A C/C++ inference engine that runs LLMs in the GGUF format on CPUs, Apple Silicon, and GPUs with low memory use. |
| GPT4All | ★ 77.4k | GPT4All is a free desktop app and Python client that runs large language models locally on your own computer, with no API calls or GPU required. |
| LocalAI | ★ 49.3k | A self-hosted server that exposes an OpenAI-compatible API for running text, vision, voice, and image models on local hardware. |
| Jan | ★ 44.7k | An open-source desktop app that runs LLMs fully offline as a ChatGPT-style assistant on your own computer. |
| Colibrì | ★ 38k | A pure-C inference engine that keeps a Mixture-of-Experts model's dense trunk resident in RAM and streams its routed experts from disk, so 744B-2.8T models run on consumer hardware. |
| llmfit | ★ 37.2k | A Rust terminal tool that inspects your CPU, RAM, GPUs and VRAM and scores which open-weight models and quantizations will actually run well on that machine, with a TUI, CLI, REST API and local-runtime integrations. |
| React Native ExecuTorch | ★ 1.7k | Software Mansion's library for running LLMs, vision, speech and embedding models on the phone inside a React Native app, built on Meta's ExecuTorch runtime |

