AI/TLDR

Ollama

Download and run open LLMs locally from your terminal

Local RuntimesOpen source
Latest
v0.34.1
Updated
14 Sep 2026
Language
Go
Coverage
3 stories
$curl -fsSL https://ollama.com/install.sh | sh

What's new

v0.34.114 Sep 2026

MLX safetensors support in ollama create is no longer experimental. Creating a GGUF model now requires llama.cpp tooling for safetensor conversion and quantization. The release also improves MLX memory handling on Apple Silicon, raises repeat-token detection to 100 tokens, and cuts the /api/tags cold load from 3.1 seconds to 294 milliseconds.

Latest news

Overview

Ollama is a tool for downloading and running open large language models on your own computer. You install it on macOS, Windows, or Linux, then pull and chat with a model straight from the terminal with a single command.

It is aimed at developers who want to run models locally instead of calling a hosted service. Beyond the CLI, Ollama exposes a REST API on localhost and ships official Python and JavaScript libraries, so you can wire local models into your own apps and scripts.

As a local runtime, Ollama handles model downloads, serving, and the request loop for you, and connects to existing coding tools and agents such as Claude Code, Codex, and Copilot CLI through its launch integrations.

What it does

  • One-line install script for macOS, Windows, and Linux, plus an official Docker image
  • Run any model from the library with a single `ollama run` command
  • Built-in REST API on http://localhost:11434 for running and managing models
  • Official Python (`pip install ollama`) and JavaScript (`npm i ollama`) client libraries
  • Launch integrations for coding tools and agents like Claude Code, Codex, and Copilot CLI
  • Built on the llama.cpp backend for local model inference

Getting started

Install Ollama, then pull and chat with a model from the terminal. The same models are reachable over a local REST API and the Python and JavaScript libraries.

Install Ollama

Run the install script on macOS or Linux. On Windows, use the PowerShell command instead.

bashbash
curl -fsSL https://ollama.com/install.sh | sh

Run and chat with a model

Pull a model from the library and start chatting in the terminal.

bashbash
ollama run gemma4

Call the REST API

Ollama serves a local REST API on port 11434 for running models from your own apps.

bashbash
curl http://localhost:11434/api/chat -d '{
  "model": "gemma4",
  "messages": [{
    "role": "user",
    "content": "Why is the sky blue?"
  }],
  "stream": false
}'

Use the Python library

Install the official client and send a chat request from Python.

pythonpython
pip install ollama

from ollama import chat

response = chat(model='gemma4', messages=[
  {
    'role': 'user',
    'content': 'Why is the sky blue?',
  },
])
print(response.message.content)

Commands and code are distilled from the project's own documentation — always check the official repo for the latest.

When to use it

  • Run open LLMs locally for privacy or offline work, without sending data to a hosted API
  • Add a local model backend to a Python or JavaScript app through the official libraries
  • Connect Ollama to coding tools and agents such as Claude Code, Codex, or Copilot CLI
  • Prototype and test prompts against different models from the library before committing to a provider

Version history

Every verified update to Ollama that AI/TLDR tracked, newest first — each links to our coverage and the official changeset.

  1. 2026-09-14v0.34.1

    MLX safetensors support in ollama create is no longer experimental. Creating a GGUF model now requires llama.cpp tooling for safetensor conversion and quantization. The release also improves MLX memory handling on Apple Silicon, raises repeat-token detection to 100 tokens, and cuts the /api/tags cold load from 3.1 seconds to 294 milliseconds.

  2. 2026-09-05v0.34.0-rc1

    Ollama models can now be used directly in ChatGPT Desktop, with setup available from the Ollama app on macOS. The release also improves structured output performance on Apple Silicon and adds support for OpenAI-compatible client tool search and response compaction.

How Ollama compares

Ollama alongside other open-source local runtimes tools AI/TLDR tracks, ranked by GitHub stars.

ToolStarsWhat it does
Ollama★ 181kDownload and run open LLMs locally from your terminal
llama.cpp★ 129kA C/C++ inference engine that runs LLMs in the GGUF format on CPUs, Apple Silicon, and GPUs with low memory use.
GPT4All★ 77.4kGPT4All is a free desktop app and Python client that runs large language models locally on your own computer, with no API calls or GPU required.
LocalAI★ 49.1kA self-hosted server that exposes an OpenAI-compatible API for running text, vision, voice, and image models on local hardware.
Jan★ 44.5kAn open-source desktop app that runs LLMs fully offline as a ChatGPT-style assistant on your own computer.
llmfit★ 36.7kA Rust terminal tool that inspects your CPU, RAM, GPUs and VRAM and scores which open-weight models and quantizations will actually run well on that machine, with a TUI, CLI, REST API and local-runtime integrations.
Colibrì★ 35.8kA pure-C inference engine that keeps a Mixture-of-Experts model's dense trunk resident in RAM and streams its routed experts from disk, so 744B-2.8T models run on consumer hardware.
AirLLM★ 34.5kA Python inference library that keeps only one transformer layer on the GPU at a time, so a 70B model runs on a single 4GB card and a 671B MoE model on about 12GB, without quantization.