Overview
Claude Haiku 5.5 is the Haiku-tier model of Anthropic's Claude 5.5 family, released on October 7, 2026. Anthropic describes it as the cheapest, fastest and most capable small model it has released, built for high-volume, latency-sensitive work such as classification, routing, extraction, summaries, live customer support and subagent tasks.
The model has a 1 million-token context window and up to 128K output tokens on the synchronous Messages API, up from 200K and 64K on Claude Haiku 4.5; the Message Batches API allows up to 300K output tokens with the output-300k-2026-03-24 beta header. It takes text and images in and returns text, and reliable knowledge runs through June 2026. Adaptive thinking is on by default and its depth is set with the effort parameter, which defaults to medium. It uses the newer tokenizer of Claude 4.7 and later, so the same text counts as about 30% more tokens than on Haiku 4.5.
Pricing depends on prompt length. For prompts up to 100,000 tokens it costs $0.10 per million input tokens and $0.50 per million output tokens; longer prompts cost $0.50 and $2.50. Cache reads cost $0.01 per million on short prompts, and Batch API requests take 50% off. Anthropic says this is about 90% cheaper than Claude Haiku 4.5 for requests under 100K tokens.
Claude Haiku 5.5 adds the browser use tool and supports computer use through the computer_toolset_20260801 toolset on the Claude API and Google Cloud. It runs safety classifiers that can decline a request with no server-side fallback. Code written for Haiku 4.5 needs changes: manual thinking budgets, non-default sampling parameters and assistant prefill all return errors, and Priority Tier is not supported.
| Released | 2026-10-07 |
|---|---|
| License | Proprietary |
| Weights | API only |
| Context | 1M |
| Max output | 128K |
| Architecture | Proprietary transformer with adaptive thinking on by default, steered by an effort parameter that defaults to medium; uses the tokenizer introduced with Claude Opus 4.7. |
| Knowledge cutoff | Jun 2026 |
| Modalities | Text, Vision |
| Status | Generally available |
Benchmarks
Claude Haiku 5.5 vs Claude Haiku 4.5, GPT-6 Luna and Claude Sonnet 5.5
| Benchmark | Claude Haiku 5.5 | Claude Haiku 4.5 | GPT-6 Luna | Claude Sonnet 5.5 |
|---|---|---|---|---|
| GDPval-AA v2.1 (Elo) | 1620 | 735 | 1437 | 1840 |
| AA-Briefcase v1.1 (Elo) | 1578 | 614 | 1336 | 1824 |
| OSWorld 2.1 | 72.4% | 15.7% | 48.9% | 83.9% |
| Humanity's Last Exam (no tools) | 45.9% | 10.2% | — | 56.9% |
| Humanity's Last Exam (with tools) | 57.4% | 18.7% | — | 64.5% |
| Terminal-Bench 4.0 | 39.2% | 0% | 16.4% | 70.6% |
| FrontierCode 1.1 | 46.4% | — | 42.4% | 52.1% |
| Chartography (no tools) | 46.4% | 6.4% | 29.1% | 61.6% |
This model's scores
- OSWorld 2.1 (computer use)72.4%
- Humanity's Last Exam (with tools)57.4%
- FrontierCode 1.146.4%
- Chartography (no tools)46.4%
- Terminal-Bench 4.0 (agentic coding)39.2%
Scores on a 0–100 scale (25-point gridlines); higher is better. Each benchmark links to its published source.
Pricing
| Input | $0.10 per million tokens |
|---|---|
| Output | $0.50 per million tokens |
For prompts up to 100K tokens; prompts over 100K cost $0.50 input / $2.50 output. Cache reads $0.01/MTok, 5m cache write $0.125/MTok, 1h cache write $0.20/MTok (short prompts). Batch API takes 50% off input and output.
Strengths
- 72.4% on OSWorld 2.1 in Anthropic's launch comparison, against 15.7% for Claude Haiku 4.5 and 48.9% for GPT-6 Luna
- From $0.10 input and $0.50 output per million tokens, about 90% below Claude Haiku 4.5 for prompts under 100K tokens
- 1M-token context with 128K max output, up from 200K and 64K on Haiku 4.5
- Browser use tool and computer-use toolset support, which Haiku 4.5 does not have
- Adaptive thinking with five effort levels, defaulting to medium
Best for
- Reach for it for high-volume classification, routing and extraction where price per call decides the design.
- Reach for it as a subagent model inside larger agent systems, where many fast, cheap calls run in parallel.
- Reach for it for latency-sensitive browser and computer-use automation, such as live customer support flows.
How to access
| Provider | Model ID |
|---|---|
| Anthropic API ↗ | claude-haiku-5-5 |
| Amazon Bedrock ↗ | anthropic.claude-haiku-5-5 |
| Google Vertex AI ↗ | claude-haiku-5-5 |
| Microsoft Foundry ↗ | claude-haiku-5-5 |
Claude Haiku — every version
The full lineage of the Claude Haiku line, newest first. Every version has its own page — click any to compare specs, benchmarks and pricing.
| Version | Released | Context | License |
|---|---|---|---|
| Claude Haiku 5.5current | 2026-10-07 | 1M | Proprietary |
| Claude Haiku 4.5 | 2025-10-15 | 200K | Proprietary |
| Claude 3.5 Haiku | 2024-10-22 | — | Proprietary |
| Claude 3 Haiku | 2024-03-13 | — | Proprietary |
FAQ
When was Claude Haiku 5.5 released?
Anthropic released Claude Haiku 5.5 on October 7, 2026, on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS, and in claude.ai. Anthropic commits to keeping it available until at least October 7, 2027. Its model ID, claude-haiku-5-5, is fixed with no date suffix.
How much does Claude Haiku 5.5 cost?
Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens, and $0.50 / $2.50 for longer prompts. Cache reads are $0.01 per million on short prompts and Batch API requests take 50% off. Anthropic says that is about 90% cheaper than Claude Haiku 4.5 under 100K tokens.
How does Claude Haiku 5.5 compare to GPT-6 Luna?
In Anthropic's launch comparison Claude Haiku 5.5 leads OpenAI's GPT-6 Luna on every shared benchmark: 72.4% to 48.9% on OSWorld 2.1, 39.2% to 16.4% on Terminal-Bench 4.0, 46.4% to 42.4% on FrontierCode 1.1 and 46.4% to 29.1% on Chartography. These are Anthropic's own published figures.
What breaks when migrating from Claude Haiku 4.5 to Claude Haiku 5.5?
Claude Haiku 5.5 rejects manual thinking budgets, non-default temperature, top_p or top_k, assistant prefill, and the computer_20250124 tool on the Claude API and Google Cloud. Changing earlier turns invalidates returned thinking blocks, the same text counts as about 30% more tokens, and Priority Tier is not supported. Anthropic's migration guide covers each change.
