AI/TLDR

Fugu Max

Sakana AI's cost-first orchestration model, released September 11, 2026 at $2 / $6 per 1M tokens.

Fugu MaxAPI onlyGenerally available — served through Sakana AI's OpenAI-compatible API.
Released
11 Sep 2026
Input
$2.00 / 1M tokens
License
Proprietary
Coverage
1 story

Overview

Fugu Max is the cost-first tier of Sakana AI's Fugu family, released on 11 September 2026 alongside the capability-first Fugu Ultra v2. The two share what Sakana calls "the same core orchestration architecture optimized for two distinct missions": Fugu Max prioritises cost-efficiency, Fugu Ultra v2 maximum capability. Fugu itself is not a trained network but "a Multi-Agent System, Delivered as One Model" — a pool of separately-trained models coordinated behind a single OpenAI-compatible endpoint.

What Fugu Max adds is breadth of pool. Sakana writes that it "expands the pool of models Sakana Fugu can orchestrate, integrating an unprecedented number of open-weights and specialized models, including NVIDIA Nemotron family through our collaboration with NVIDIA". On results, Sakana reports that "Fugu Max achieves best overall score on six benchmarks", naming Terminal Bench 2.1, GPQAD, AA-LCR and AutomationBench among them; it does not publish a full scored table for those six, so no benchmark bars are listed on this page.

The pitch is price at that level of capability: $2 per million input tokens and $6 per million output, which Sakana describes as output pricing "40-60% lower than Sonnet 5, GPT 5.6 Terra, and Kimi K3". Cached input is $0.25 per million and web-service calls are billed at $0.007 each; unlike the Ultra tier, the rate does not change with context length. Sakana says both models are "available today via our standard OpenAI-compatible API", reachable with a single-line parameter change for existing Fugu users.

Released2026-09-11
LicenseProprietary
WeightsAPI only
ArchitectureMulti-agent orchestration — a pool of separately-trained models coordinated behind one endpoint rather than a single network. Sakana states the pool membership and routing policy are proprietary and not exposed.
ModalitiesText
StatusGenerally available — served through Sakana AI's OpenAI-compatible API.

Pricing

Input$2.00 / 1M tokens
Cached input$0.25 / 1M tokens
Output$6.00 / 1M tokens

Flat regardless of context length. Sakana's product page also lists web-service calls at $0.007 per call.

Pricing source ↗

Strengths

  • $2 per million input and $6 per million output tokens, which Sakana positions as 40–60% below Sonnet 5, GPT 5.6 Terra and Kimi K3 on output
  • Flat pricing regardless of context length, unlike the Fugu Ultra tier's rate step above 272K tokens
  • Cheap cached input at $0.25 per million tokens for repeated-prefix workloads
  • Orchestrates a broadened pool of open-weight and specialist models, including the NVIDIA Nemotron family
  • Sakana reports best overall score on six benchmarks, among them Terminal Bench 2.1, GPQAD, AA-LCR and AutomationBench

Best for

  • High-volume agent traffic where cost per token decides whether the workload is viable
  • Terminal and automation agents, the ground Terminal Bench 2.1 and AutomationBench cover
  • Long-document workloads that want a flat rate rather than a long-context surcharge
  • Teams already calling Fugu who want a cheaper tier without changing their client

How to access

ProviderModel ID
Sakana AI API ↗

FAQ

What is Fugu Max?

It is the cost-first tier of Sakana AI's Fugu family, released on 11 September 2026. Fugu is a multi-agent system delivered as one model: it coordinates a pool of separately-trained models behind a single OpenAI-compatible endpoint.

What does Fugu Max cost?

$2 per million input tokens, $6 per million output tokens and $0.25 per million cached input tokens, flat regardless of context length. Sakana's product page also lists web-service calls at $0.007 each.

How is it different from Fugu Ultra v2?

Same orchestration architecture, different mission. Fugu Max optimises for cost; Fugu Ultra v2 optimises for capability on complex reasoning and is priced at $5 / $30 per million tokens with a higher rate above 272K tokens of context.

Which models does it orchestrate?

Sakana says Fugu Max integrates "an unprecedented number of open-weights and specialized models, including NVIDIA Nemotron family", but states that the specific models Fugu selects and how it coordinates them are proprietary and not exposed.

Does Sakana publish benchmark scores for Fugu Max?

Sakana reports that Fugu Max "achieves best overall score on six benchmarks", naming Terminal Bench 2.1, GPQAD, AA-LCR and AutomationBench, but does not publish the scored table for them in the launch post.