AI/TLDR

OpenRouter · 2026-08-21 · notable

OpenRouter Image Benchmarks — 39 image models on one page of hard prompts

OpenRouter Image Benchmarks runs the same set of deliberately hard prompts through every image model it hosts and shows the raw pictures side by side, sortable by cost or generation time. Free to view, no account needed.

OpenRouter Image Benchmarks comparison grid

Every image model OpenRouter hosts, run through the same hard prompts, with the raw outputs shown in one grid.

Key specs

Image models39
Challenge families7

What is it?

Image Benchmarks is a free public page from OpenRouter that puts the same deliberately hard prompts through every image model on the platform and shows the actual pictures in a grid. The announcement counts 39 image models. You judge the outputs with your own eyes instead of reading a single score.

How does it work?

Prompts are grouped into seven challenge families — improbable scenes, counting, text, spatial relations, negation, editing and consistency — each aimed at a known weak spot, with named tests such as Full Wine Glass, Three Fingers, Long Exact String, Occlusion, Minimal Diff and Four References. OpenRouter notes that 'LLM-as-a-judge evals can't yet capture the details a human would notice,' which is why the page shows raw output rather than a graded number.

Why does it matter?

OpenRouter opens by saying that next to the many LLM benchmarks, 'picking an image model can feel arbitrary.' The grid sorts by cost or by time, and the listed models span roughly $0.006 to $0.25 per image and 3.6 to 208.2 seconds per generation, so a team can see what the cheap and slow ends of that range actually produce before committing.

Who is it for?

developers and designers choosing an image model

Try it

https://openrouter.ai/benchmarks/media/images

Sources

Tags

  • openrouter
  • image-generation
  • benchmark
  • evaluation
  • text-to-image
  • model-comparison
  • image-editing

← All releases · Learn AI