GenAI Benchmarks
Model rankings
Headline criteria for every model in the benchmark. Click a model for per-category scores and its full result gallery. The same table answers the routing question for an AI design agent generating on a user's behalf.
“Overall” gives each scored criterion one equal vote: the auto
measurements (resolution, latency, cost, transparency, color accuracy)
plus the interim Claude-VLM quality tier (text, prompt adherence,
composition). Equal votes mean a model that fails one criterion
outright cannot buy its way back with cost and speed. To rank by what
a specific job needs, use the
use-case pages, which reweight the same numbers.
| # | Model | $ / image | p50 latency | Native output | Overall | Details |
|---|---|---|---|---|---|---|
| 1 | GPT Image 1.5 OpenAI | $0.140 | 34.0s | 1024×1024 | 3.74 / 5 | See scores |
| 2 | Recraft V4.1 Recraft | $0.035 | 10.7s | 1024×1024 | 3.63 / 5 | See scores |
| 3 | Muse Image Meta | $0.010 | 17.8s | 1920×1280 | 3.62 / 5 | See scores |
| 4 | Nano Banana 2 Lite Google | $0.042 | 4.0s | 1408×768 | 3.58 / 5 | See scores |
| 5 | Seedream 4.5 ByteDance | $0.040 | 12.4s | 2048×2048 | 3.53 / 5 | See scores |
| 6 | Seedream 5.0 Lite ByteDance | $0.035 | 30.7s | 2048×2048 | 3.52 / 5 | See scores |
| 7 | FLUX.2 Black Forest Labs | $0.012 | 2.0s | 1024×768 | 3.50 / 5 | See scores |
| 8 | Seedream 5.0 Pro ByteDance | $0.135 | 65.3s | 2048×2048 | 3.49 / 5 | See scores |
| 9 | FLUX.2 [dev] Turbo Black Forest Labs | $0.008 | 2.0s | 1024×768 | 3.47 / 5 | See scores |
| 10 | Nano Banana 2 Google | $0.080 | 13.2s | 1408×768 | 3.34 / 5 | See scores |
| 11 | Gemini 2.5 Flash Image Google | $0.040 | 7.4s | 1024×1024 | 3.27 / 5 | See scores |
| 12 | GPT Image 2.5 Sunburst OpenAI | $0.036 | 30.4s | 1024×768 | 3.27 / 5 | See scores |
| 13 | GPT Image 2.5 Flare OpenAI | $0.042 | 21.3s | 1024×768 | 3.26 / 5 | See scores |
| 14 | $0.060 | 65.9s | 1024×1024 | 3.22 / 5 | See scores | |
| 15 | FLUX.2 [pro] Black Forest Labs | $0.030 | 11.5s | 1024×768 | 3.18 / 5 | See scores |
| 16 | Qwen-Image Alibaba | $0.020 | 6.9s | 1024×768 | 3.18 / 5 | See scores |
| 17 | Nano Banana Pro Google | $0.150 | 23.3s | 1024×1024 | 3.13 / 5 | See scores |
| 18 | Luma Photon Luma AI | $0.021 | 16.1s | 1536×1536 | 3.10 / 5 | See scores |
| 19 | Ideogram 3.0 Ideogram | $0.060 | 17.9s | 1024×1024 | 3.08 / 5 | See scores |
| 20 | Recraft V3 Recraft | $0.040 | 7.4s | 1024×1024 | 3.08 / 5 | See scores |
| 21 | Qwen Image 3.0 Alibaba | $0.040 | 141.1s | 1024×1024 | 3.06 / 5 | See scores |
| 22 | GPT Image 2 OpenAI | $0.155 | 110.1s | 1024×768 | 3.04 / 5 | See scores |
| 23 | Stable Diffusion 1.5 Stability AI | $0.002 | 2.1s | 512×512 | 2.30 / 5 | See scores |
Want a head-to-head? Pick two or three models and compare every score and every prompt side by side.
Open the comparison toolCost vs quality
The buyer's chart: models on the Pareto frontier are undominated; for every faded dot there is a model that is both cheaper and better on the measured criteria.
New model? New benchmarks.
We re-run the identical 37-prompt suite on every model release. Get the scores, and what changed in the rankings, in your inbox.
You're on the list.
We'll email you when new benchmark results ship.

