GenAI Benchmarks

Model rankings

Headline criteria for every model in the benchmark. Click a model for per-category scores and its full result gallery. The same table answers the routing question for an AI design agent generating on a user's behalf.

“Overall” gives each scored criterion one equal vote: the auto measurements (resolution, latency, cost, transparency, color accuracy) plus the interim Claude-VLM quality tier (text, prompt adherence, composition). Equal votes mean a model that fails one criterion outright cannot buy its way back with cost and speed. To rank by what a specific job needs, use the use-case pages, which reweight the same numbers.
Three cartoon avocados in model jerseys on a winners' podium
Pilot dataset (pilot-0): auto criteria + interim Claude-VLM quality tier, unweighted. Leader pack: within 0.25 of the top overall score.
# Model $ / image p50 latency Native output Overall Details
1 $0.140 34.0s 1024×1024 3.74 / 5 See scores
2 $0.035 10.7s 1024×1024 3.63 / 5 See scores
3 $0.010 17.8s 1920×1280 3.62 / 5 See scores
4 $0.042 4.0s 1408×768 3.58 / 5 See scores
5
Seedream 4.5 ByteDance
$0.040 12.4s 2048×2048 3.53 / 5 See scores
6 $0.035 30.7s 2048×2048 3.52 / 5 See scores
7
FLUX.2 Black Forest Labs
$0.012 2.0s 1024×768 3.50 / 5 See scores
8 $0.135 65.3s 2048×2048 3.49 / 5 See scores
9
FLUX.2 [dev] Turbo Black Forest Labs
$0.008 2.0s 1024×768 3.47 / 5 See scores
10 $0.080 13.2s 1408×768 3.34 / 5 See scores
11 $0.040 7.4s 1024×1024 3.27 / 5 See scores
12 $0.036 30.4s 1024×768 3.27 / 5 See scores
13 $0.042 21.3s 1024×768 3.26 / 5 See scores
14 $0.060 65.9s 1024×1024 3.22 / 5 See scores
15
FLUX.2 [pro] Black Forest Labs
$0.030 11.5s 1024×768 3.18 / 5 See scores
16
Qwen-Image Alibaba
$0.020 6.9s 1024×768 3.18 / 5 See scores
17 $0.150 23.3s 1024×1024 3.13 / 5 See scores
18 $0.021 16.1s 1536×1536 3.10 / 5 See scores
19 $0.060 17.9s 1024×1024 3.08 / 5 See scores
20
Recraft V3 Recraft
$0.040 7.4s 1024×1024 3.08 / 5 See scores
21 $0.040 141.1s 1024×1024 3.06 / 5 See scores
22 $0.155 110.1s 1024×768 3.04 / 5 See scores
23 $0.002 2.1s 512×512 2.30 / 5 See scores

Want a head-to-head? Pick two or three models and compare every score and every prompt side by side.

Open the comparison tool

Cost vs quality

The buyer's chart: models on the Pareto frontier are undominated; for every faded dot there is a model that is both cheaper and better on the measured criteria.

2 2.5 3 3.5 4 $0.002 $0.005 $0.01 $0.02 $0.05 $0.10 $0.20 Cost per image (log scale) Overall score (all eight criteria) GPT Image 1.5 Recraft V4.1 Muse Image Nano Banana 2 Lite Seedream 4.5 FLUX.2 Seedream 5.0 Pro Seedream 5.0 Lite FLUX.2 [dev] Turbo Nano Banana 2 Gemini 2.5 Flash Image GPT Image 2.5 Sunburst GPT Image 2.5 Flare Qwen-Image Grok Imagine Image 2.0 FLUX.2 [pro] Nano Banana Pro Luma Photon Recraft V3 Ideogram 3.0 GPT Image 2 Qwen Image 3.0 Stable Diffusion 1.5
Score axis is zoomed to 2–4, the band every model falls in. Models on the highlighted Pareto frontier (Stable Diffusion 1.5, FLUX.2 [dev] Turbo, Muse Image, Recraft V4.1, GPT Image 1.5) are undominated: no other model is both cheaper and better on these criteria. Faded dots are dominated. Hover or tap a point for its exact cost and score and a link to the full scorecard; the table above carries the same data.