GenAI Benchmarks

Model rankings

Headline criteria for every model in the benchmark. Click a model for per-category scores and its full result gallery.

“Overall” blends every scored criterion as an unweighted mean: the auto measurements (resolution, latency, cost, transparency, color accuracy) plus the interim Claude-VLM quality tier (text, prompt adherence, composition). Because it is unweighted, cost and latency pull cheap, fast models to the top and the premium flagship to the bottom. To rank by what a specific job needs, use the use-case pages, which reweight the same numbers.
Three cartoon avocados in model jerseys on a winners' podium: gold for Nano Banana, silver for Seedream, bronze for FLUX
Pilot dataset (pilot-0): auto criteria + interim Claude-VLM quality tier, unweighted. Leader pack: within 0.25 of the top overall score.
# Model Tier $ / image p50 latency Native output Overall Details
1
Budget
$0.020 4.0s 1408×768 3.87 / 5 See scores
2
Seedream 4.5 ByteDance
Gateway
$0.048 12.4s 2048×2048 3.80 / 5 See scores
3
Extended
$0.020 30.7s 2048×2048 3.65 / 5 See scores
4
FLUX.2 Black Forest Labs
Gateway
$0.013 2.0s 1024×768 3.64 / 5 See scores
5
FLUX.2 [dev] Turbo Black Forest Labs
Extended
$0.015 2.0s 1024×768 3.61 / 5 See scores
6
Budget
$0.039 7.4s 1024×1024 3.60 / 5 See scores
7
Extended
$0.019 16.1s 1536×1536 3.39 / 5 See scores
8
Recraft V3 Recraft
Extended
$0.040 7.4s 1024×1024 3.29 / 5 See scores
9
Extended
$0.100 13.2s 1408×768 3.24 / 5 See scores
10
FLUX.2 [pro] Black Forest Labs
Extended
$0.040 11.5s 1024×768 3.23 / 5 See scores
11
Qwen-Image Alibaba
Extended
$0.030 6.9s 1024×768 3.21 / 5 See scores
12
Extended
$0.120 34.0s 1024×1024 3.07 / 5 See scores
13
Extended
$0.060 17.9s 1024×1024 2.97 / 5 See scores
14
Extended
$0.010 2.1s 512×512 2.84 / 5 See scores
15
Gateway
$0.360 23.3s 1024×1024 2.55 / 5 See scores

Want a head-to-head? Pick two or three models and compare every score and every prompt side by side.

Open the comparison tool

Cost vs quality

The buyer's chart: models on the Pareto frontier are undominated; for every faded dot there is a model that is both cheaper and better on the measured criteria.

2 2.5 3 3.5 4 4.5 $0.01 $0.02 $0.05 $0.10 $0.20 $0.50 Cost per image (log scale) Overall score (auto criteria) Nano Banana 2 Lite Seedream 4.5 Seedream 5.0 Lite FLUX.2 Gemini 2.5 Flash Image FLUX.2 [dev] Turbo Luma Photon Recraft V3 FLUX.2 [pro] Nano Banana 2 Qwen-Image GPT Image 1.5 Ideogram 3.0 Stable Diffusion 1.5 Nano Banana Pro
Score axis is zoomed to 2–4.5, the band every model falls in. Models on the highlighted Pareto frontier (Stable Diffusion 1.5, FLUX.2, Nano Banana 2 Lite) are undominated: no other model is both cheaper and better on these criteria. Faded dots are dominated. Hover or tap a point for its exact cost and score and a link to the full scorecard; the table above carries the same data.