Structured JSON Scene
Structured JSON prompts are how creative-automation pipelines, agents and CE.SDK integrations actually drive image models, and nobody benchmarks it. Each spec field becomes a checklist item plus an automated color check, the modality closest to IMG.LY's product reality.
Render the following scene exactly: {"subject":"a lighthouse keeper","gaze":"toward the sea","wardrobe":"yellow raincoat","lighting":"golden hour from the left","composition":"rule of thirds, subject on the right","colorPlate":[{"name":"coat","hex":"#F4C430","target_pct":15},{"name":"sky","hex":"#F6A15C","target_pct":40}]}
Results
FLUX.2
FLUX.2 [pro]
FLUX.2 [dev] Turbo
Gemini 2.5 Flash Image
GPT Image 1.5
Ideogram 3.0
Luma Photon
Nano Banana 2
Nano Banana 2 Lite
Nano Banana Pro
Qwen-Image
Recraft V3
Seedream 4.5
Seedream 5.0 Lite
Stable Diffusion 1.5
Measured color drift
FLUX.2 · seed 1111
ΔE 13.7 off brand
#f4c430 → #fec689
ΔE 1.7 on brand
#f6a15c → #f9a766
FLUX.2 · seed 2222
ΔE 9.5 visible shift
#f4c430 → #fdc777
ΔE 3.1 close
#f6a15c → #fba858
FLUX.2 · seed 3333
ΔE 18.5 off brand
#f4c430 → #f9b888
ΔE 5.4 visible shift
#f6a15c → #f8a778
FLUX.2 [pro] · seed 1111
ΔE 16.3 off brand
#f4c430 → #fdd5a8
ΔE 10.8 off brand
#f6a15c → #fbc89a
FLUX.2 [pro] · seed 2222
ΔE 16.2 off brand
#f4c430 → #e6a76a
ΔE 4.5 close
#f6a15c → #e6a76a
FLUX.2 [pro] · seed 3333
ΔE 15.3 off brand
#f4c430 → #fdd6a7
ΔE 7.7 visible shift
#f6a15c → #d7a577
FLUX.2 [dev] Turbo · seed 1111
ΔE 15.9 off brand
#f4c430 → #fab679
ΔE 4.7 close
#f6a15c → #f3a978
FLUX.2 [dev] Turbo · seed 2222
ΔE 14.9 off brand
#f4c430 → #f8b777
ΔE 5.5 visible shift
#f6a15c → #e9a579
FLUX.2 [dev] Turbo · seed 3333
ΔE 15.0 off brand
#f4c430 → #fbd6a6
ΔE 4.6 close
#f6a15c → #f5a877
Gemini 2.5 Flash Image · sample 1
ΔE 14.6 off brand
#f4c430 → #fab979
ΔE 4.0 close
#f6a15c → #f4a774
Gemini 2.5 Flash Image · sample 2
ΔE 38.9 off brand
#f4c430 → #976657
ΔE 26.6 off brand
#f6a15c → #976657
Gemini 2.5 Flash Image · sample 3
ΔE 51.4 off brand
#f4c430 → #585459
ΔE 44.2 off brand
#f6a15c → #585459
GPT Image 1.5 · sample 1
ΔE 13.1 off brand
#f4c430 → #fcb668
ΔE 3.4 close
#f6a15c → #faa958
GPT Image 1.5 · sample 2
ΔE 13.7 off brand
#f4c430 → #f8b977
ΔE 3.6 close
#f6a15c → #e89658
GPT Image 1.5 · sample 3
ΔE 15.9 off brand
#f4c430 → #f9a858
ΔE 3.1 close
#f6a15c → #f9a858
Ideogram 3.0 · seed 1111
ΔE 2.7 close
#f4c430 → #ffc700
ΔE 15.0 off brand
#f6a15c → #f7d3ab
Ideogram 3.0 · seed 2222
ΔE 20.3 off brand
#f4c430 → #fee8ca
ΔE 20.6 off brand
#f6a15c → #fee8ca
Ideogram 3.0 · seed 3333
ΔE 15.0 off brand
#f4c430 → #fed9a7
ΔE 15.7 off brand
#f6a15c → #fed9a7
Luma Photon · seed 1111
ΔE 11.3 off brand
#f4c430 → #f8c886
ΔE 12.1 off brand
#f6a15c → #f8c886
Luma Photon · seed 2222
ΔE 12.5 off brand
#f4c430 → #f9d599
ΔE 15.6 off brand
#f6a15c → #f9d599
Luma Photon · seed 3333
ΔE 21.3 off brand
#f4c430 → #d6cbb6
ΔE 20.3 off brand
#f6a15c → #d6cbb6
Nano Banana 2 · sample 1
ΔE 25.0 off brand
#f4c430 → #b79985
ΔE 15.9 off brand
#f6a15c → #b79985
Nano Banana 2 · sample 2
ΔE 15.8 off brand
#f4c430 → #fdd6a8
ΔE 10.6 off brand
#f6a15c → #fcc799
Nano Banana 2 · sample 3
ΔE 16.5 off brand
#f4c430 → #fde7b7
ΔE 5.2 visible shift
#f6a15c → #e7a878
Nano Banana 2 Lite · sample 1
ΔE 12.2 off brand
#f4c430 → #f9c787
ΔE 7.2 visible shift
#f6a15c → #d9a577
Nano Banana 2 Lite · sample 2
ΔE 13.3 off brand
#f4c430 → #f7c58a
ΔE 10.4 off brand
#f6a15c → #f7c58a
Nano Banana 2 Lite · sample 3
ΔE 9.1 visible shift
#f4c430 → #fac879
ΔE 7.8 visible shift
#f6a15c → #d8a676
Nano Banana Pro · sample 1
ΔE 16.5 off brand
#f4c430 → #fdc493
ΔE 6.2 visible shift
#f6a15c → #e7a67c
Nano Banana Pro · sample 2
ΔE 32.5 off brand
#f4c430 → #b57b67
ΔE 18.3 off brand
#f6a15c → #b57b67
Nano Banana Pro · sample 3
ΔE 12.9 off brand
#f4c430 → #fec686
ΔE 2.1 close
#f6a15c → #f7a965
Qwen-Image · seed 1111
ΔE 16.0 off brand
#f4c430 → #f6c698
ΔE 11.1 off brand
#f6a15c → #f6c698
Qwen-Image · seed 2222
ΔE 15.8 off brand
#f4c430 → #f9d5a9
ΔE 15.2 off brand
#f6a15c → #f9d5a9
Qwen-Image · seed 3333
ΔE 23.1 off brand
#f4c430 → #e7d8c9
ΔE 21.1 off brand
#f6a15c → #e7d8c9
Recraft V3 · seed 1111
ΔE 17.1 off brand
#f4c430 → #fae6b9
ΔE 17.3 off brand
#f6a15c → #b8a789
Recraft V3 · seed 2222
ΔE 9.3 visible shift
#f4c430 → #f6c679
ΔE 12.5 off brand
#f6a15c → #f6c679
Recraft V3 · seed 3333
ΔE 29.1 off brand
#f4c430 → #a6aaa7
ΔE 26.2 off brand
#f6a15c → #a6aaa7
Seedream 4.5 · seed 1111
ΔE 12.8 off brand
#f4c430 → #f9d399
ΔE 6.7 visible shift
#f6a15c → #ea8739
Seedream 4.5 · seed 2222
ΔE 6.9 visible shift
#f4c430 → #fdc667
ΔE 6.0 visible shift
#f6a15c → #f7962b
Seedream 4.5 · seed 3333
ΔE 6.1 visible shift
#f4c430 → #f6b63a
ΔE 6.7 visible shift
#f6a15c → #e89734
Seedream 5.0 Lite · seed 1111
ΔE 15.0 off brand
#f4c430 → #fdd7a6
ΔE 5.8 visible shift
#f6a15c → #fbb778
Seedream 5.0 Lite · seed 2222
ΔE 8.0 visible shift
#f4c430 → #fec56a
ΔE 1.6 on brand
#f6a15c → #f59c58
Seedream 5.0 Lite · seed 3333
ΔE 12.7 off brand
#f4c430 → #feb664
ΔE 2.0 on brand
#f6a15c → #fca75c
Stable Diffusion 1.5 · seed 1111
ΔE 26.2 off brand
#f4c430 → #988668
ΔE 21.2 off brand
#f6a15c → #988668
Stable Diffusion 1.5 · seed 2222
ΔE 16.8 off brand
#f4c430 → #ffff17
ΔE 29.6 off brand
#f6a15c → #a899a6
Stable Diffusion 1.5 · seed 3333
ΔE 18.4 off brand
#f4c430 → #f6d9b7
ΔE 17.2 off brand
#f6a15c → #f6d9b7
ΔE2000 between each requested brand color and the nearest dominant swatch in the output. Under 2 is imperceptible; above 10 is a clearly different color.
Run-to-run consistency
Scores
| Prompt Adherence | Color Accuracy | Resolution | Latency | Cost | |
|---|---|---|---|---|---|
| FLUX.2 | 4.8 / 5 | 2.7 / 5 | 2.0 / 5 | 4.0 / 5 | 4.0 / 5 |
| FLUX.2 [pro] | 4.8 / 5 | 2.0 / 5 | 2.0 / 5 | 1.7 / 5 | 4.0 / 5 |
| FLUX.2 [dev] Turbo | 5.0 / 5 | 2.3 / 5 | 2.0 / 5 | 4.0 / 5 | 4.0 / 5 |
| Gemini 2.5 Flash Image | 4.8 / 5 | 1.0 / 5 | 3.0 / 5 | 3.0 / 5 | 4.0 / 5 |
| GPT Image 1.5 | 4.8 / 5 | 3.0 / 5 | 3.0 / 5 | 1.0 / 5 | 3.0 / 5 |
| Ideogram 3.0 | 4.3 / 5 | 1.7 / 5 | 3.0 / 5 | 2.0 / 5 | 3.0 / 5 |
| Luma Photon | 4.3 / 5 | 1.3 / 5 | 4.0 / 5 | 1.7 / 5 | 4.0 / 5 |
| Nano Banana 2 | 5.0 / 5 | 1.3 / 5 | 3.0 / 5 | 2.0 / 5 | 3.0 / 5 |
| Nano Banana 2 Lite | 5.0 / 5 | 2.7 / 5 | 3.0 / 5 | 4.0 / 5 | 4.0 / 5 |
| Nano Banana Pro | 4.8 / 5 | 1.7 / 5 | 3.0 / 5 | 1.0 / 5 | 1.0 / 5 |
| Qwen-Image | 4.8 / 5 | 1.3 / 5 | 2.0 / 5 | 3.0 / 5 | 4.0 / 5 |
| Recraft V3 | 4.3 / 5 | 1.3 / 5 | 3.0 / 5 | 2.7 / 5 | 4.0 / 5 |
| Seedream 4.5 | 4.5 / 5 | 3.0 / 5 | 5.0 / 5 | 1.3 / 5 | 4.0 / 5 |
| Seedream 5.0 Lite | 5.0 / 5 | 3.0 / 5 | 5.0 / 5 | 1.0 / 5 | 4.0 / 5 |
| Stable Diffusion 1.5 | 1.9 / 5 | 0.7 / 5 | 1.0 / 5 | 4.0 / 5 | 5.0 / 5 |
New model? New benchmarks.
We re-run the identical 37-prompt suite on every model release. Get the scores, and what changed in the rankings, in your inbox.
You're on the list.
We'll email you when new benchmark results ship.
Frequently asked questions
Image generation is stochastic: the same prompt produces different images on every run, so a single sample measures luck, not ability. Every model runs each benchmark three times with fixed seeds (1111, 2222, 3333), or three unseeded samples where the API accepts no seed. Scores average all three samples, and the consistency benchmarks measure the variation itself.
Yes. Every model receives the same prompt text with default parameters and no per-model tuning, so differences in output reflect the model, not prompt engineering. The suite is versioned and prompts are append-only, which keeps historical scores comparable.
The results on this page are scored on this exact prompt, so the strongest model for it is easy to spot; this benchmark sits in the Structured Spec Adherence category. Scores are per model version and reflect this benchmark only. For a ranking across every benchmark see the model rankings, and to weight the numbers by a specific job see the use-case pages.
Each image is scored 0 to 5 per criterion. Measured criteria (resolution, latency, cost, transparency, color accuracy) are computed automatically; quality criteria (text accuracy, prompt adherence, composition) are judged by an automated Claude vision tier against each prompt’s checklist. Page scores are unweighted means over all of a model’s runs in that scope. Blind expert-panel review has not run yet; the dataset is pilot-0.
Model output is a starting point, not a finished asset. Production work usually needs background removal, exact brand colors or editable text, none of which generation guarantees on every run. The IMG.LY AI Editor gives your users those controls to refine any model’s output to production quality.
The suite re-runs on notable model releases so the rankings stay current. The results shown are the pilot-0 dataset, scored by measured criteria plus an automated Claude vision tier, with blind expert review planned. Prompts are versioned and append-only, so scores stay comparable across runs.

