Benchmark prompts
Structured Spec Adherence

Structured JSON Scene

Structured JSON prompts are how creative-automation pipelines, agents and CE.SDK integrations actually drive image models, and nobody benchmarks it. Each spec field becomes a checklist item plus an automated color check, the modality closest to IMG.LY's product reality.

The exact prompt, sent to every model:
Render the following scene exactly: {"subject":"a lighthouse keeper","gaze":"toward the sea","wardrobe":"yellow raincoat","lighting":"golden hour from the left","composition":"rule of thirds, subject on the right","colorPlate":[{"name":"coat","hex":"#F4C430","target_pct":15},{"name":"sky","hex":"#F6A15C","target_pct":40}]}
Suite pilot-0 · identical prompt and default parameters for every model · seeds 1111, 2222, 3333 (or repeated samples for seedless APIs)

Results

Click any image to enlarge. Checkerboard backing shows true transparency; a rendered checkerboard pattern inside the image is a model faking it.

FLUX.2

FLUX.2 [pro]

FLUX.2 [dev] Turbo

Gemini 2.5 Flash Image

GPT Image 1.5

Ideogram 3.0

Luma Photon

Nano Banana 2

Nano Banana 2 Lite

Nano Banana Pro

Qwen-Image

Recraft V3

Seedream 4.5

Seedream 5.0 Lite

Stable Diffusion 1.5

Measured color drift

The requested brand color next to the swatch each model actually rendered, with the CIEDE2000 distance between them.

FLUX.2 · seed 1111

ΔE 13.7 off brand

#f4c430 → #fec689

ΔE 1.7 on brand

#f6a15c → #f9a766

FLUX.2 · seed 2222

ΔE 9.5 visible shift

#f4c430 → #fdc777

ΔE 3.1 close

#f6a15c → #fba858

FLUX.2 · seed 3333

ΔE 18.5 off brand

#f4c430 → #f9b888

ΔE 5.4 visible shift

#f6a15c → #f8a778

FLUX.2 [pro] · seed 1111

ΔE 16.3 off brand

#f4c430 → #fdd5a8

ΔE 10.8 off brand

#f6a15c → #fbc89a

FLUX.2 [pro] · seed 2222

ΔE 16.2 off brand

#f4c430 → #e6a76a

ΔE 4.5 close

#f6a15c → #e6a76a

FLUX.2 [pro] · seed 3333

ΔE 15.3 off brand

#f4c430 → #fdd6a7

ΔE 7.7 visible shift

#f6a15c → #d7a577

FLUX.2 [dev] Turbo · seed 1111

ΔE 15.9 off brand

#f4c430 → #fab679

ΔE 4.7 close

#f6a15c → #f3a978

FLUX.2 [dev] Turbo · seed 2222

ΔE 14.9 off brand

#f4c430 → #f8b777

ΔE 5.5 visible shift

#f6a15c → #e9a579

FLUX.2 [dev] Turbo · seed 3333

ΔE 15.0 off brand

#f4c430 → #fbd6a6

ΔE 4.6 close

#f6a15c → #f5a877

Gemini 2.5 Flash Image · sample 1

ΔE 14.6 off brand

#f4c430 → #fab979

ΔE 4.0 close

#f6a15c → #f4a774

Gemini 2.5 Flash Image · sample 2

ΔE 38.9 off brand

#f4c430 → #976657

ΔE 26.6 off brand

#f6a15c → #976657

Gemini 2.5 Flash Image · sample 3

ΔE 51.4 off brand

#f4c430 → #585459

ΔE 44.2 off brand

#f6a15c → #585459

GPT Image 1.5 · sample 1

ΔE 13.1 off brand

#f4c430 → #fcb668

ΔE 3.4 close

#f6a15c → #faa958

GPT Image 1.5 · sample 2

ΔE 13.7 off brand

#f4c430 → #f8b977

ΔE 3.6 close

#f6a15c → #e89658

GPT Image 1.5 · sample 3

ΔE 15.9 off brand

#f4c430 → #f9a858

ΔE 3.1 close

#f6a15c → #f9a858

Ideogram 3.0 · seed 1111

ΔE 2.7 close

#f4c430 → #ffc700

ΔE 15.0 off brand

#f6a15c → #f7d3ab

Ideogram 3.0 · seed 2222

ΔE 20.3 off brand

#f4c430 → #fee8ca

ΔE 20.6 off brand

#f6a15c → #fee8ca

Ideogram 3.0 · seed 3333

ΔE 15.0 off brand

#f4c430 → #fed9a7

ΔE 15.7 off brand

#f6a15c → #fed9a7

Luma Photon · seed 1111

ΔE 11.3 off brand

#f4c430 → #f8c886

ΔE 12.1 off brand

#f6a15c → #f8c886

Luma Photon · seed 2222

ΔE 12.5 off brand

#f4c430 → #f9d599

ΔE 15.6 off brand

#f6a15c → #f9d599

Luma Photon · seed 3333

ΔE 21.3 off brand

#f4c430 → #d6cbb6

ΔE 20.3 off brand

#f6a15c → #d6cbb6

Nano Banana 2 · sample 1

ΔE 25.0 off brand

#f4c430 → #b79985

ΔE 15.9 off brand

#f6a15c → #b79985

Nano Banana 2 · sample 2

ΔE 15.8 off brand

#f4c430 → #fdd6a8

ΔE 10.6 off brand

#f6a15c → #fcc799

Nano Banana 2 · sample 3

ΔE 16.5 off brand

#f4c430 → #fde7b7

ΔE 5.2 visible shift

#f6a15c → #e7a878

Nano Banana 2 Lite · sample 1

ΔE 12.2 off brand

#f4c430 → #f9c787

ΔE 7.2 visible shift

#f6a15c → #d9a577

Nano Banana 2 Lite · sample 2

ΔE 13.3 off brand

#f4c430 → #f7c58a

ΔE 10.4 off brand

#f6a15c → #f7c58a

Nano Banana 2 Lite · sample 3

ΔE 9.1 visible shift

#f4c430 → #fac879

ΔE 7.8 visible shift

#f6a15c → #d8a676

Nano Banana Pro · sample 1

ΔE 16.5 off brand

#f4c430 → #fdc493

ΔE 6.2 visible shift

#f6a15c → #e7a67c

Nano Banana Pro · sample 2

ΔE 32.5 off brand

#f4c430 → #b57b67

ΔE 18.3 off brand

#f6a15c → #b57b67

Nano Banana Pro · sample 3

ΔE 12.9 off brand

#f4c430 → #fec686

ΔE 2.1 close

#f6a15c → #f7a965

Qwen-Image · seed 1111

ΔE 16.0 off brand

#f4c430 → #f6c698

ΔE 11.1 off brand

#f6a15c → #f6c698

Qwen-Image · seed 2222

ΔE 15.8 off brand

#f4c430 → #f9d5a9

ΔE 15.2 off brand

#f6a15c → #f9d5a9

Qwen-Image · seed 3333

ΔE 23.1 off brand

#f4c430 → #e7d8c9

ΔE 21.1 off brand

#f6a15c → #e7d8c9

Recraft V3 · seed 1111

ΔE 17.1 off brand

#f4c430 → #fae6b9

ΔE 17.3 off brand

#f6a15c → #b8a789

Recraft V3 · seed 2222

ΔE 9.3 visible shift

#f4c430 → #f6c679

ΔE 12.5 off brand

#f6a15c → #f6c679

Recraft V3 · seed 3333

ΔE 29.1 off brand

#f4c430 → #a6aaa7

ΔE 26.2 off brand

#f6a15c → #a6aaa7

Seedream 4.5 · seed 1111

ΔE 12.8 off brand

#f4c430 → #f9d399

ΔE 6.7 visible shift

#f6a15c → #ea8739

Seedream 4.5 · seed 2222

ΔE 6.9 visible shift

#f4c430 → #fdc667

ΔE 6.0 visible shift

#f6a15c → #f7962b

Seedream 4.5 · seed 3333

ΔE 6.1 visible shift

#f4c430 → #f6b63a

ΔE 6.7 visible shift

#f6a15c → #e89734

Seedream 5.0 Lite · seed 1111

ΔE 15.0 off brand

#f4c430 → #fdd7a6

ΔE 5.8 visible shift

#f6a15c → #fbb778

Seedream 5.0 Lite · seed 2222

ΔE 8.0 visible shift

#f4c430 → #fec56a

ΔE 1.6 on brand

#f6a15c → #f59c58

Seedream 5.0 Lite · seed 3333

ΔE 12.7 off brand

#f4c430 → #feb664

ΔE 2.0 on brand

#f6a15c → #fca75c

Stable Diffusion 1.5 · seed 1111

ΔE 26.2 off brand

#f4c430 → #988668

ΔE 21.2 off brand

#f6a15c → #988668

Stable Diffusion 1.5 · seed 2222

ΔE 16.8 off brand

#f4c430 → #ffff17

ΔE 29.6 off brand

#f6a15c → #a899a6

Stable Diffusion 1.5 · seed 3333

ΔE 18.4 off brand

#f4c430 → #f6d9b7

ΔE 17.2 off brand

#f6a15c → #f6d9b7

ΔE2000 between each requested brand color and the nearest dominant swatch in the output. Under 2 is imperceptible; above 10 is a clearly different color.

Run-to-run consistency

Same model, same prompt, different seed or sample. Click a card to flip between the two generations: the less it moves, the safer the model is for templated, variable-data production.

Scores

Mean per model across this prompt's runs. Auto-measured criteria plus an interim Claude-VLM tier for prompt adherence and composition; any criterion still awaiting review is labeled pending, never zeroed.
  Prompt Adherence Color Accuracy Resolution Latency Cost
FLUX.2 4.8 / 5 2.7 / 5 2.0 / 5 4.0 / 5 4.0 / 5
FLUX.2 [pro] 4.8 / 5 2.0 / 5 2.0 / 5 1.7 / 5 4.0 / 5
FLUX.2 [dev] Turbo 5.0 / 5 2.3 / 5 2.0 / 5 4.0 / 5 4.0 / 5
Gemini 2.5 Flash Image 4.8 / 5 1.0 / 5 3.0 / 5 3.0 / 5 4.0 / 5
GPT Image 1.5 4.8 / 5 3.0 / 5 3.0 / 5 1.0 / 5 3.0 / 5
Ideogram 3.0 4.3 / 5 1.7 / 5 3.0 / 5 2.0 / 5 3.0 / 5
Luma Photon 4.3 / 5 1.3 / 5 4.0 / 5 1.7 / 5 4.0 / 5
Nano Banana 2 5.0 / 5 1.3 / 5 3.0 / 5 2.0 / 5 3.0 / 5
Nano Banana 2 Lite 5.0 / 5 2.7 / 5 3.0 / 5 4.0 / 5 4.0 / 5
Nano Banana Pro 4.8 / 5 1.7 / 5 3.0 / 5 1.0 / 5 1.0 / 5
Qwen-Image 4.8 / 5 1.3 / 5 2.0 / 5 3.0 / 5 4.0 / 5
Recraft V3 4.3 / 5 1.3 / 5 3.0 / 5 2.7 / 5 4.0 / 5
Seedream 4.5 4.5 / 5 3.0 / 5 5.0 / 5 1.3 / 5 4.0 / 5
Seedream 5.0 Lite 5.0 / 5 3.0 / 5 5.0 / 5 1.0 / 5 4.0 / 5
Stable Diffusion 1.5 1.9 / 5 0.7 / 5 1.0 / 5 4.0 / 5 5.0 / 5

Frequently asked questions

Image generation is stochastic: the same prompt produces different images on every run, so a single sample measures luck, not ability. Every model runs each benchmark three times with fixed seeds (1111, 2222, 3333), or three unseeded samples where the API accepts no seed. Scores average all three samples, and the consistency benchmarks measure the variation itself.

Yes. Every model receives the same prompt text with default parameters and no per-model tuning, so differences in output reflect the model, not prompt engineering. The suite is versioned and prompts are append-only, which keeps historical scores comparable.

The results on this page are scored on this exact prompt, so the strongest model for it is easy to spot; this benchmark sits in the Structured Spec Adherence category. Scores are per model version and reflect this benchmark only. For a ranking across every benchmark see the model rankings, and to weight the numbers by a specific job see the use-case pages.

Each image is scored 0 to 5 per criterion. Measured criteria (resolution, latency, cost, transparency, color accuracy) are computed automatically; quality criteria (text accuracy, prompt adherence, composition) are judged by an automated Claude vision tier against each prompt’s checklist. Page scores are unweighted means over all of a model’s runs in that scope. Blind expert-panel review has not run yet; the dataset is pilot-0.

Model output is a starting point, not a finished asset. Production work usually needs background removal, exact brand colors or editable text, none of which generation guarantees on every run. The IMG.LY AI Editor gives your users those controls to refine any model’s output to production quality.

The suite re-runs on notable model releases so the rankings stay current. The results shown are the pilot-0 dataset, scored by measured criteria plus an automated Claude vision tier, with blind expert review planned. Prompts are versioned and append-only, so scores stay comparable across runs.