Benchmark prompts
Brand-Color Fidelity

Two-Color Illustration

This prompt names two exact brand colors by hex and asks the model to use only those. We compare the color each model actually rendered against the requested values with the CIEDE2000 distance, the same perceptual metric print shops use. Brand-color fidelity is non-negotiable for marketing and packaging work, and it is one of the few quality criteria we can measure automatically and objectively, which is why it anchors the pilot.

The exact prompt, sent to every model:
Flat illustration of a delivery scooter rider, using exactly two brand colors: #FF3B30 red and #1D1D1F near-black, on pure white background.
Suite pilot-0 · identical prompt and default parameters for every model · seeds 1111, 2222, 3333 (or repeated samples for seedless APIs)

Results

Click any image to enlarge. Checkerboard backing shows true transparency; a rendered checkerboard pattern inside the image is a model faking it.

FLUX.2

FLUX.2 [pro]

FLUX.2 [dev] Turbo

Gemini 2.5 Flash Image

GPT Image 1.5

Ideogram 3.0

Luma Photon

Nano Banana 2

Nano Banana 2 Lite

Nano Banana Pro

Qwen-Image

Recraft V3

Seedream 4.5

Seedream 5.0 Lite

Stable Diffusion 1.5

Measured color drift

The requested brand color next to the swatch each model actually rendered, with the CIEDE2000 distance between them.

FLUX.2 · seed 1111

ΔE 5.9 visible shift

#ff3b30 → #f90102

ΔE 3.7 close

#1d1d1f → #282827

FLUX.2 · seed 2222

ΔE 6.4 visible shift

#ff3b30 → #f60101

ΔE 3.7 close

#1d1d1f → #14141a

FLUX.2 · seed 3333

ΔE 5.8 visible shift

#ff3b30 → #fa0101

ΔE 6.5 visible shift

#1d1d1f → #050101

FLUX.2 [pro] · seed 1111

ΔE 5.5 visible shift

#ff3b30 → #fd0401

ΔE 6.3 visible shift

#1d1d1f → #020202

FLUX.2 [pro] · seed 2222

ΔE 5.7 visible shift

#ff3b30 → #fb0401

ΔE 5.9 visible shift

#1d1d1f → #030507

FLUX.2 [pro] · seed 3333

ΔE 3.8 close

#ff3b30 → #f7271a

ΔE 2.1 close

#1d1d1f → #171718

FLUX.2 [dev] Turbo · seed 1111

ΔE 7.7 visible shift

#ff3b30 → #e71319

ΔE 2.4 close

#1d1d1f → #15171a

FLUX.2 [dev] Turbo · seed 2222

ΔE 4.5 close

#ff3b30 → #f81516

ΔE 3.8 close

#1d1d1f → #25242a

FLUX.2 [dev] Turbo · seed 3333

ΔE 8.1 visible shift

#ff3b30 → #e70b15

ΔE 2.7 close

#1d1d1f → #16171b

Gemini 2.5 Flash Image · sample 1

ΔE 4.8 close

#ff3b30 → #f7534c

ΔE 9.6 visible shift

#1d1d1f → #182c37

Gemini 2.5 Flash Image · sample 2

ΔE 33.7 off brand

#ff3b30 → #a6abab

ΔE 8.3 visible shift

#1d1d1f → #1a292c

Gemini 2.5 Flash Image · sample 3

ΔE 7.3 visible shift

#ff3b30 → #f8625a

ΔE 10.1 off brand

#1d1d1f → #26353c

GPT Image 1.5 · sample 1

ΔE 4.2 close

#ff3b30 → #fd3414

ΔE 2.4 close

#1d1d1f → #1b1d22

GPT Image 1.5 · sample 2

ΔE 1.7 on brand

#ff3b30 → #fc3d38

ΔE 7.3 visible shift

#1d1d1f → #232c37

GPT Image 1.5 · sample 3

ΔE 3.8 close

#ff3b30 → #fc2615

ΔE 5.3 visible shift

#1d1d1f → #24232c

Ideogram 3.0 · seed 1111

ΔE 3.9 close

#ff3b30 → #fd0622

ΔE 3.3 close

#1d1d1f → #262729

Ideogram 3.0 · seed 2222

ΔE 5.3 visible shift

#ff3b30 → #fd0303

ΔE 2.5 close

#1d1d1f → #171616

Ideogram 3.0 · seed 3333

ΔE 4.2 close

#ff3b30 → #fb1516

ΔE 2.0 close

#1d1d1f → #191919

Luma Photon · seed 1111

ΔE 3.7 close

#ff3b30 → #fd1c18

ΔE 2.4 close

#1d1d1f → #171717

Luma Photon · seed 2222

ΔE 2.3 close

#ff3b30 → #fb2a2a

ΔE 2.6 close

#1d1d1f → #161717

Luma Photon · seed 3333

ΔE 1.4 on brand

#ff3b30 → #fd4639

ΔE 6.7 visible shift

#1d1d1f → #282b36

Nano Banana 2 · sample 1

ΔE 1.0 on brand

#ff3b30 → #fb4435

ΔE 2.3 close

#1d1d1f → #252426

Nano Banana 2 · sample 2

ΔE 1.3 on brand

#ff3b30 → #fd3a29

ΔE 2.1 close

#1d1d1f → #161618

Nano Banana 2 · sample 3

ΔE 1.6 on brand

#ff3b30 → #f93c2a

ΔE 2.3 close

#1d1d1f → #161619

Nano Banana 2 Lite · sample 1

ΔE 1.3 on brand

#ff3b30 → #f93c33

ΔE 2.1 close

#1d1d1f → #161619

Nano Banana 2 Lite · sample 2

ΔE 30.2 off brand

#ff3b30 → #959596

ΔE 1.9 on brand

#1d1d1f → #181719

Nano Banana 2 Lite · sample 3

ΔE 1.2 on brand

#ff3b30 → #fe4138

ΔE 3.2 close

#1d1d1f → #191d23

Nano Banana Pro · sample 1

ΔE 0.8 on brand

#ff3b30 → #fd4333

ΔE 2.5 close

#1d1d1f → #252326

Nano Banana Pro · sample 2

ΔE 1.5 on brand

#ff3b30 → #fb3728

ΔE 1.9 on brand

#1d1d1f → #171719

Nano Banana Pro · sample 3

ΔE 1.2 on brand

#ff3b30 → #fc412e

ΔE 2.0 close

#1d1d1f → #161719

Qwen-Image · seed 1111

ΔE 1.7 on brand

#ff3b30 → #fb3228

ΔE 2.9 close

#1d1d1f → #151615

Qwen-Image · seed 2222

ΔE 4.5 close

#ff3b30 → #fc4948

ΔE 5.8 visible shift

#1d1d1f → #2c2d2c

Qwen-Image · seed 3333

ΔE 1.9 on brand

#ff3b30 → #fe453c

ΔE 1.4 on brand

#1d1d1f → #19191b

Recraft V3 · seed 1111

ΔE 8.6 visible shift

#ff3b30 → #e71a34

ΔE 4.0 close

#1d1d1f → #1a1515

Recraft V3 · seed 2222

ΔE 29.4 off brand

#ff3b30 → #7b7676

ΔE 2.9 close

#1d1d1f → #191716

Recraft V3 · seed 3333

ΔE 34.6 off brand

#ff3b30 → #cbc7c4

ΔE 4.1 close

#1d1d1f → #1b1615

Seedream 4.5 · seed 1111

ΔE 3.3 close

#ff3b30 → #f6292c

ΔE 1.9 on brand

#1d1d1f → #181719

Seedream 4.5 · seed 2222

ΔE 2.5 close

#ff3b30 → #fa3233

ΔE 2.2 close

#1d1d1f → #161619

Seedream 4.5 · seed 3333

ΔE 3.2 close

#ff3b30 → #f52b2b

ΔE 2.6 close

#1d1d1f → #151516

Seedream 5.0 Lite · seed 1111

ΔE 3.0 close

#ff3b30 → #fd1b25

ΔE 6.1 visible shift

#1d1d1f → #131b26

Seedream 5.0 Lite · seed 2222

ΔE 31.1 off brand

#ff3b30 → #96979a

ΔE 4.0 close

#1d1d1f → #23242b

Seedream 5.0 Lite · seed 3333

ΔE 3.8 close

#ff3b30 → #fc141b

ΔE 6.7 visible shift

#1d1d1f → #131d29

Stable Diffusion 1.5 · seed 1111

ΔE 5.4 visible shift

#ff3b30 → #fd0103

ΔE 2.6 close

#1d1d1f → #161516

Stable Diffusion 1.5 · seed 2222

ΔE 14.1 off brand

#ff3b30 → #cc0116

ΔE 6.3 visible shift

#1d1d1f → #030303

Stable Diffusion 1.5 · seed 3333

ΔE 7.5 visible shift

#ff3b30 → #e43139

ΔE 6.3 visible shift

#1d1d1f → #070c17

ΔE2000 between each requested brand color and the nearest dominant swatch in the output. Under 2 is imperceptible; above 10 is a clearly different color.

Run-to-run consistency

Same model, same prompt, different seed or sample. Click a card to flip between the two generations: the less it moves, the safer the model is for templated, variable-data production.

Scores

Mean per model across this prompt's runs. Auto-measured criteria plus an interim Claude-VLM tier for prompt adherence and composition; any criterion still awaiting review is labeled pending, never zeroed.
  Prompt Adherence Color Accuracy Resolution Latency Cost
FLUX.2 3.8 / 5 3.3 / 5 2.0 / 5 5.0 / 5 4.0 / 5
FLUX.2 [pro] 4.6 / 5 3.3 / 5 2.0 / 5 2.3 / 5 4.0 / 5
FLUX.2 [dev] Turbo 3.8 / 5 3.3 / 5 2.0 / 5 4.3 / 5 4.0 / 5
Gemini 2.5 Flash Image 5.0 / 5 2.0 / 5 3.0 / 5 3.0 / 5 4.0 / 5
GPT Image 1.5 5.0 / 5 4.0 / 5 3.0 / 5 1.0 / 5 3.0 / 5
Ideogram 3.0 4.2 / 5 4.0 / 5 3.0 / 5 2.0 / 5 3.0 / 5
Luma Photon 3.8 / 5 4.0 / 5 4.0 / 5 2.0 / 5 4.0 / 5
Nano Banana 2 4.6 / 5 5.0 / 5 3.0 / 5 2.0 / 5 3.0 / 5
Nano Banana 2 Lite 5.0 / 5 3.7 / 5 3.0 / 5 3.7 / 5 4.0 / 5
Nano Banana Pro 5.0 / 5 5.0 / 5 3.0 / 5 1.3 / 5 1.0 / 5
Qwen-Image 2.5 / 5 4.0 / 5 2.0 / 5 3.0 / 5 4.0 / 5
Recraft V3 2.5 / 5 2.3 / 5 3.0 / 5 3.0 / 5 4.0 / 5
Seedream 4.5 4.2 / 5 4.0 / 5 5.0 / 5 1.7 / 5 4.0 / 5
Seedream 5.0 Lite 4.6 / 5 3.0 / 5 5.0 / 5 1.0 / 5 4.0 / 5
Stable Diffusion 1.5 2.5 / 5 3.0 / 5 1.0 / 5 4.3 / 5 5.0 / 5

Frequently asked questions

Image generation is stochastic: the same prompt produces different images on every run, so a single sample measures luck, not ability. Every model runs each benchmark three times with fixed seeds (1111, 2222, 3333), or three unseeded samples where the API accepts no seed. Scores average all three samples, and the consistency benchmarks measure the variation itself.

Yes. Every model receives the same prompt text with default parameters and no per-model tuning, so differences in output reflect the model, not prompt engineering. The suite is versioned and prompts are append-only, which keeps historical scores comparable.

The results on this page are scored on this exact prompt, so the strongest model for it is easy to spot; this benchmark sits in the Brand-Color Fidelity category. Scores are per model version and reflect this benchmark only. For a ranking across every benchmark see the model rankings, and to weight the numbers by a specific job see the use-case pages.

Each image is scored 0 to 5 per criterion. Measured criteria (resolution, latency, cost, transparency, color accuracy) are computed automatically; quality criteria (text accuracy, prompt adherence, composition) are judged by an automated Claude vision tier against each prompt’s checklist. Page scores are unweighted means over all of a model’s runs in that scope. Blind expert-panel review has not run yet; the dataset is pilot-0.

Model output is a starting point, not a finished asset. Production work usually needs background removal, exact brand colors or editable text, none of which generation guarantees on every run. The IMG.LY AI Editor gives your users those controls to refine any model’s output to production quality.

The suite re-runs on notable model releases so the rankings stay current. The results shown are the pilot-0 dataset, scored by measured criteria plus an automated Claude vision tier, with blind expert review planned. Prompts are versioned and append-only, so scores stay comparable across runs.