Benchmark prompts
Brand-Color Fidelity

Duotone Portrait

Duotone is a brand-styling staple for posters and editorial. Naming both tones by hex makes color fidelity measurable with the same CIEDE2000 pipeline, testing whether a model holds an exact two-color mapping.

Use cases:
Marketing & ad creative
Template & asset packs
Core set
The exact prompt, sent to every model:
Duotone portrait of a saxophone player, highlights in #FDCB6E and shadows in #2D3436, poster style.
Suite pilot-0 · identical prompt and default parameters for every model · seeds 1111, 2222, 3333 (or repeated samples for seedless APIs)

Results

Click any image to enlarge. Checkerboard backing shows true transparency; a rendered checkerboard pattern inside the image is a model faking it.

FLUX.2

FLUX.2 [pro]

FLUX.2 [dev] Turbo

Gemini 2.5 Flash Image

GPT Image 1.5

Ideogram 3.0

Luma Photon

Nano Banana 2

Nano Banana 2 Lite

Nano Banana Pro

Qwen-Image

Recraft V3

Seedream 4.5

Seedream 5.0 Lite

Stable Diffusion 1.5

Measured color drift

The requested brand color next to the swatch each model actually rendered, with the CIEDE2000 distance between them.

FLUX.2 · seed 1111

ΔE 17.8 off brand

#fdcb6e → #b98d0e

ΔE 5.9 visible shift

#2d3436 → #272625

FLUX.2 · seed 2222

ΔE 20.7 off brand

#fdcb6e → #a78809

ΔE 5.6 visible shift

#2d3436 → #252b26

FLUX.2 · seed 3333

ΔE 7.2 visible shift

#fdcb6e → #ffc621

ΔE 9.7 visible shift

#2d3436 → #0b2322

FLUX.2 [pro] · seed 1111

ΔE 29.2 off brand

#fdcb6e → #916d29

ΔE 2.7 close

#2d3436 → #333a3e

FLUX.2 [pro] · seed 2222

ΔE 45.3 off brand

#fdcb6e → #674b04

ΔE 7.5 visible shift

#2d3436 → #172522

FLUX.2 [pro] · seed 3333

ΔE 7.8 visible shift

#fdcb6e → #ffc40c

ΔE 6.1 visible shift

#2d3436 → #1b2121

FLUX.2 [dev] Turbo · seed 1111

ΔE 11.7 off brand

#fdcb6e → #ffebb7

ΔE 7.1 visible shift

#2d3436 → #282625

FLUX.2 [dev] Turbo · seed 2222

ΔE 21.6 off brand

#fdcb6e → #fffeea

ΔE 6.7 visible shift

#2d3436 → #282624

FLUX.2 [dev] Turbo · seed 3333

ΔE 17.9 off brand

#fdcb6e → #fffad7

ΔE 6.0 visible shift

#2d3436 → #383534

Gemini 2.5 Flash Image · sample 1

ΔE 38.4 off brand

#fdcb6e → #745c39

ΔE 4.4 close

#2d3436 → #343433

Gemini 2.5 Flash Image · sample 2

ΔE 4.5 close

#fdcb6e → #f4b954

ΔE 10.9 off brand

#2d3436 → #131311

Gemini 2.5 Flash Image · sample 3

ΔE 12.0 off brand

#fdcb6e → #e8d5ae

ΔE 4.7 close

#2d3436 → #3b3c3b

GPT Image 1.5 · sample 1

ΔE 6.9 visible shift

#fdcb6e → #ffdb61

ΔE 4.3 close

#2d3436 → #353c39

GPT Image 1.5 · sample 2

ΔE 7.6 visible shift

#fdcb6e → #fbcb29

ΔE 3.9 close

#2d3436 → #273139

GPT Image 1.5 · sample 3

ΔE 5.4 visible shift

#fdcb6e → #fac942

ΔE 4.0 close

#2d3436 → #353b42

Ideogram 3.0 · seed 1111

ΔE 8.8 visible shift

#fdcb6e → #fce48c

ΔE 7.8 visible shift

#2d3436 → #292627

Ideogram 3.0 · seed 2222

ΔE 5.5 visible shift

#fdcb6e → #f5b843

ΔE 13.3 off brand

#2d3436 → #1a1309

Ideogram 3.0 · seed 3333

ΔE 21.0 off brand

#fdcb6e → #fcf8e7

ΔE 2.5 close

#2d3436 → #2c3339

Luma Photon · seed 1111

ΔE 47.2 off brand

#fdcb6e → #585449

ΔE 4.7 close

#2d3436 → #323e46

Luma Photon · seed 2222

ΔE 11.0 off brand

#fdcb6e → #f4a429

ΔE 6.3 visible shift

#2d3436 → #2a2825

Luma Photon · seed 3333

ΔE 47.3 off brand

#fdcb6e → #595247

ΔE 4.7 close

#2d3436 → #393837

Nano Banana 2 · sample 1

ΔE 14.0 off brand

#fdcb6e → #c49849

ΔE 4.7 close

#2d3436 → #252926

Nano Banana 2 · sample 2

ΔE 1.5 on brand

#fdcb6e → #fcc664

ΔE 4.4 close

#2d3436 → #262a27

Nano Banana 2 · sample 3

ΔE 15.8 off brand

#fdcb6e → #b89548

ΔE 2.9 close

#2d3436 → #262c2c

Nano Banana 2 Lite · sample 1

ΔE 25.0 off brand

#fdcb6e → #e7e7e8

ΔE 1.1 on brand

#2d3436 → #2c3233

Nano Banana 2 Lite · sample 2

ΔE 24.6 off brand

#fdcb6e → #e9e8e8

ΔE 3.3 close

#2d3436 → #242a2b

Nano Banana 2 Lite · sample 3

ΔE 16.3 off brand

#fdcb6e → #b99458

ΔE 0.8 on brand

#2d3436 → #2c3233

Nano Banana Pro · sample 1

ΔE 2.4 close

#fdcb6e → #fac55a

ΔE 0.3 on brand

#2d3436 → #2c3335

Nano Banana Pro · sample 2

ΔE 2.8 close

#fdcb6e → #f7c357

ΔE 0.9 on brand

#2d3436 → #2c3334

Nano Banana Pro · sample 3

ΔE 15.8 off brand

#fdcb6e → #ba9556

ΔE 1.2 on brand

#2d3436 → #2a3234

Qwen-Image · seed 1111

ΔE 51.9 off brand

#fdcb6e → #d73668

ΔE 2.1 close

#2d3436 → #2e3235

Qwen-Image · seed 2222

ΔE 50.0 off brand

#fdcb6e → #fd417c

ΔE 9.2 visible shift

#2d3436 → #191918

Qwen-Image · seed 3333

ΔE 21.6 off brand

#fdcb6e → #efe8e0

ΔE 11.2 off brand

#2d3436 → #0e100e

Recraft V3 · seed 1111

ΔE 16.7 off brand

#fdcb6e → #d7cab5

ΔE 8.6 visible shift

#2d3436 → #173847

Recraft V3 · seed 2222

ΔE 20.3 off brand

#fdcb6e → #b5a998

ΔE 3.9 close

#2d3436 → #373839

Recraft V3 · seed 3333

ΔE 18.1 off brand

#fdcb6e → #f6e9d6

ΔE 10.9 off brand

#2d3436 → #072539

Seedream 4.5 · seed 1111

ΔE 11.6 off brand

#fdcb6e → #faeba5

ΔE 10.8 off brand

#2d3436 → #041918

Seedream 4.5 · seed 2222

ΔE 55.3 off brand

#fdcb6e → #49433b

ΔE 6.6 visible shift

#2d3436 → #282724

Seedream 4.5 · seed 3333

ΔE 10.1 off brand

#fdcb6e → #e4b888

ΔE 4.4 close

#2d3436 → #283839

Seedream 5.0 Lite · seed 1111

ΔE 30.5 off brand

#fdcb6e → #976725

ΔE 12.6 off brand

#2d3436 → #19130c

Seedream 5.0 Lite · seed 2222

ΔE 32.4 off brand

#fdcb6e → #886735

ΔE 5.3 visible shift

#2d3436 → #272825

Seedream 5.0 Lite · seed 3333

ΔE 13.5 off brand

#fdcb6e → #e79819

ΔE 4.2 close

#2d3436 → #262a28

Stable Diffusion 1.5 · seed 1111

ΔE 6.3 visible shift

#fdcb6e → #f3be31

ΔE 11.7 off brand

#2d3436 → #17140d

Stable Diffusion 1.5 · seed 2222

ΔE 12.5 off brand

#fdcb6e → #fcc298

ΔE 4.2 close

#2d3436 → #373737

Stable Diffusion 1.5 · seed 3333

ΔE 4.9 close

#fdcb6e → #f5b94c

ΔE 12.3 off brand

#2d3436 → #19140c

ΔE2000 between each requested brand color and the nearest dominant swatch in the output. Under 2 is imperceptible; above 10 is a clearly different color.

Run-to-run consistency

Same model, same prompt, different seed or sample. Click a card to flip between the two generations: the less it moves, the safer the model is for templated, variable-data production.

Scores

Mean per model across this prompt's runs. Auto-measured criteria plus an interim Claude-VLM tier for prompt adherence and composition; any criterion still awaiting review is labeled pending, never zeroed.
  Prompt Adherence Color Accuracy Resolution Latency Cost
FLUX.2 4.2 / 5 2.3 / 5 2.0 / 5 4.0 / 5 4.0 / 5
FLUX.2 [pro] 4.6 / 5 1.7 / 5 2.0 / 5 2.7 / 5 4.0 / 5
FLUX.2 [dev] Turbo 4.6 / 5 2.3 / 5 2.0 / 5 4.0 / 5 4.0 / 5
Gemini 2.5 Flash Image 4.6 / 5 2.0 / 5 3.0 / 5 3.0 / 5 4.0 / 5
GPT Image 1.5 4.6 / 5 3.3 / 5 3.0 / 5 1.0 / 5 3.0 / 5
Ideogram 3.0 3.3 / 5 2.7 / 5 3.0 / 5 2.0 / 5 3.0 / 5
Luma Photon 4.6 / 5 1.0 / 5 4.0 / 5 1.7 / 5 4.0 / 5
Nano Banana 2 5.0 / 5 3.3 / 5 3.0 / 5 2.0 / 5 3.0 / 5
Nano Banana 2 Lite 4.2 / 5 2.3 / 5 3.0 / 5 4.0 / 5 4.0 / 5
Nano Banana Pro 5.0 / 5 4.3 / 5 3.0 / 5 1.0 / 5 1.0 / 5
Qwen-Image 1.3 / 5 0.7 / 5 2.0 / 5 3.0 / 5 4.0 / 5
Recraft V3 2.5 / 5 2.0 / 5 3.0 / 5 3.0 / 5 4.0 / 5
Seedream 4.5 3.8 / 5 1.7 / 5 5.0 / 5 1.7 / 5 4.0 / 5
Seedream 5.0 Lite 4.6 / 5 1.7 / 5 5.0 / 5 1.0 / 5 4.0 / 5
Stable Diffusion 1.5 4.2 / 5 3.0 / 5 1.0 / 5 4.0 / 5 5.0 / 5

Frequently asked questions

Image generation is stochastic: the same prompt produces different images on every run, so a single sample measures luck, not ability. Every model runs each benchmark three times with fixed seeds (1111, 2222, 3333), or three unseeded samples where the API accepts no seed. Scores average all three samples, and the consistency benchmarks measure the variation itself.

Yes. Every model receives the same prompt text with default parameters and no per-model tuning, so differences in output reflect the model, not prompt engineering. The suite is versioned and prompts are append-only, which keeps historical scores comparable.

The results on this page are scored on this exact prompt, so the strongest model for it is easy to spot; this benchmark sits in the Brand-Color Fidelity category. Scores are per model version and reflect this benchmark only. For a ranking across every benchmark see the model rankings, and to weight the numbers by a specific job see the use-case pages.

Each image is scored 0 to 5 per criterion. Measured criteria (resolution, latency, cost, transparency, color accuracy) are computed automatically; quality criteria (text accuracy, prompt adherence, composition) are judged by an automated Claude vision tier against each prompt’s checklist. Page scores are unweighted means over all of a model’s runs in that scope. Blind expert-panel review has not run yet; the dataset is pilot-0.

Model output is a starting point, not a finished asset. Production work usually needs background removal, exact brand colors or editable text, none of which generation guarantees on every run. The IMG.LY AI Editor gives your users those controls to refine any model’s output to production quality.

The suite re-runs on notable model releases so the rankings stay current. The results shown are the pilot-0 dataset, scored by measured criteria plus an automated Claude vision tier, with blind expert review planned. Prompts are versioned and append-only, so scores stay comparable across runs.