GenAI Benchmarks

Why AI image models miss your brand colors

We gave 15 models prompts with exact hex requirements and measured the returned pixels with CIEDE2000, the perceptual color-difference standard. Nobody passes cleanly: the best model reaches 3.67 of 5, 12 of 15 score below 3, and the field averages 2.6. Close-enough color is easy; the color in your brand book is not.

The benchmark avocado mascot as a painter holding a green swatch fan, frowning at an off-color portrait of itself
Scores

What every model scored

Mean color-fidelity score over 4 brand-color benchmarks, three seeds each. Scores band the CIEDE2000 distance between the required hex and the nearest dominant color actually rendered; the raw per-hex deltas are recorded with every score.
Color-fidelity criterion, full 15-model run (pilot-0).
  Color fidelity Runs scored $ / image
1. Nano Banana Pro
Closest to spec
3.67 / 5 9 $0.360
2. GPT Image 1.5 3.44 / 5 9 $0.120
3. Nano Banana 2 3.22 / 5 9 $0.100
4. Nano Banana 2 Lite 2.89 / 5 9 $0.020
5. Seedream 4.5 2.89 / 5 9 $0.048
6. FLUX.2 2.78 / 5 9 $0.013
7. Ideogram 3.0 2.78 / 5 9 $0.060
8. FLUX.2 [dev] Turbo 2.67 / 5 9 $0.015
9. Seedream 5.0 Lite 2.56 / 5 9 $0.020
10. FLUX.2 [pro] 2.33 / 5 9 $0.040
11. Stable Diffusion 1.5 2.22 / 5 9 $0.010
12. Luma Photon 2.11 / 5 9 $0.019
13. Qwen-Image 2.00 / 5 9 $0.030
14. Recraft V3 1.89 / 5 9 $0.040
15. Gemini 2.5 Flash Image 1.67 / 5 9 $0.039
Evidence

The hardest test: a strict two-color duotone

The duotone prompt allows exactly two colors, specified by hex. It is the lowest-scoring brand-color benchmark in the suite: models drift the hues, add gradient shading, or introduce third colors.
Full scores and prompt text: Duotone Portrait
Takeaway

What this means if you are building with AI images

Brand color is a deterministic requirement: your palette is a set of exact values, not a vibe. The data says treating the model as the enforcement mechanism will fail a measurable fraction of the time on every model available today.

  • Enforce the palette after generation, not in the prompt. Prompting with hex codes improves the odds but never guarantees the value. Snapping fills, overlays and text to a brand kit on an editable canvas turns a near-miss into an exact match without a regeneration lottery.
  • Route color-critical jobs to the top of this table. The spread between the best and worst model is more than two full points. A multi-model gateway can send duotones and brand-heavy work to the leaders and volume work to cheaper generalists.

The pattern repeats across the benchmark: generation is probabilistic, brand requirements are deterministic, and the gap is closed in an editing step where colors are values, not suggestions. Building that editing step is the job of a human-in-the-loop editor like the IMG.LY AI Editor: generation next to the manual controls the models cannot deliver, brand kits, exact color values and snapping on an editable canvas, so your users get the control they need to be productive with generative AI. See it working in the live AI Editor demo.

AI Editor

Ship on-brand AI imagery

Give your users the control they need to be productive with generative AI. The IMG.LY AI Editor pairs generation with brand kits and exact color controls on an editable canvas, with the AI Gateway routing color-critical jobs to the best model.

IMG.LY AI Editor: generative image tools next to manual editing controls on a canvas