GenAI Benchmarks

What the data says

Findings from benchmarking 15 models on 37 design-production prompts: the places where every model falls short of production requirements, measured rather than asserted. Each finding names the gap, shows the evidence, and spells out what the editing step has to cover.

The benchmark avocado mascot as a detective, examining an evidence board of pinned images through a magnifying glass
Transparency

13 of 15 models fail at transparency

We asked every model for transparent PNGs and measured the alpha channel of what came back. Almost none of it survives contact with a real sticker, merch or cut-out pipeline.

Brand color

No model hits your exact hex

Measured with CIEDE2000 against required brand colors, the best model scores 3.67 of 5 and most of the field lands below 3. Close-enough color is not the color in your brand book.

Text

Rendered text: close is not shippable

Even the best model occasionally breaks a headline, and the model famous for text lands mid-field. One wrong character means regenerating the whole image, unless the words are editable layers.

Consistency

No model holds a character across scenes

The same described character drifts between scenes on every model we tested. Series work needs identity as a reusable asset, not a regeneration lottery.

AI Editor

The pattern behind every finding

Generation is probabilistic; production requirements are deterministic. The IMG.LY AI Editor closes the gap with manual controls for background removal, exact colors and editable text, giving your users the control they need to be productive with generative AI.

IMG.LY AI Editor: generative image tools next to manual editing controls on a canvas