What the data says
Findings from benchmarking 15 models on 37 design-production prompts: the places where every model falls short of production requirements, measured rather than asserted. Each finding names the gap, shows the evidence, and spells out what the editing step has to cover.
13 of 15 models fail at transparency
We asked every model for transparent PNGs and measured the alpha channel of what came back. Almost none of it survives contact with a real sticker, merch or cut-out pipeline.
No model hits your exact hex
Measured with CIEDE2000 against required brand colors, the best model scores 3.67 of 5 and most of the field lands below 3. Close-enough color is not the color in your brand book.
Rendered text: close is not shippable
Even the best model occasionally breaks a headline, and the model famous for text lands mid-field. One wrong character means regenerating the whole image, unless the words are editable layers.
No model holds a character across scenes
The same described character drifts between scenes on every model we tested. Series work needs identity as a reusable asset, not a regeneration lottery.
The pattern behind every finding
Generation is probabilistic; production requirements are deterministic. The IMG.LY AI Editor closes the gap with manual controls for background removal, exact colors and editable text, giving your users the control they need to be productive with generative AI.


