GenAI Benchmarks

The transparency problem in AI image generation

We asked 15 models for die-cut stickers, isolated objects and shirt graphics on transparent backgrounds, and measured the alpha channel of every file they returned. 13 of 15 scored exactly zero: they paint the subject on a white or checkerboard background instead of emitting real transparency. The field averages 0.4 of 5 on this criterion, the lowest of any criterion in the benchmark.

The benchmark avocado mascot comparing a truly transparent sticker of itself with one printed on a fake checkerboard background
Scores

What every model scored

Mean transparency score over 3 cut-out benchmarks, three seeds each. The score measures whether the returned file carries a genuine alpha channel with clean edges, not whether the image merely looks isolated.
Transparency criterion, full 15-model run (pilot-0).
  Transparency score Runs scored $ / image
1. GPT Image 1.5
Only real alpha
4.22 / 5 9 $0.120
2. Recraft V3 2.00 / 5 9 $0.040
3. FLUX.2 [pro] 0.00 / 5 9 $0.040
4. FLUX.2 [dev] Turbo 0.00 / 5 9 $0.015
5. FLUX.2 0.00 / 5 9 $0.013
6. Gemini 2.5 Flash Image 0.00 / 5 9 $0.039
7. Ideogram 3.0 0.00 / 5 9 $0.060
8. Luma Photon 0.00 / 5 9 $0.019
9. Nano Banana 2 Lite 0.00 / 5 9 $0.020
10. Nano Banana 2 0.00 / 5 9 $0.100
11. Nano Banana Pro 0.00 / 5 9 $0.360
12. Qwen-Image 0.00 / 5 9 $0.030
13. Seedream 4.5 0.00 / 5 9 $0.048
14. Seedream 5.0 Lite 0.00 / 5 9 $0.020
15. Stable Diffusion 1.5 0.00 / 5 9 $0.010
Evidence

The same sticker prompt, every model

Identical die-cut sticker prompt across the field. Most results look like stickers. Almost none of the files can actually be placed on a shirt, a phone case or a web page without background removal first.
Full scores and prompt text: Die-Cut Sticker
Takeaway

What this means if you are building with AI images

If your product promises stickers, apparel graphics, product cut-outs or any asset that composites onto another surface, the model output is not the deliverable. It is the input to a pipeline. Two consequences follow from the data:

  • Plan for background removal, not for luck. With one exception, no amount of prompt engineering produces a usable alpha channel. A deterministic background-removal and edge-cleanup step turns every model on this page into a viable source, and it removes the single biggest quality gamble from the workflow.
  • Route transparency-critical jobs deliberately. One model does emit real alpha. If a job cannot tolerate a cleanup step, a multi-model gateway can route that job to the model that delivers it and send everything else to cheaper generalists.

This is the pattern the whole benchmark keeps pointing at: generation is probabilistic, production requirements are deterministic, and the gap between them is closed by an editing step with a human in the loop. Building that editing step is a product job of its own, and it is exactly the job the IMG.LY AI Editor is built for: it pairs generation with the manual controls the models cannot deliver, one-click background removal, edge cleanup, masking and compositing on a real canvas. The extra step the data says is unavoidable becomes a feature of your product instead of a support ticket, and your users get the control they need to be productive with generative AI. See it working in the live AI Editor demo.

AI Editor

Close the transparency gap in your product

Give your users the control they need to be productive with generative AI. The IMG.LY AI Editor pairs every model on this page with the manual last mile they lack: background removal, edge cleanup and on-canvas compositing, with the AI Gateway routing each job to the right model.

IMG.LY AI Editor: generative image tools next to manual editing controls on a canvas