GenAI Benchmarks

“I want to create merch and stickers”

T-shirt graphics, die-cut stickers and POD assets, generated at catalog scale. The dealbreakers: no real alpha channel, halo edges, and designs that fall apart at production size.

Apparel, merch & POD
The benchmark avocado mascot holding a life-size die-cut sticker of itself, next to a mug and a t-shirt with the same design
Scores

Recommended models

Same measurements as the main rankings, weighted for this job. The weight matrix is 100% covered by scored criteria (auto tier plus the interim Claude-VLM quality tier); any criterion still awaiting review is excluded and the weights renormalized, never faked.
Pilot dataset (pilot-0), weighted for apparel, merch & pod.
  Use-case score Criteria covered $ / image
1. GPT Image 1.5
Best fit
3.77 / 5 100% $0.140
2. Recraft V4.1 3.17 / 5 100% $0.035
3. Recraft V3 2.82 / 5 100% $0.040
4. Muse Image 2.78 / 5 100% $0.010
5. Seedream 4.5 2.70 / 5 100% $0.040
6. Seedream 5.0 Lite 2.67 / 5 100% $0.035
7. Seedream 5.0 Pro 2.63 / 5 100% $0.135
8. Nano Banana 2 Lite 2.52 / 5 100% $0.042
9. FLUX.2 2.46 / 5 100% $0.012
10. FLUX.2 [dev] Turbo 2.46 / 5 100% $0.008
11. Nano Banana 2 2.41 / 5 100% $0.080
12. Luma Photon 2.41 / 5 100% $0.021
13. GPT Image 2.5 Sunburst 2.41 / 5 100% $0.036
14. GPT Image 2.5 Flare 2.39 / 5 100% $0.042
15. Grok Imagine Image 2.0 2.37 / 5 100% $0.060
16. Qwen Image 3.0 2.34 / 5 100% $0.040
17. Gemini 2.5 Flash Image 2.33 / 5 100% $0.040
18. Nano Banana Pro 2.32 / 5 100% $0.150
19. FLUX.2 [pro] 2.25 / 5 100% $0.030
20. Ideogram 3.0 2.24 / 5 100% $0.060
21. Qwen-Image 2.19 / 5 100% $0.020
22. GPT Image 2 2.19 / 5 100% $0.155
23. Stable Diffusion 1.5 1.74 / 5 100% $0.002
Want the raw numbers instead? See the full model rankings
Methodology

How this is weighted

A POD asset that arrives without a real alpha channel is not a product, it is a background-removal ticket, so transparency dominates the matrix. Resolution, prompt adherence and cost share the next tier because catalog-scale generation multiplies every per-image weakness; color, text and latency trail for graphics that are typically cleaned up or upscaled before printing anyway.

Weights are provisional (pilot-0) and published in full; the reviewed matrix ships with the frozen suite. Read the methodology
  • Transparency 35%
  • Resolution 15%
  • Prompt Adherence 15%
  • Cost 15%
  • Color Accuracy 10%
  • Text Accuracy 5%
  • Latency 5%
Evidence

The evidence: GPT Image 1.5

GPT Image 1.5 takes the top spot where this job puts its weight: transparency at 4.2 / 5 (35% of the matrix), prompt adherence at 4.6 / 5 (15% of the matrix) and cost at 3.1 / 5 (15% of the matrix).

The best-fit model's results on the benchmarks that matter for this use case. Click through for the full cross-model comparison.
Takeaway

The last mile: what no model delivers

Transparency is the clearest gap in the whole benchmark: across every cut-out prompt the average alpha score is 0.36 of 5, and only one model emits a usable alpha channel. If your product promises stickers or apparel graphics, plan for background removal and edge cleanup in an editing step, not for the model to deliver a production file.

That editing step is what the AI editor for print-on-demand provides.

AI Editor

Generate with any model. Finish in the AI Editor.

Give your users the control they need to be productive with generative AI. The IMG.LY AI Editor turns apparel, merch & pod generations into finished, on-spec assets: background removal, brand kits and editable text on a real canvas.

IMG.LY AI Editor: generative image tools next to manual editing controls on a canvas