Die-Cut Sticker
A die-cut sticker needs a genuinely transparent background: real alpha, not a checkerboard pattern painted into the pixels. This benchmark checks whether the model emits a true alpha channel that a production pipeline can composite. We included it because print-on-demand and merch workflows depend on it, and because it exposes a gap the whole field currently shares. In the pilot, no model produced usable transparency, which is a finding worth publishing on its own.
Die-cut sticker of a smiling cartoon avocado wearing sunglasses, thick white sticker border, transparent background.
Results
FLUX.2
FLUX.2 [pro]
FLUX.2 [dev] Turbo
Gemini 2.5 Flash Image
GPT Image 1.5
Ideogram 3.0
Luma Photon
Nano Banana 2
Nano Banana 2 Lite
Nano Banana Pro
Qwen-Image
Recraft V3
Seedream 4.5
Seedream 5.0 Lite
Stable Diffusion 1.5
Run-to-run consistency
Scores
| Prompt Adherence | Transparency | Resolution | Latency | Cost | |
|---|---|---|---|---|---|
| FLUX.2 | 3.9 / 5 | 0.0 / 5 | 2.0 / 5 | 4.3 / 5 | 4.0 / 5 |
| FLUX.2 [pro] | 4.4 / 5 | 0.0 / 5 | 2.0 / 5 | 2.7 / 5 | 4.0 / 5 |
| FLUX.2 [dev] Turbo | 4.4 / 5 | 0.0 / 5 | 2.0 / 5 | 4.3 / 5 | 4.0 / 5 |
| Gemini 2.5 Flash Image | 4.4 / 5 | 0.0 / 5 | 3.0 / 5 | 3.0 / 5 | 4.0 / 5 |
| GPT Image 1.5 | 3.9 / 5 | 5.0 / 5 | 3.0 / 5 | 1.0 / 5 | 3.0 / 5 |
| Ideogram 3.0 | 3.3 / 5 | 0.0 / 5 | 3.0 / 5 | 1.7 / 5 | 3.0 / 5 |
| Luma Photon | 4.4 / 5 | 0.0 / 5 | 4.0 / 5 | 2.0 / 5 | 4.0 / 5 |
| Nano Banana 2 | 5.0 / 5 | 0.0 / 5 | 3.0 / 5 | 2.0 / 5 | 3.0 / 5 |
| Nano Banana 2 Lite | 5.0 / 5 | 0.0 / 5 | 3.0 / 5 | 4.0 / 5 | 4.0 / 5 |
| Nano Banana Pro | 5.0 / 5 | 0.0 / 5 | 3.0 / 5 | 1.0 / 5 | 1.0 / 5 |
| Qwen-Image | 3.3 / 5 | 0.0 / 5 | 2.0 / 5 | 3.0 / 5 | 4.0 / 5 |
| Recraft V3 | 3.3 / 5 | 2.0 / 5 | 3.0 / 5 | 3.0 / 5 | 4.0 / 5 |
| Seedream 4.5 | 3.9 / 5 | 0.0 / 5 | 5.0 / 5 | 2.0 / 5 | 4.0 / 5 |
| Seedream 5.0 Lite | 3.3 / 5 | 0.0 / 5 | 5.0 / 5 | 1.0 / 5 | 4.0 / 5 |
| Stable Diffusion 1.5 | 2.8 / 5 | 0.0 / 5 | 1.0 / 5 | 4.0 / 5 | 5.0 / 5 |
New model? New benchmarks.
We re-run the identical 37-prompt suite on every model release. Get the scores, and what changed in the rankings, in your inbox.
You're on the list.
We'll email you when new benchmark results ship.
Frequently asked questions
Image generation is stochastic: the same prompt produces different images on every run, so a single sample measures luck, not ability. Every model runs each benchmark three times with fixed seeds (1111, 2222, 3333), or three unseeded samples where the API accepts no seed. Scores average all three samples, and the consistency benchmarks measure the variation itself.
Yes. Every model receives the same prompt text with default parameters and no per-model tuning, so differences in output reflect the model, not prompt engineering. The suite is versioned and prompts are append-only, which keeps historical scores comparable.
The results on this page are scored on this exact prompt, so the strongest model for it is easy to spot; this benchmark sits in the Transparency & Cutouts category. Scores are per model version and reflect this benchmark only. For a ranking across every benchmark see the model rankings, and to weight the numbers by a specific job see the use-case pages.
Each image is scored 0 to 5 per criterion. Measured criteria (resolution, latency, cost, transparency, color accuracy) are computed automatically; quality criteria (text accuracy, prompt adherence, composition) are judged by an automated Claude vision tier against each prompt’s checklist. Page scores are unweighted means over all of a model’s runs in that scope. Blind expert-panel review has not run yet; the dataset is pilot-0.
Model output is a starting point, not a finished asset. Production work usually needs background removal, exact brand colors or editable text, none of which generation guarantees on every run. The IMG.LY AI Editor gives your users those controls to refine any model’s output to production quality.
The suite re-runs on notable model releases so the rankings stay current. The results shown are the pilot-0 dataset, scored by measured criteria plus an automated Claude vision tier, with blind expert review planned. Prompts are versioned and append-only, so scores stay comparable across runs.

