GenAI Benchmarks

“I want to generate product imagery”

Staged product shots, seasonal backgrounds and PDP variants generated per SKU. What matters: faithful product staging, composition that survives a listing grid, resolution for zoom, and clean cut-outs for composite layouts.

E-commerce imagery
The benchmark avocado mascot staging a sneaker on a turntable inside a studio lightbox
Scores

Recommended models

Same measurements as the main rankings, weighted for this job. The weight matrix is 100% covered by scored criteria (auto tier plus the interim Claude-VLM quality tier); any criterion still awaiting review is excluded and the weights renormalized, never faked.
Pilot dataset (pilot-0), weighted for e-commerce imagery.
  Use-case score Criteria covered $ / image
1. Muse Image
Best fit
4.00 / 5 100% $0.010
2. Seedream 5.0 Lite 3.96 / 5 100% $0.035
3. Nano Banana 2 Lite 3.92 / 5 100% $0.042
4. Seedream 5.0 Pro 3.90 / 5 100% $0.135
5. Seedream 4.5 3.88 / 5 100% $0.040
6. Recraft V4.1 3.81 / 5 100% $0.035
7. GPT Image 1.5 3.75 / 5 100% $0.140
8. FLUX.2 3.73 / 5 100% $0.012
9. FLUX.2 [dev] Turbo 3.68 / 5 100% $0.008
10. Nano Banana 2 3.66 / 5 100% $0.080
11. Grok Imagine Image 2.0 3.62 / 5 100% $0.060
12. Gemini 2.5 Flash Image 3.60 / 5 100% $0.040
13. GPT Image 2.5 Sunburst 3.58 / 5 100% $0.036
14. GPT Image 2.5 Flare 3.54 / 5 100% $0.042
15. FLUX.2 [pro] 3.48 / 5 100% $0.030
16. Luma Photon 3.44 / 5 100% $0.021
17. Qwen-Image 3.44 / 5 100% $0.020
18. Qwen Image 3.0 3.40 / 5 100% $0.040
19. Nano Banana Pro 3.35 / 5 100% $0.150
20. GPT Image 2 3.34 / 5 100% $0.155
21. Ideogram 3.0 3.31 / 5 100% $0.060
22. Recraft V3 3.12 / 5 100% $0.040
23. Stable Diffusion 1.5 2.57 / 5 100% $0.002
Want the raw numbers instead? See the full model rankings
Methodology

How this is weighted

The product has to stay the product, so prompt adherence leads: a staged shot that subtly alters the SKU is a return waiting to happen. Composition and resolution follow because PDP images crop across breakpoints and zoom on hover. Cost matters at catalog scale, while alpha only covers the occasional cut-out treatment.

Weights are provisional (pilot-0) and published in full; the reviewed matrix ships with the frozen suite. Read the methodology
  • Prompt Adherence 25%
  • Composition 20%
  • Resolution 15%
  • Cost 15%
  • Color Accuracy 10%
  • Latency 10%
  • Transparency 5%
Evidence

The evidence: Muse Image

Muse Image takes the top spot where this job puts its weight: prompt adherence at 4.8 / 5 (25% of the matrix), composition at 4.3 / 5 (20% of the matrix) and cost at 5.0 / 5 (15% of the matrix).

The best-fit model's results on the benchmarks that matter for this use case. Click through for the full cross-model comparison.
Takeaway

The last mile: what no model delivers

Product staging is one of the most mature capabilities in the suite: the leaders hit five out of five on our cosmetics and sneaker prompts. The gaps are on either side of the shot. Real product fidelity requires compositing your actual product photo, not trusting a generated lookalike, and cut-outs need background removal because almost no model emits a real alpha channel. The shipping pattern is generate the scene, composite the SKU photo and price text as layers, and export per marketplace spec.

AI Editor

Generate with any model. Finish in the AI Editor.

Give your users the control they need to be productive with generative AI. The IMG.LY AI Editor turns e-commerce imagery generations into finished, on-spec assets: background removal, brand kits and editable text on a real canvas.

IMG.LY AI Editor: generative image tools next to manual editing controls on a canvas