GenAI Benchmarks

“I want to generate product ads”

Hero images, campaign backgrounds and ad variants, generated per SKU or per audience inside your platform. What matters: composition that leaves room for copy, brand-color fidelity, and unit cost at campaign volume.

Marketing & ad creative
The benchmark avocado mascot directing a perfume-bottle photo shoot from a director’s chair with a megaphone
Scores

Recommended models

Same measurements as the main rankings, weighted for this job. The weight matrix is 100% covered by scored criteria (auto tier plus the interim Claude-VLM quality tier); any criterion still awaiting review is excluded and the weights renormalized, never faked.
Pilot dataset (pilot-0), weighted for marketing & ad creative.
  Use-case score Criteria covered $ / image
1. Muse Image
Best fit
4.21 / 5 100% $0.010
2. Nano Banana 2 Lite 4.17 / 5 100% $0.042
3. Seedream 5.0 Lite 4.12 / 5 100% $0.035
4. Seedream 5.0 Pro 4.09 / 5 100% $0.135
5. FLUX.2 4.04 / 5 100% $0.012
6. Seedream 4.5 4.03 / 5 100% $0.040
7. FLUX.2 [dev] Turbo 3.96 / 5 100% $0.008
8. Recraft V4.1 3.95 / 5 100% $0.035
9. Nano Banana 2 3.91 / 5 100% $0.080
10. GPT Image 2.5 Sunburst 3.89 / 5 100% $0.036
11. Grok Imagine Image 2.0 3.86 / 5 100% $0.060
12. GPT Image 2.5 Flare 3.84 / 5 100% $0.042
13. GPT Image 1.5 3.83 / 5 100% $0.140
14. Gemini 2.5 Flash Image 3.80 / 5 100% $0.040
15. FLUX.2 [pro] 3.77 / 5 100% $0.030
16. Qwen-Image 3.77 / 5 100% $0.020
17. GPT Image 2 3.66 / 5 100% $0.155
18. Nano Banana Pro 3.60 / 5 100% $0.150
19. Qwen Image 3.0 3.60 / 5 100% $0.040
20. Ideogram 3.0 3.59 / 5 100% $0.060
21. Luma Photon 3.54 / 5 100% $0.021
22. Recraft V3 3.18 / 5 100% $0.040
23. Stable Diffusion 1.5 2.64 / 5 100% $0.002
Want the raw numbers instead? See the full model rankings
Methodology

How this is weighted

Ad creative lives or dies on composition: the image has to leave room for the offer, the logo and the CTA, so it leads the matrix. Adherence, brand color and cost share the second tier because campaign variants are generated in volume against a brief; resolution, text and latency matter less for assets that are finished in an editor before trafficking.

Weights are provisional (pilot-0) and published in full; the reviewed matrix ships with the frozen suite. Read the methodology
  • Composition 25%
  • Prompt Adherence 15%
  • Color Accuracy 15%
  • Cost 15%
  • Resolution 10%
  • Text Accuracy 10%
  • Latency 10%
Evidence

The evidence: Muse Image

Muse Image takes the top spot where this job puts its weight: composition at 4.3 / 5 (25% of the matrix), cost at 5.0 / 5 (15% of the matrix) and prompt adherence at 4.8 / 5 (15% of the matrix).

The best-fit model's results on the benchmarks that matter for this use case. Click through for the full cross-model comparison.
Takeaway

The last mile: what no model delivers

Models reserve negative space well, but they miss exact brand hexes and cannot place your logo, legal line or localized copy. Ad production at volume works as a pipeline: generate the background, then compose brand elements and text as deterministic layers on top, so every variant stays on brand without a regeneration lottery.

AI Editor

Generate with any model. Finish in the AI Editor.

Give your users the control they need to be productive with generative AI. The IMG.LY AI Editor turns marketing & ad creative generations into finished, on-spec assets: background removal, brand kits and editable text on a real canvas.

IMG.LY AI Editor: generative image tools next to manual editing controls on a canvas