GenAI Benchmarks

“I want to personalize creatives at scale”

Thousands of variants per campaign, generated by automation rather than by hand: per-segment backgrounds, per-market scenes, per-recipient artwork. What matters: obedience to structured, machine-generated prompts, and unit cost that survives multiplication.

Personalization at scale
The benchmark avocado mascot stamping different names onto a conveyor belt of identical greeting cards
Scores

Recommended models

Same measurements as the main rankings, weighted for this job. The weight matrix is 100% covered by scored criteria (auto tier plus the interim Claude-VLM quality tier); any criterion still awaiting review is excluded and the weights renormalized, never faked.
Pilot dataset (pilot-0), weighted for personalization at scale.
  Use-case score Criteria covered $ / image
1. Nano Banana 2 Lite
Best fit
4.26 / 5 100% $0.020
2. FLUX.2 4.10 / 5 100% $0.013
3. FLUX.2 [dev] Turbo 4.05 / 5 100% $0.015
4. Gemini 2.5 Flash Image 3.97 / 5 100% $0.039
5. Seedream 4.5 3.91 / 5 100% $0.048
6. Seedream 5.0 Lite 3.88 / 5 100% $0.020
7. FLUX.2 [pro] 3.86 / 5 100% $0.040
8. Qwen-Image 3.75 / 5 100% $0.030
9. Nano Banana 2 3.73 / 5 100% $0.100
10. GPT Image 1.5 3.58 / 5 100% $0.120
11. Luma Photon 3.53 / 5 100% $0.019
12. Recraft V3 3.38 / 5 100% $0.040
13. Ideogram 3.0 3.34 / 5 100% $0.060
14. Nano Banana Pro 2.98 / 5 100% $0.360
15. Stable Diffusion 1.5 2.98 / 5 100% $0.010
Want the raw numbers instead? See the full model rankings
Methodology

How this is weighted

Automation has no human retry loop, so prompt adherence leads: every variant that ignores its data-driven brief is a silent defect in someone's mailbox. Cost and latency follow because the same job runs thousands of times per campaign; per-image quality criteria stay light because variants share one reviewed template and inherit its typography and colors.

Weights are provisional (pilot-0) and published in full; the reviewed matrix ships with the frozen suite. Read the methodology
  • Prompt Adherence 30%
  • Cost 25%
  • Latency 15%
  • Text Accuracy 10%
  • Composition 10%
  • Color Accuracy 5%
  • Resolution 5%
Evidence

The evidence: Nano Banana 2 Lite

Nano Banana 2 Lite takes the top spot where this job puts its weight: prompt adherence at 4.7 / 5 (30% of the matrix), cost at 4.0 / 5 (25% of the matrix) and latency at 3.9 / 5 (15% of the matrix).

The best-fit model's results on the benchmarks that matter for this use case. Click through for the full cross-model comparison.
Takeaway

The last mile: what no model delivers

Our structured-spec benchmark shows the strongest models follow machine-generated prompts remarkably well, and the budget tier makes per-variant cost viable. What generation cannot do is guarantee the deterministic parts of a variant: the recipient's name spelled correctly, the legal line, the logo placement, the brand hex. Automation pipelines that ship put those elements on the canvas as template layers and let the model supply only the interchangeable visual underneath.

AI Editor

Generate with any model. Finish in the AI Editor.

Give your users the control they need to be productive with generative AI. The IMG.LY AI Editor turns personalization at scale generations into finished, on-spec assets: background removal, brand kits and editable text on a real canvas.

IMG.LY AI Editor: generative image tools next to manual editing controls on a canvas