GenAI Benchmarks

“I want to personalize creatives at scale”

Thousands of variants per campaign, generated by automation rather than by hand: per-segment backgrounds, per-market scenes, per-recipient artwork. What matters: obedience to structured, machine-generated prompts, and unit cost that survives multiplication.

Personalization at scale
The benchmark avocado mascot stamping different names onto a conveyor belt of identical greeting cards
Scores

Recommended models

Same measurements as the main rankings, weighted for this job. The weight matrix is 100% covered by scored criteria (auto tier plus the interim Claude-VLM quality tier); any criterion still awaiting review is excluded and the weights renormalized, never faked.
Pilot dataset (pilot-0), weighted for personalization at scale.
  Use-case score Criteria covered $ / image
1. FLUX.2
Best fit
4.41 / 5 100% $0.012
2. FLUX.2 [dev] Turbo 4.41 / 5 100% $0.008
3. Muse Image 4.35 / 5 100% $0.010
4. Nano Banana 2 Lite 4.34 / 5 100% $0.042
5. Recraft V4.1 4.13 / 5 100% $0.035
6. Seedream 5.0 Lite 4.08 / 5 100% $0.035
7. Gemini 2.5 Flash Image 4.08 / 5 100% $0.040
8. Seedream 4.5 4.07 / 5 100% $0.040
9. GPT Image 2.5 Sunburst 4.03 / 5 100% $0.036
10. FLUX.2 [pro] 4.02 / 5 100% $0.030
11. GPT Image 2.5 Flare 4.02 / 5 100% $0.042
12. Nano Banana 2 3.98 / 5 100% $0.080
13. Qwen-Image 3.98 / 5 100% $0.020
14. Grok Imagine Image 2.0 3.91 / 5 100% $0.060
15. Seedream 5.0 Pro 3.85 / 5 100% $0.135
16. Qwen Image 3.0 3.79 / 5 100% $0.040
17. Luma Photon 3.76 / 5 100% $0.021
18. GPT Image 1.5 3.73 / 5 100% $0.140
19. Nano Banana Pro 3.65 / 5 100% $0.150
20. Ideogram 3.0 3.64 / 5 100% $0.060
21. GPT Image 2 3.61 / 5 100% $0.155
22. Recraft V3 3.47 / 5 100% $0.040
23. Stable Diffusion 1.5 3.09 / 5 100% $0.002
Want the raw numbers instead? See the full model rankings
Methodology

How this is weighted

Automation has no human retry loop, so prompt adherence leads: every variant that ignores its data-driven brief is a silent defect in someone's mailbox. Cost and latency follow because the same job runs thousands of times per campaign; per-image quality criteria stay light because variants share one reviewed template and inherit its typography and colors.

Weights are provisional (pilot-0) and published in full; the reviewed matrix ships with the frozen suite. Read the methodology
  • Prompt Adherence 30%
  • Cost 25%
  • Latency 15%
  • Text Accuracy 10%
  • Composition 10%
  • Color Accuracy 5%
  • Resolution 5%
Evidence

The evidence: FLUX.2

FLUX.2 takes the top spot where this job puts its weight: prompt adherence at 4.4 / 5 (30% of the matrix), cost at 4.9 / 5 (25% of the matrix) and latency at 4.8 / 5 (15% of the matrix).

The best-fit model's results on the benchmarks that matter for this use case. Click through for the full cross-model comparison.
Takeaway

The last mile: what no model delivers

Our structured-spec benchmark shows the strongest models follow machine-generated prompts remarkably well, and the budget tier makes per-variant cost viable. What generation cannot do is guarantee the deterministic parts of a variant: the recipient's name spelled correctly, the legal line, the logo placement, the brand hex. Automation pipelines that ship put those elements on the canvas as template layers and let the model supply only the interchangeable visual underneath.

That editing step is what the print personalization editor provides.

AI Editor

Generate with any model. Finish in the AI Editor.

Give your users the control they need to be productive with generative AI. The IMG.LY AI Editor turns personalization at scale generations into finished, on-spec assets: background removal, brand kits and editable text on a real canvas.

IMG.LY AI Editor: generative image tools next to manual editing controls on a canvas