GenAI Benchmarks

“I want designs with real text”

Posters, flyers, social thumbnails and wordmarks where rendered text is the content, produced by your users or your automation. One misspelled headline makes the asset unusable.

The benchmark avocado mascot on a step ladder, straightening the letter A in a giant AURELIA headline
Scores

Recommended models

Same measurements as the main rankings, weighted for this job. The weight matrix is 100% covered by scored criteria (auto tier plus the interim Claude-VLM quality tier); any criterion still awaiting review is excluded and the weights renormalized, never faked.
Pilot dataset (pilot-0), weighted for text-heavy layouts.
  Use-case score Criteria covered $ / image
1. Muse Image
Best fit
4.49 / 5 100% $0.010
2. Seedream 5.0 Pro 4.43 / 5 100% $0.135
3. Seedream 5.0 Lite 4.42 / 5 100% $0.035
4. Nano Banana 2 Lite 4.37 / 5 100% $0.042
5. Seedream 4.5 4.36 / 5 100% $0.040
6. Recraft V4.1 4.28 / 5 100% $0.035
7. FLUX.2 4.25 / 5 100% $0.012
8. Nano Banana 2 4.24 / 5 100% $0.080
9. Grok Imagine Image 2.0 4.24 / 5 100% $0.060
10. GPT Image 1.5 4.21 / 5 100% $0.140
11. GPT Image 2.5 Sunburst 4.21 / 5 100% $0.036
12. FLUX.2 [dev] Turbo 4.20 / 5 100% $0.008
13. GPT Image 2.5 Flare 4.19 / 5 100% $0.042
14. Gemini 2.5 Flash Image 4.15 / 5 100% $0.040
15. FLUX.2 [pro] 4.13 / 5 100% $0.030
16. GPT Image 2 4.09 / 5 100% $0.155
17. Qwen Image 3.0 4.08 / 5 100% $0.040
18. Nano Banana Pro 4.07 / 5 100% $0.150
19. Qwen-Image 4.05 / 5 100% $0.020
20. Ideogram 3.0 3.95 / 5 100% $0.060
21. Luma Photon 3.77 / 5 100% $0.021
22. Recraft V3 3.45 / 5 100% $0.040
23. Stable Diffusion 1.5 2.07 / 5 100% $0.002
Want the raw numbers instead? See the full model rankings
Methodology

How this is weighted

When the text is the content, nothing else can compensate for a broken glyph, so text accuracy takes several times the weight of any other criterion. Composition and adherence follow because a poster with perfect type still fails when the layout ignores the brief; the remaining criteria stay light because typography failures are the reason these jobs get regenerated.

Weights are provisional (pilot-0) and published in full; the reviewed matrix ships with the frozen suite. Read the methodology
  • Text Accuracy 40%
  • Composition 15%
  • Prompt Adherence 10%
  • Color Accuracy 10%
  • Resolution 10%
  • Cost 10%
  • Latency 5%
Evidence

The evidence: Muse Image

Muse Image takes the top spot where this job puts its weight: text accuracy at 5.0 / 5 (40% of the matrix), composition at 4.3 / 5 (15% of the matrix) and cost at 5.0 / 5 (10% of the matrix).

The best-fit model's results on the benchmarks that matter for this use case. Click through for the full cross-model comparison.
Takeaway

The last mile: what no model delivers

Nine models render every required string in this suite exactly, and everything below that top tier misspells routinely. Baked-in AI text cannot be corrected without regenerating the whole image. The reliable pattern is generating the visual and keeping headlines as editable text layers on a canvas, where a human or a template fixes the last word.

AI Editor

Generate with any model. Finish in the AI Editor.

Give your users the control they need to be productive with generative AI. The IMG.LY AI Editor turns text-heavy layouts generations into finished, on-spec assets: background removal, brand kits and editable text on a real canvas.

IMG.LY AI Editor: generative image tools next to manual editing controls on a canvas