GenAI Benchmarks

AI text rendering: close is not shippable

We required exact strings in 8 benchmarks and transcribed what 23 models actually rendered. 9 models rendered every required string in this suite exactly. Below the top tier the errors are dropped words, stray characters and invented punctuation rather than misspellings, and the legacy baseline manages 1.17. One wrong character makes a headline unusable, and baked-in text cannot be fixed without regenerating the whole image.

The benchmark avocado mascot scratching its head at a hand-painted sign that misspells the word avocado
Scores

What every model scored

Mean text-accuracy score across the typography suite, three seeds each. Each required string is scored by how faithfully it appears in a transcription of the image. Case and line breaks are forgiven, so points are lost only to real character errors: misspellings, dropped words, stray glyphs.
Text-accuracy criterion, full 23-model run (pilot-0).
  Text accuracy Runs scored $ / image
1. FLUX.2 [pro]
Most reliable
5.00 / 5 24 $0.030
2. GPT Image 1.5 5.00 / 5 24 $0.140
3. GPT Image 2.5 Flare 5.00 / 5 24 $0.042
4. GPT Image 2 5.00 / 5 24 $0.155
5. Grok Imagine Image 2.0 5.00 / 5 24 $0.060
6. Muse Image 5.00 / 5 24 $0.010
7. Nano Banana Pro 5.00 / 5 24 $0.150
8. Qwen Image 3.0 5.00 / 5 24 $0.040
9. Recraft V4.1 5.00 / 5 24 $0.035
10. Gemini 2.5 Flash Image 4.96 / 5 24 $0.040
11. GPT Image 2.5 Sunburst 4.96 / 5 24 $0.036
12. Nano Banana 2 4.96 / 5 24 $0.080
13. FLUX.2 [dev] Turbo 4.92 / 5 24 $0.008
14. FLUX.2 4.92 / 5 24 $0.012
15. Nano Banana 2 Lite 4.92 / 5 24 $0.042
16. Seedream 5.0 Pro 4.92 / 5 24 $0.135
17. Seedream 4.5 4.88 / 5 24 $0.040
18. Seedream 5.0 Lite 4.88 / 5 24 $0.035
19. Qwen-Image 4.83 / 5 24 $0.020
20. Ideogram 3.0 4.71 / 5 24 $0.060
21. Luma Photon 4.13 / 5 24 $0.021
22. Recraft V3 4.00 / 5 24 $0.040
23. Stable Diffusion 1.5 1.17 / 5 24 $0.002
The reputation check: Ideogram, the model known for text rendering, lands at #20 of 23 here. The general flagships have absorbed the lead that made text specialists special.
Evidence

One word, every model

The wordmark benchmark asks for a single word in a clean geometric sans serif. Every failure mode in the suite shows up here at a glance: perfect renders, all-caps drift, broken line splits, and fully garbled letterforms.
Full scores and prompt text: Wordmark
Takeaway

What this means if you are building with AI images

Text is the highest-stakes element in a generated asset: it is the part users read, the part legal reviews, and the part that makes a design unusable when a single character is wrong. The data supports one production pattern:

  • Keep headlines as editable text layers. Generate the visual, set the words as real typography on a canvas. A wrong word becomes a two-second fix instead of a regeneration lottery, fonts stay licensed and brand-exact, and localization means swapping a string rather than re-prompting per language.
  • When text must be in the image, route to the top of this table. The gap between the leader and mid-field is the difference between an occasional retry and a review queue full of typos.

Even at the top, treat rendered text as a draft. The models are close; products that ship are exact. Closing that gap is the job of a human-in-the-loop editor like the IMG.LY AI Editor: generated visuals next to real, editable typography, so a wrong word is a keystroke fix instead of a regeneration lottery and your users get the control they need to be productive with generative AI. See it working in the live AI Editor demo.

AI Editor

Generate the visual. Keep the words editable.

Give your users the control they need to be productive with generative AI. The IMG.LY AI Editor generates the visual and keeps the words as real, editable text layers, with the AI Gateway routing text-critical jobs to the models at the top of this table.

IMG.LY AI Editor: generative image tools next to manual editing controls on a canvas