AI text rendering: close is not shippable
We required exact strings in 4 typography benchmarks and transcribed what 15 models actually rendered. 3 models, led by FLUX.2 [pro], rendered every required string in this suite exactly. Below the top tier outright misspelling is routine, dropped words and stray characters follow, and the legacy baseline manages 1.17. One wrong character makes a headline unusable, and baked-in text cannot be fixed without regenerating the whole image.
What every model scored
| Text accuracy | Runs scored | $ / image | |
|---|---|---|---|
| 1. FLUX.2 [pro] Most reliable | 5.00 / 5 | 24 | $0.040 |
| 2. GPT Image 1.5 | 5.00 / 5 | 24 | $0.120 |
| 3. Nano Banana Pro | 5.00 / 5 | 24 | $0.360 |
| 4. Gemini 2.5 Flash Image | 4.96 / 5 | 24 | $0.039 |
| 5. Nano Banana 2 | 4.96 / 5 | 24 | $0.100 |
| 6. FLUX.2 [dev] Turbo | 4.92 / 5 | 24 | $0.015 |
| 7. FLUX.2 | 4.92 / 5 | 24 | $0.013 |
| 8. Nano Banana 2 Lite | 4.92 / 5 | 24 | $0.020 |
| 9. Seedream 4.5 | 4.88 / 5 | 24 | $0.048 |
| 10. Seedream 5.0 Lite | 4.88 / 5 | 24 | $0.020 |
| 11. Qwen-Image | 4.83 / 5 | 24 | $0.030 |
| 12. Ideogram 3.0 | 4.71 / 5 | 24 | $0.060 |
| 13. Luma Photon | 4.13 / 5 | 24 | $0.019 |
| 14. Recraft V3 | 4.00 / 5 | 24 | $0.040 |
| 15. Stable Diffusion 1.5 | 1.17 / 5 | 24 | $0.010 |
One word, every model
What this means if you are building with AI images
Text is the highest-stakes element in a generated asset: it is the part users read, the part legal reviews, and the part that makes a design unusable when a single character is wrong. The data supports one production pattern:
- Keep headlines as editable text layers. Generate the visual, set the words as real typography on a canvas. A wrong word becomes a two-second fix instead of a regeneration lottery, fonts stay licensed and brand-exact, and localization means swapping a string rather than re-prompting per language.
- When text must be in the image, route to the top of this table. The gap between the leader and mid-field is the difference between an occasional retry and a review queue full of typos.
Even at the top, treat rendered text as a draft. The models are close; products that ship are exact. Closing that gap is the job of a human-in-the-loop editor like the IMG.LY AI Editor: generated visuals next to real, editable typography, so a wrong word is a keystroke fix instead of a regeneration lottery and your users get the control they need to be productive with generative AI. See it working in the live AI Editor demo.
Get the 2026 Benchmark Report
15 models, 37 prompts, 1,662 measured images. Every finding and ranking from this benchmark, with the methodology behind the numbers, as a PDF in your inbox.
Check your inbox.
Your report is on the way.
Generate the visual. Keep the words editable.
Give your users the control they need to be productive with generative AI. The IMG.LY AI Editor generates the visual and keeps the words as real, editable text layers, with the AI Gateway routing text-critical jobs to the models at the top of this table.


