AI cannot hold a character across scenes yet
Our consistency series describes one character precisely and asks for the same person in three different scenes. It is the lowest-ceiling capability in the whole benchmark: the best model, Nano Banana Pro, averages 3.89 of 5 adherence, and every model drifts somewhere: outfits change, faces shift, props disappear. Identity survives a scene change only approximately.
What every model scored
| Consistency adherence | Runs scored | $ / image | |
|---|---|---|---|
| 1. Nano Banana Pro Holds identity best | 3.89 / 5 | 9 | $0.360 |
| 2. Nano Banana 2 Lite | 3.89 / 5 | 9 | $0.020 |
| 3. FLUX.2 | 3.75 / 5 | 9 | $0.013 |
| 4. Seedream 4.5 | 3.75 / 5 | 9 | $0.048 |
| 5. Nano Banana 2 | 3.75 / 5 | 9 | $0.100 |
| 6. FLUX.2 [pro] | 3.75 / 5 | 9 | $0.040 |
| 7. FLUX.2 [dev] Turbo | 3.75 / 5 | 9 | $0.015 |
| 8. GPT Image 1.5 | 3.75 / 5 | 9 | $0.120 |
| 9. Seedream 5.0 Lite | 3.75 / 5 | 9 | $0.020 |
| 10. Luma Photon | 3.75 / 5 | 9 | $0.019 |
| 11. Ideogram 3.0 | 3.75 / 5 | 9 | $0.060 |
| 12. Qwen-Image | 3.75 / 5 | 9 | $0.030 |
| 13. Gemini 2.5 Flash Image | 3.75 / 5 | 9 | $0.039 |
| 14. Stable Diffusion 1.5 | 3.47 / 5 | 9 | $0.010 |
| 15. Recraft V3 | 2.50 / 5 | 9 | $0.040 |
Best against weakest, scene by scene
Best: Nano Banana Pro
Weakest: Recraft V3
What this means if you are building with AI images
Series work runs on identity: a mascot across a campaign, a character through a storybook, a brand persona across a feed. The data says no amount of prompt detail makes regeneration a reliable identity mechanism today.
- Make identity an asset, not a prompt. Generate the character once, cut it out, and reuse it as a placed asset or template element across scenes. The background can be generated per scene; the identity is composited deterministically.
- Budget for human review on anything serial. Drift is gradual and easy to miss frame by frame. A canvas where an editor can swap a hairline or fix a wardrobe color beats a regeneration loop that changes everything at once.
This is the strongest version of the benchmark's recurring pattern: the more deterministic your requirement, the earlier the model has to hand off to the editing step. The IMG.LY AI Editor is built as that handoff: characters become reusable assets your users can cut out, place and fix manually on a canvas, giving them the control they need to be productive with generative AI. See it working in the live AI Editor demo.
Get the 2026 Benchmark Report
15 models, 37 prompts, 1,662 measured images. Every finding and ranking from this benchmark, with the methodology behind the numbers, as a PDF in your inbox.
Check your inbox.
Your report is on the way.
Keep your characters consistent
Give your users the control they need to be productive with generative AI. The IMG.LY AI Editor keeps identity deterministic with reusable assets, templates and manual fixes on the canvas, with the AI Gateway generating scenes with any model.


