GenAI Benchmarks

AI cannot hold a character across scenes yet

Our consistency series describes one character precisely and asks for the same person in three different scenes. No capability in the benchmark has a lower ceiling. The best model, Seedream 5.0 Pro, averages 4.86 of 5 adherence, and every model drifts somewhere: outfits change, faces shift, props disappear. Identity survives a scene change only approximately.

The benchmark avocado mascot comparing its reference photo against a police lineup of slightly wrong avocado look-alikes
Scores

What every model scored

Mean prompt adherence over the 3-scene consistency series, three seeds each, judged against a fixed checklist of identity details (hair, wardrobe, accessories, props) per scene.
Character-consistency adherence, full 23-model run (pilot-0).
  Consistency adherence Runs scored $ / image
1. Seedream 5.0 Pro
Holds identity best
4.86 / 5 9 $0.135
2. GPT Image 2.5 Sunburst 4.72 / 5 9 $0.036
3. Qwen Image 3.0 4.72 / 5 9 $0.040
4. Recraft V4.1 4.72 / 5 9 $0.035
5. Grok Imagine Image 2.0 4.58 / 5 9 $0.060
6. GPT Image 2 4.44 / 5 9 $0.155
7. GPT Image 2.5 Flare 4.17 / 5 9 $0.042
8. Muse Image 4.17 / 5 9 $0.010
9. Nano Banana Pro 3.89 / 5 9 $0.150
10. Nano Banana 2 Lite 3.89 / 5 9 $0.042
11. FLUX.2 3.75 / 5 9 $0.012
12. Seedream 4.5 3.75 / 5 9 $0.040
13. Nano Banana 2 3.75 / 5 9 $0.080
14. FLUX.2 [pro] 3.75 / 5 9 $0.030
15. FLUX.2 [dev] Turbo 3.75 / 5 9 $0.008
16. GPT Image 1.5 3.75 / 5 9 $0.140
17. Seedream 5.0 Lite 3.75 / 5 9 $0.035
18. Luma Photon 3.75 / 5 9 $0.021
19. Ideogram 3.0 3.75 / 5 9 $0.060
20. Qwen-Image 3.75 / 5 9 $0.020
21. Gemini 2.5 Flash Image 3.75 / 5 9 $0.040
22. Stable Diffusion 1.5 3.47 / 5 9 $0.002
23. Recraft V3 2.50 / 5 9 $0.040
Evidence

Best against weakest, scene by scene

The same described character in all three scenes. Even the leader's character subtly drifts between settings; the weakest renders three different people.
Takeaway

What this means if you are building with AI images

Series work runs on identity: a mascot across a campaign, a character through a storybook, a brand persona across a feed. The data says no amount of prompt detail makes regeneration a reliable identity mechanism today.

  • Make identity an asset, not a prompt. Generate the character once, cut it out, and reuse it as a placed asset or template element across scenes. The background can be generated per scene; the identity is composited deterministically.
  • Budget for human review on anything serial. Drift is gradual and easy to miss frame by frame. A canvas where an editor can swap a hairline or fix a wardrobe color beats a regeneration loop that changes everything at once.

This is the strongest version of the benchmark's recurring pattern: the more deterministic your requirement, the earlier the model has to hand off to the editing step. The IMG.LY AI Editor is built as that handoff: characters become reusable assets your users can cut out, place and fix manually on a canvas, giving them the control they need to be productive with generative AI. See it working in the live AI Editor demo.

AI Editor

Keep your characters consistent

Give your users the control they need to be productive with generative AI. The IMG.LY AI Editor keeps identity deterministic with reusable assets, templates and manual fixes on the canvas, with the AI Gateway generating scenes with any model.

IMG.LY AI Editor: generative image tools next to manual editing controls on a canvas