GenAI Benchmarks

“I want my users to create designs in my app”

Social posts, invitations and personal designs created by your users inside your product. What matters when a person is watching a spinner: latency first, unit cost at user volume, and output that lands close enough to edit.

In-app UGC creation
The benchmark avocado mascot drawing on a giant smartphone screen, surrounded by floating sticker panels
Scores

Recommended models

Same measurements as the main rankings, weighted for this job. The weight matrix is 100% covered by scored criteria (auto tier plus the interim Claude-VLM quality tier); any criterion still awaiting review is excluded and the weights renormalized, never faked.
Pilot dataset (pilot-0), weighted for in-app ugc creation.
  Use-case score Criteria covered $ / image
1. Nano Banana 2 Lite
Best fit
4.12 / 5 100% $0.020
2. FLUX.2 4.03 / 5 100% $0.013
3. FLUX.2 [dev] Turbo 3.96 / 5 100% $0.015
4. Gemini 2.5 Flash Image 3.70 / 5 100% $0.039
5. Seedream 4.5 3.60 / 5 100% $0.048
6. FLUX.2 [pro] 3.55 / 5 100% $0.040
7. Qwen-Image 3.54 / 5 100% $0.030
8. Nano Banana 2 3.47 / 5 100% $0.100
9. Seedream 5.0 Lite 3.46 / 5 100% $0.020
10. GPT Image 1.5 3.24 / 5 100% $0.120
11. Luma Photon 3.23 / 5 100% $0.019
12. Recraft V3 3.23 / 5 100% $0.040
13. Ideogram 3.0 3.13 / 5 100% $0.060
14. Stable Diffusion 1.5 3.01 / 5 100% $0.010
15. Nano Banana Pro 2.77 / 5 100% $0.360
Want the raw numbers instead? See the full model rankings
Methodology

How this is weighted

A person is watching the spinner, so latency leads and cost follows: consumer products generate on every tap and pay for every retry. Prompt adherence weighs the same as cost because a result that ignores the prompt is a retry. The fidelity criteria stay light: UGC is judged by its creator, not by a brand team.

Weights are provisional (pilot-0) and published in full; the reviewed matrix ships with the frozen suite. Read the methodology
  • Latency 25%
  • Cost 20%
  • Prompt Adherence 20%
  • Composition 10%
  • Text Accuracy 10%
  • Color Accuracy 10%
  • Resolution 5%
Evidence

The evidence: Nano Banana 2 Lite

Nano Banana 2 Lite takes the top spot where this job puts its weight: latency at 3.9 / 5 (25% of the matrix), prompt adherence at 4.7 / 5 (20% of the matrix) and cost at 4.0 / 5 (20% of the matrix).

The best-fit model's results on the benchmarks that matter for this use case. Click through for the full cross-model comparison.
Takeaway

The last mile: what no model delivers

In-app generation is where the editing canvas stops being a nice-to-have and becomes the product surface: the model proposes, the user refines. The data supports routing for speed here, since the fast tier holds 4.0-plus quality at around two seconds while slow endpoints run over thirty. What no model provides is the refinement step itself: swapping the headline, nudging the layout, applying the user's colors. That loop is the canvas, and it is also what turns one generation into a kept, personalized design instead of a re-roll.

AI Editor

Generate with any model. Finish in the AI Editor.

Give your users the control they need to be productive with generative AI. The IMG.LY AI Editor turns in-app ugc creation generations into finished, on-spec assets: background removal, brand kits and editable text on a real canvas.

IMG.LY AI Editor: generative image tools next to manual editing controls on a canvas