GenAI Benchmarks

“I want my users to create designs in my app”

Social posts, invitations and personal designs created by your users inside your product. What matters when a person is watching a spinner: latency first, unit cost at user volume, and output that lands close enough to edit.

In-app UGC creation
The benchmark avocado mascot drawing on a giant smartphone screen, surrounded by floating sticker panels
Scores

Recommended models

Same measurements as the main rankings, weighted for this job. The weight matrix is 100% covered by scored criteria (auto tier plus the interim Claude-VLM quality tier); any criterion still awaiting review is excluded and the weights renormalized, never faked.
Pilot dataset (pilot-0), weighted for in-app ugc creation.
  Use-case score Criteria covered $ / image
1. FLUX.2
Best fit
4.35 / 5 100% $0.012
2. FLUX.2 [dev] Turbo 4.35 / 5 100% $0.008
3. Nano Banana 2 Lite 4.23 / 5 100% $0.042
4. Muse Image 4.03 / 5 100% $0.010
5. Recraft V4.1 3.89 / 5 100% $0.035
6. Gemini 2.5 Flash Image 3.84 / 5 100% $0.040
7. Seedream 4.5 3.83 / 5 100% $0.040
8. Qwen-Image 3.81 / 5 100% $0.020
9. Nano Banana 2 3.77 / 5 100% $0.080
10. FLUX.2 [pro] 3.75 / 5 100% $0.030
11. Seedream 5.0 Lite 3.75 / 5 100% $0.035
12. GPT Image 2.5 Flare 3.75 / 5 100% $0.042
13. GPT Image 2.5 Sunburst 3.73 / 5 100% $0.036
14. Seedream 5.0 Pro 3.56 / 5 100% $0.135
15. Grok Imagine Image 2.0 3.53 / 5 100% $0.060
16. Luma Photon 3.50 / 5 100% $0.021
17. GPT Image 1.5 3.48 / 5 100% $0.140
18. Nano Banana Pro 3.45 / 5 100% $0.150
19. Ideogram 3.0 3.45 / 5 100% $0.060
20. Recraft V3 3.36 / 5 100% $0.040
21. Qwen Image 3.0 3.36 / 5 100% $0.040
22. GPT Image 2 3.29 / 5 100% $0.155
23. Stable Diffusion 1.5 3.19 / 5 100% $0.002
Want the raw numbers instead? See the full model rankings
Methodology

How this is weighted

A person is watching the spinner, so latency leads and cost follows: consumer products generate on every tap and pay for every retry. Prompt adherence weighs the same as cost because a result that ignores the prompt is a retry. The fidelity criteria stay light: UGC is judged by its creator, not by a brand team.

Weights are provisional (pilot-0) and published in full; the reviewed matrix ships with the frozen suite. Read the methodology
  • Latency 25%
  • Cost 20%
  • Prompt Adherence 20%
  • Composition 10%
  • Text Accuracy 10%
  • Color Accuracy 10%
  • Resolution 5%
Evidence

The evidence: FLUX.2

FLUX.2 takes the top spot where this job puts its weight: latency at 4.8 / 5 (25% of the matrix), cost at 4.9 / 5 (20% of the matrix) and prompt adherence at 4.4 / 5 (20% of the matrix).

The best-fit model's results on the benchmarks that matter for this use case. Click through for the full cross-model comparison.
Takeaway

The last mile: what no model delivers

In-app generation is where the editing canvas stops being a nice-to-have and becomes the product surface: the model proposes, the user refines. The data supports routing for speed here, since the fastest tier returns in around two seconds and still ranks seventh and ninth overall, while slow endpoints run over thirty. What no model provides is the refinement step itself: swapping the headline, nudging the layout, applying the user's colors. That loop is the canvas, and it is also what turns one generation into a kept, personalized design instead of a re-roll.

AI Editor

Generate with any model. Finish in the AI Editor.

Give your users the control they need to be productive with generative AI. The IMG.LY AI Editor turns in-app ugc creation generations into finished, on-spec assets: background removal, brand kits and editable text on a real canvas.

IMG.LY AI Editor: generative image tools next to manual editing controls on a canvas