Benchmark prompts
Consistency & Repeatability

Character Series: Desk

Run two of the character-consistency set. Holding the same face, hair, apron and glasses across a new pose and setting is exactly what breaks when teams try to generate a coherent illustrated series.

Use cases:
Template & asset packs
Personalization at scale
The exact prompt, sent to every model:
Maya, a red-haired botanist in a green apron and round glasses, reading a field guide at a wooden desk, consistent character design, storybook illustration style.
Suite pilot-0 · identical prompt and default parameters for every model · seeds 1111, 2222, 3333 (or repeated samples for seedless APIs)

Results

Click any image to enlarge. Checkerboard backing shows true transparency; a rendered checkerboard pattern inside the image is a model faking it.

FLUX.2

FLUX.2 [pro]

FLUX.2 [dev] Turbo

Gemini 2.5 Flash Image

GPT Image 1.5

Ideogram 3.0

Luma Photon

Nano Banana 2

Nano Banana 2 Lite

Nano Banana Pro

Qwen-Image

Recraft V3

Seedream 4.5

Seedream 5.0 Lite

Stable Diffusion 1.5

Run-to-run consistency

Same model, same prompt, different seed or sample. Click a card to flip between the two generations: the less it moves, the safer the model is for templated, variable-data production.

Scores

Mean per model across this prompt's runs. Auto-measured criteria plus an interim Claude-VLM tier for prompt adherence and composition; any criterion still awaiting review is labeled pending, never zeroed.
  Prompt Adherence Resolution Latency Cost
FLUX.2 3.8 / 5 2.0 / 5 4.0 / 5 4.0 / 5
FLUX.2 [pro] 3.8 / 5 2.0 / 5 2.0 / 5 4.0 / 5
FLUX.2 [dev] Turbo 3.8 / 5 2.0 / 5 4.0 / 5 4.0 / 5
Gemini 2.5 Flash Image 3.8 / 5 3.0 / 5 2.7 / 5 4.0 / 5
GPT Image 1.5 3.8 / 5 3.0 / 5 1.0 / 5 3.0 / 5
Ideogram 3.0 3.8 / 5 3.0 / 5 2.0 / 5 3.0 / 5
Luma Photon 3.8 / 5 4.0 / 5 2.0 / 5 4.0 / 5
Nano Banana 2 3.8 / 5 3.0 / 5 2.0 / 5 3.0 / 5
Nano Banana 2 Lite 3.8 / 5 3.0 / 5 3.7 / 5 4.0 / 5
Nano Banana Pro 3.8 / 5 3.0 / 5 1.0 / 5 1.0 / 5
Qwen-Image 3.8 / 5 2.0 / 5 3.0 / 5 4.0 / 5
Recraft V3 2.5 / 5 3.0 / 5 3.0 / 5 4.0 / 5
Seedream 4.5 3.8 / 5 5.0 / 5 2.0 / 5 4.0 / 5
Seedream 5.0 Lite 3.8 / 5 5.0 / 5 1.0 / 5 4.0 / 5
Stable Diffusion 1.5 3.8 / 5 1.0 / 5 4.0 / 5 5.0 / 5

Frequently asked questions

Image generation is stochastic: the same prompt produces different images on every run, so a single sample measures luck, not ability. Every model runs each benchmark three times with fixed seeds (1111, 2222, 3333), or three unseeded samples where the API accepts no seed. Scores average all three samples, and the consistency benchmarks measure the variation itself.

Yes. Every model receives the same prompt text with default parameters and no per-model tuning, so differences in output reflect the model, not prompt engineering. The suite is versioned and prompts are append-only, which keeps historical scores comparable.

The results on this page are scored on this exact prompt, so the strongest model for it is easy to spot; this benchmark sits in the Consistency & Repeatability category. Scores are per model version and reflect this benchmark only. For a ranking across every benchmark see the model rankings, and to weight the numbers by a specific job see the use-case pages.

Each image is scored 0 to 5 per criterion. Measured criteria (resolution, latency, cost, transparency, color accuracy) are computed automatically; quality criteria (text accuracy, prompt adherence, composition) are judged by an automated Claude vision tier against each prompt’s checklist. Page scores are unweighted means over all of a model’s runs in that scope. Blind expert-panel review has not run yet; the dataset is pilot-0.

Model output is a starting point, not a finished asset. Production work usually needs background removal, exact brand colors or editable text, none of which generation guarantees on every run. The IMG.LY AI Editor gives your users those controls to refine any model’s output to production quality.

The suite re-runs on notable model releases so the rankings stay current. The results shown are the pilot-0 dataset, scored by measured criteria plus an automated Claude vision tier, with blind expert review planned. Prompts are versioned and append-only, so scores stay comparable across runs.