Benchmark prompts
Composition & Negative Space

Centered Product Margin

Luxury ad layouts rely on symmetry and generous margins around a single hero. This tests compositional control in the opposite direction from N01: precise centering and empty space on all sides for framing and copy.

The exact prompt, sent to every model:
Perfectly centered symmetrical photo of a wristwatch on deep green velvet, generous empty margins on all sides, luxury ad composition.
Suite pilot-0 · identical prompt and default parameters for every model · seeds 1111, 2222, 3333 (or repeated samples for seedless APIs)

Results

Click any image to enlarge. Checkerboard backing shows true transparency; a rendered checkerboard pattern inside the image is a model faking it.

FLUX.2

FLUX.2 [pro]

FLUX.2 [dev] Turbo

Gemini 2.5 Flash Image

GPT Image 1.5

Ideogram 3.0

Luma Photon

Nano Banana 2

Nano Banana 2 Lite

Nano Banana Pro

Qwen-Image

Recraft V3

Seedream 4.5

Seedream 5.0 Lite

Stable Diffusion 1.5

Run-to-run consistency

Same model, same prompt, different seed or sample. Click a card to flip between the two generations: the less it moves, the safer the model is for templated, variable-data production.

Scores

Mean per model across this prompt's runs. Auto-measured criteria plus an interim Claude-VLM tier for prompt adherence and composition; any criterion still awaiting review is labeled pending, never zeroed.
  Prompt Adherence Composition Resolution Latency Cost
FLUX.2 4.6 / 5 4.7 / 5 2.0 / 5 4.0 / 5 4.0 / 5
FLUX.2 [pro] 4.6 / 5 4.7 / 5 2.0 / 5 1.7 / 5 4.0 / 5
FLUX.2 [dev] Turbo 5.0 / 5 5.0 / 5 2.0 / 5 4.0 / 5 4.0 / 5
Gemini 2.5 Flash Image 4.6 / 5 4.7 / 5 3.0 / 5 3.0 / 5 4.0 / 5
GPT Image 1.5 5.0 / 5 5.0 / 5 3.0 / 5 1.0 / 5 3.0 / 5
Ideogram 3.0 2.9 / 5 3.3 / 5 3.0 / 5 2.0 / 5 3.0 / 5
Luma Photon 4.2 / 5 4.0 / 5 4.0 / 5 1.7 / 5 4.0 / 5
Nano Banana 2 5.0 / 5 5.0 / 5 3.0 / 5 2.0 / 5 3.0 / 5
Nano Banana 2 Lite 4.6 / 5 4.7 / 5 3.0 / 5 4.0 / 5 4.0 / 5
Nano Banana Pro 4.6 / 5 4.3 / 5 3.0 / 5 1.0 / 5 1.0 / 5
Qwen-Image 5.0 / 5 5.0 / 5 2.0 / 5 3.0 / 5 4.0 / 5
Recraft V3 2.5 / 5 3.0 / 5 3.0 / 5 3.0 / 5 4.0 / 5
Seedream 4.5 4.2 / 5 4.0 / 5 5.0 / 5 1.7 / 5 4.0 / 5
Seedream 5.0 Lite 4.6 / 5 4.7 / 5 5.0 / 5 1.0 / 5 4.0 / 5
Stable Diffusion 1.5 2.1 / 5 1.7 / 5 1.0 / 5 4.0 / 5 5.0 / 5

Frequently asked questions

Image generation is stochastic: the same prompt produces different images on every run, so a single sample measures luck, not ability. Every model runs each benchmark three times with fixed seeds (1111, 2222, 3333), or three unseeded samples where the API accepts no seed. Scores average all three samples, and the consistency benchmarks measure the variation itself.

Yes. Every model receives the same prompt text with default parameters and no per-model tuning, so differences in output reflect the model, not prompt engineering. The suite is versioned and prompts are append-only, which keeps historical scores comparable.

The results on this page are scored on this exact prompt, so the strongest model for it is easy to spot; this benchmark sits in the Composition & Negative Space category. Scores are per model version and reflect this benchmark only. For a ranking across every benchmark see the model rankings, and to weight the numbers by a specific job see the use-case pages.

Each image is scored 0 to 5 per criterion. Measured criteria (resolution, latency, cost, transparency, color accuracy) are computed automatically; quality criteria (text accuracy, prompt adherence, composition) are judged by an automated Claude vision tier against each prompt’s checklist. Page scores are unweighted means over all of a model’s runs in that scope. Blind expert-panel review has not run yet; the dataset is pilot-0.

Model output is a starting point, not a finished asset. Production work usually needs background removal, exact brand colors or editable text, none of which generation guarantees on every run. The IMG.LY AI Editor gives your users those controls to refine any model’s output to production quality.

The suite re-runs on notable model releases so the rankings stay current. The results shown are the pilot-0 dataset, scored by measured criteria plus an automated Claude vision tier, with blind expert review planned. Prompts are versioned and append-only, so scores stay comparable across runs.