Mid-Century Travel Poster
Art-directed style adherence is what asset-pack and agency work demands: a specific era, flat shapes, a limited palette and screen-print texture, held together. This tests whether a model can commit to a named style instead of defaulting to generic realism.
Mid-century travel poster of Lisbon: tram, tiled facades and hills, flat shapes, warm limited palette, screen-print texture, 1950s style.
Results
FLUX.2
FLUX.2 [pro]
FLUX.2 [dev] Turbo
Gemini 2.5 Flash Image
GPT Image 1.5
Ideogram 3.0
Luma Photon
Nano Banana 2
Nano Banana 2 Lite
Nano Banana Pro
Qwen-Image
Recraft V3
Seedream 4.5
Seedream 5.0 Lite
Stable Diffusion 1.5
Run-to-run consistency
Scores
| Prompt Adherence | Resolution | Latency | Cost | |
|---|---|---|---|---|
| FLUX.2 | 5.0 / 5 | 2.0 / 5 | 4.0 / 5 | 4.0 / 5 |
| FLUX.2 [pro] | 5.0 / 5 | 2.0 / 5 | 2.0 / 5 | 4.0 / 5 |
| FLUX.2 [dev] Turbo | 5.0 / 5 | 2.0 / 5 | 4.0 / 5 | 4.0 / 5 |
| Gemini 2.5 Flash Image | 5.0 / 5 | 3.0 / 5 | 3.0 / 5 | 4.0 / 5 |
| GPT Image 1.5 | 5.0 / 5 | 3.0 / 5 | 1.0 / 5 | 3.0 / 5 |
| Ideogram 3.0 | 4.2 / 5 | 3.0 / 5 | 2.0 / 5 | 3.0 / 5 |
| Luma Photon | 5.0 / 5 | 4.0 / 5 | 1.3 / 5 | 4.0 / 5 |
| Nano Banana 2 | 5.0 / 5 | 3.0 / 5 | 2.0 / 5 | 3.0 / 5 |
| Nano Banana 2 Lite | 5.0 / 5 | 3.0 / 5 | 4.0 / 5 | 4.0 / 5 |
| Nano Banana Pro | 5.0 / 5 | 3.0 / 5 | 1.0 / 5 | 1.0 / 5 |
| Qwen-Image | 5.0 / 5 | 2.0 / 5 | 3.0 / 5 | 4.0 / 5 |
| Recraft V3 | 1.3 / 5 | 3.0 / 5 | 3.0 / 5 | 4.0 / 5 |
| Seedream 4.5 | 5.0 / 5 | 5.0 / 5 | 1.7 / 5 | 4.0 / 5 |
| Seedream 5.0 Lite | 4.6 / 5 | 5.0 / 5 | 1.0 / 5 | 4.0 / 5 |
| Stable Diffusion 1.5 | 4.2 / 5 | 1.0 / 5 | 4.0 / 5 | 5.0 / 5 |
New model? New benchmarks.
We re-run the identical 37-prompt suite on every model release. Get the scores, and what changed in the rankings, in your inbox.
You're on the list.
We'll email you when new benchmark results ship.
Frequently asked questions
Image generation is stochastic: the same prompt produces different images on every run, so a single sample measures luck, not ability. Every model runs each benchmark three times with fixed seeds (1111, 2222, 3333), or three unseeded samples where the API accepts no seed. Scores average all three samples, and the consistency benchmarks measure the variation itself.
Yes. Every model receives the same prompt text with default parameters and no per-model tuning, so differences in output reflect the model, not prompt engineering. The suite is versioned and prompts are append-only, which keeps historical scores comparable.
The results on this page are scored on this exact prompt, so the strongest model for it is easy to spot; this benchmark sits in the Style Adherence & Art Direction category. Scores are per model version and reflect this benchmark only. For a ranking across every benchmark see the model rankings, and to weight the numbers by a specific job see the use-case pages.
Each image is scored 0 to 5 per criterion. Measured criteria (resolution, latency, cost, transparency, color accuracy) are computed automatically; quality criteria (text accuracy, prompt adherence, composition) are judged by an automated Claude vision tier against each prompt’s checklist. Page scores are unweighted means over all of a model’s runs in that scope. Blind expert-panel review has not run yet; the dataset is pilot-0.
Model output is a starting point, not a finished asset. Production work usually needs background removal, exact brand colors or editable text, none of which generation guarantees on every run. The IMG.LY AI Editor gives your users those controls to refine any model’s output to production quality.
The suite re-runs on notable model releases so the rankings stay current. The results shown are the pilot-0 dataset, scored by measured criteria plus an automated Claude vision tier, with blind expert review planned. Prompts are versioned and append-only, so scores stay comparable across runs.

