GenAI Benchmarks

Methodology

The benchmark answers one question: which image model should you build on for production design work? Everything below exists so the answer is defensible: measurable criteria, identical conditions, published method, and honest labels on anything not yet scored.

Current dataset: suite pilot-0 · methodology pilot-1 · 23 models · 37 prompts · 2550 generations · last run Sep 10, 2026 · scores computed Sep 14, 2026

The rules every run follows

  • Every model gets the identical canonical prompt: same text, default parameters, fixed seeds (or repeated samples where an API is seedless). Prompt tuning per model is never allowed.
  • Prompts derive from real production use cases (web-to-print, merch, creative automation, template tooling), not from what makes models look good.
  • The prompt suite is versioned and append-only; published scores carry their suite and methodology version and are never silently restated.
  • Every result ships with its raw measurements: seed, latency, cost, native resolution, measured color distances, and the original file's content hash.
  • No model gets special treatment. IMG.LY partner models run under exactly the same conditions, and the AI Gateway promotion stays visually separated from rankings.
  • Anything not yet scored is labeled pending: it is never rendered as a zero and never silently dropped from averages.

Criteria and scoring tiers

Each criterion is scored by the cheapest tier that is reliable for it. Community voting arrives later as a separate signal and never overwrites expert scores.
  Tier How it is measured Pilot status
Resolution Auto Native output dimensions per run Live
Latency Auto Submit-to-complete wall clock, p50 reported Live
Cost Auto Provider price per generation Live
Transparency Auto Alpha channel present; stray semi-transparent pixel share Live
Color accuracy Auto CIEDE2000 between requested hexes and dominant swatches Live
Text accuracy VLM, then OCR Rendered text vs required strings, similarity-scored Claude-VLM transcription (pilot)
Prompt adherence VLM Pass/fail checklist per prompt assertion Claude-VLM (pilot)
Composition VLM Requested negative space exists and is clean Claude-VLM (pilot)
Design-readiness Expert panel Blind rubric 1–5: "ship this asset with ≤5 min cleanup?" pending

What each score means

Five of the eight criteria are measured from the file and mapped onto 0-5. Three of them use the fixed thresholds below, and the two that span orders of magnitude, cost and latency, are interpolated continuously between the same values. Either way any measured score in this benchmark can be reproduced from the output itself. A score of 5 means the file cleared the top of the scale, not that it was the best in the field.
  What is measured 5 4 3 2 1
Resolution Megapixels, higher is better 4 MP and up 2 MP 1 MP 0.5 MP below 0.5 MP
Cost USD per image, lower is better up to $0.01 $0.05 $0.15 $0.30 $0.60 and above
Latency Submit to complete, lower is better up to 2s 5s 10s 30s 150s and above
Color accuracy Mean ΔE2000, lower is better up to 2 5 10 20 not awarded, above 20 scores 0
Transparency Alpha channel of the file clean cutout, under 10% partial not awarded 10% to 25% partial over 25% partial channel present, image opaque
Text accuracy Similarity to the required string 99.5% or better 85% or better 65% or better 40% or better anything above 0
Adherence, composition Judged per prompt, not measured — — — — —

Cost and latency are scored continuously between the anchors above, because both span orders of magnitude and a fixed threshold in the middle of that range would let a fraction of a cent decide a whole point. The remaining measured criteria keep fixed bands, and those are deliberately coarse: on resolution, everything from 2 to 4 megapixels scores the same 4. Read the price and size columns on the model pages, not just the score.

Worked example: measured color drift

The two-color illustration benchmark requests exact brand hexes. We extract each output's dominant swatches and compute the CIEDE2000 distance to the nearest one. Under 2 is imperceptible; above 10 reads as a different color.

FLUX.2 · seed 1111

ΔE 5.9 visible shift

#ff3b30 → #f90102

ΔE 3.7 close

#1d1d1f → #282827

FLUX.2 · seed 2222

ΔE 6.4 visible shift

#ff3b30 → #f60101

ΔE 3.7 close

#1d1d1f → #14141a

FLUX.2 · seed 3333

ΔE 5.8 visible shift

#ff3b30 → #fa0101

ΔE 6.5 visible shift

#1d1d1f → #050101

ΔE2000 between each requested brand color and the nearest dominant swatch in the output. Under 2 is imperceptible; above 10 is a clearly different color.

Worked example: text accuracy

Typography benchmarks require exact strings. Rendered text is transcribed (by the Claude vision judge in the pilot; OCR at full-suite scale) and each required string is scored independently by best-match containment, so legitimate extra text in the image costs nothing.

Expected

Aurelia

Rendered

  • FLUX.2 · seed 1111 A u r e l i a 100% match · 5/5
  • FLUX.2 [pro] · seed 1111 A u r e l i a 100% match · 5/5
  • FLUX.2 [dev] Turbo · seed 1111 A u r e l i a 100% match · 5/5
  • Gemini 2.5 Flash Image · sample 1 A U R E L I A 100% match · 5/5
  • GPT Image 1.5 · sample 1 A u r e l i a 100% match · 5/5
  • GPT Image 2 · sample 1 A u r e l i a 100% match · 5/5
  • GPT Image 2.5 Flare · sample 1 A u r e l i a 100% match · 5/5
  • GPT Image 2.5 Sunburst · sample 1 A u r e l i a 100% match · 5/5
  • Grok Imagine Image 2.0 · sample 1 A u r e l i a 100% match · 5/5
  • Ideogram 3.0 · seed 1111 A U R E L I A 100% match · 5/5
  • Luma Photon · seed 1111 A U R E L i A 100% match · 5/5
  • Muse Image · sample 1 A u r e l i a 100% match · 5/5
  • Nano Banana 2 · sample 1 A U R e L I A 100% match · 5/5
  • Nano Banana 2 Lite · sample 1 A U R E L I A 100% match · 5/5
  • Nano Banana Pro · sample 1 A U R E L I A 100% match · 5/5
  • Qwen-Image · seed 1111 A u r e l i a 100% match · 5/5
  • Qwen Image 3.0 · seed 1111 A u r e l i a 100% match · 5/5
  • Recraft V3 · seed 1111 A u r e l i a 100% match · 5/5
  • Recraft V4.1 · sample 1 A u r e l i a 100% match · 5/5
  • Seedream 4.5 · seed 1111 A u r e l i a 86% match · 4/5
  • Seedream 5.0 Lite · seed 1111 A U R E L I A 100% match · 5/5
  • Seedream 5.0 Pro · seed 1111 A u r e l i a 100% match · 5/5
  • Stable Diffusion 1.5 · seed 1111 A u r l l i l a A l r j i r a 71% match · 3/5
Full comparison: Wordmark benchmark

Worked example: transparency

The die-cut sticker benchmark requests a transparent background. The auto tier checks for a real alpha channel and stray semi-transparent pixels. Swap the backing below: almost no model ships real alpha, and some render a fake checkerboard into the pixels.
Background:

Methodology revisions

Every score row carries the version of the recipe that computed it and the date it was computed. When the recipe changes, the version changes with it, so a restated number is always labeled as restated.

pilot-1 (September 14, 2026). The current recipe, and the one behind every score on the site today. Three things changed against pilot-0. Cost and latency moved from stepped bands to the continuous log scale described above, so a fraction of a cent no longer decides a whole point. The blended overall and category scores now give each criterion one equal vote instead of averaging raw score rows. Cost figures were restated from provider usage exports: the earlier per-image estimates were wrong for 15 of 23 models, several by a factor of two to four, and every stored cost now records the billing unit, output size and date it was verified at. All earlier runs were recomputed under this recipe alongside the September model additions.

pilot-0 (July 2026). The initial recipe of the pilot: stepped bands for every measured criterion, estimated per-image costs, and a blended overall averaged over raw score rows.

Current status and caveats

This is the pilot dataset (pilot-0): a deliberately small run that exercises the full pipeline end to end. The auto tier is live, and the VLM-judged quality criteria (prompt adherence, composition) are scored by an interim Claude-VLM tier, not yet the blind expert panel that arrives with the frozen suite. The blended overall gives each of the eight criteria one equal vote, rather than averaging the raw measurements, which would let the four criteria scored on every prompt outweigh the four scored on a subset of them. One equal vote each is still a generic ranking, which is why the use-case pages reweight the same measurements by job.

When the frozen suite replaces the pilot, scores restate once, the suite version changes, and from that point published scores are never restated without a new, clearly labeled methodology version.