Methodology
The benchmark answers one question: which image model should you build on for production design work? Everything below exists so the answer is defensible: measurable criteria, identical conditions, published method, and honest labels on anything not yet scored.
The rules every run follows
- Every model gets the identical canonical prompt: same text, default parameters, fixed seeds (or repeated samples where an API is seedless). Prompt tuning per model is never allowed.
- Prompts derive from real production use cases (web-to-print, merch, creative automation, template tooling), not from what makes models look good.
- The prompt suite is versioned and append-only; published scores carry their suite and methodology version and are never silently restated.
- Every result ships with its raw measurements: seed, latency, cost, native resolution, measured color distances, and the original file's content hash.
- No model gets special treatment. IMG.LY partner models run under exactly the same conditions, and the AI Gateway promotion stays visually separated from rankings.
- Anything not yet scored is labeled pending: it is never rendered as a zero and never silently dropped from averages.
Criteria and scoring tiers
| Tier | How it is measured | Pilot status | |
|---|---|---|---|
| Resolution | Auto | Native output dimensions per run | Live |
| Latency | Auto | Submit-to-complete wall clock, p50 reported | Live |
| Cost | Auto | Provider price per generation | Live |
| Transparency | Auto | Alpha channel present; stray semi-transparent pixel share | Live |
| Color accuracy | Auto | CIEDE2000 between requested hexes and dominant swatches | Live |
| Text accuracy | VLM, then OCR | Rendered text vs required strings, similarity-scored | Claude-VLM transcription (pilot) |
| Prompt adherence | VLM | Pass/fail checklist per prompt assertion | Claude-VLM (pilot) |
| Composition | VLM | Requested negative space exists and is clean | Claude-VLM (pilot) |
| Design-readiness | Expert panel | Blind rubric 1–5: "ship this asset with ≤5 min cleanup?" | pending |
What each score means
| What is measured | 5 | 4 | 3 | 2 | 1 | |
|---|---|---|---|---|---|---|
| Resolution | Megapixels, higher is better | 4 MP and up | 2 MP | 1 MP | 0.5 MP | below 0.5 MP |
| Cost | USD per image, lower is better | up to $0.01 | $0.05 | $0.15 | $0.30 | $0.60 and above |
| Latency | Submit to complete, lower is better | up to 2s | 5s | 10s | 30s | 150s and above |
| Color accuracy | Mean ΔE2000, lower is better | up to 2 | 5 | 10 | 20 | not awarded, above 20 scores 0 |
| Transparency | Alpha channel of the file | clean cutout, under 10% partial | not awarded | 10% to 25% partial | over 25% partial | channel present, image opaque |
| Text accuracy | Similarity to the required string | 99.5% or better | 85% or better | 65% or better | 40% or better | anything above 0 |
| Adherence, composition | Judged per prompt, not measured | — | — | — | — | — |
Cost and latency are scored continuously between the anchors above, because both span orders of magnitude and a fixed threshold in the middle of that range would let a fraction of a cent decide a whole point. The remaining measured criteria keep fixed bands, and those are deliberately coarse: on resolution, everything from 2 to 4 megapixels scores the same 4. Read the price and size columns on the model pages, not just the score.
Worked example: measured color drift
FLUX.2 · seed 1111
ΔE 5.9 visible shift
#ff3b30 → #f90102
ΔE 3.7 close
#1d1d1f → #282827
FLUX.2 · seed 2222
ΔE 6.4 visible shift
#ff3b30 → #f60101
ΔE 3.7 close
#1d1d1f → #14141a
FLUX.2 · seed 3333
ΔE 5.8 visible shift
#ff3b30 → #fa0101
ΔE 6.5 visible shift
#1d1d1f → #050101
ΔE2000 between each requested brand color and the nearest dominant swatch in the output. Under 2 is imperceptible; above 10 is a clearly different color.
Worked example: text accuracy
Expected
Aurelia
Rendered
- FLUX.2 · seed 1111 A u r e l i a 100% match · 5/5
- FLUX.2 [pro] · seed 1111 A u r e l i a 100% match · 5/5
- FLUX.2 [dev] Turbo · seed 1111 A u r e l i a 100% match · 5/5
- Gemini 2.5 Flash Image · sample 1 A U R E L I A 100% match · 5/5
- GPT Image 1.5 · sample 1 A u r e l i a 100% match · 5/5
- GPT Image 2 · sample 1 A u r e l i a 100% match · 5/5
- GPT Image 2.5 Flare · sample 1 A u r e l i a 100% match · 5/5
- GPT Image 2.5 Sunburst · sample 1 A u r e l i a 100% match · 5/5
- Grok Imagine Image 2.0 · sample 1 A u r e l i a 100% match · 5/5
- Ideogram 3.0 · seed 1111 A U R E L I A 100% match · 5/5
- Luma Photon · seed 1111 A U R E L i A 100% match · 5/5
- Muse Image · sample 1 A u r e l i a 100% match · 5/5
- Nano Banana 2 · sample 1 A U R e L I A 100% match · 5/5
- Nano Banana 2 Lite · sample 1 A U R E L I A 100% match · 5/5
- Nano Banana Pro · sample 1 A U R E L I A 100% match · 5/5
- Qwen-Image · seed 1111 A u r e l i a 100% match · 5/5
- Qwen Image 3.0 · seed 1111 A u r e l i a 100% match · 5/5
- Recraft V3 · seed 1111 A u r e l i a 100% match · 5/5
- Recraft V4.1 · sample 1 A u r e l i a 100% match · 5/5
- Seedream 4.5 · seed 1111 A u r e l i a 86% match · 4/5
- Seedream 5.0 Lite · seed 1111 A U R E L I A 100% match · 5/5
- Seedream 5.0 Pro · seed 1111 A u r e l i a 100% match · 5/5
- Stable Diffusion 1.5 · seed 1111 A u r l l i l a A l r j i r a 71% match · 3/5
Worked example: transparency
Methodology revisions
pilot-1 (September 14, 2026). The current recipe, and the one behind every score on the site today. Three things changed against pilot-0. Cost and latency moved from stepped bands to the continuous log scale described above, so a fraction of a cent no longer decides a whole point. The blended overall and category scores now give each criterion one equal vote instead of averaging raw score rows. Cost figures were restated from provider usage exports: the earlier per-image estimates were wrong for 15 of 23 models, several by a factor of two to four, and every stored cost now records the billing unit, output size and date it was verified at. All earlier runs were recomputed under this recipe alongside the September model additions.
pilot-0 (July 2026). The initial recipe of the pilot: stepped bands for every measured criterion, estimated per-image costs, and a blended overall averaged over raw score rows.
Current status and caveats
This is the pilot dataset (pilot-0): a deliberately small run that exercises the full pipeline end to end. The auto tier is live, and the VLM-judged quality criteria (prompt adherence, composition) are scored by an interim Claude-VLM tier, not yet the blind expert panel that arrives with the frozen suite. The blended overall gives each of the eight criteria one equal vote, rather than averaging the raw measurements, which would let the four criteria scored on every prompt outweigh the four scored on a subset of them. One equal vote each is still a generic ranking, which is why the use-case pages reweight the same measurements by job.
When the frozen suite replaces the pilot, scores restate once, the suite version changes, and from that point published scores are never restated without a new, clearly labeled methodology version.

