Methodology
The benchmark answers one question: which image model should you build on for production design work? Everything below exists so the answer is defensible: measurable criteria, identical conditions, published method, and honest labels on anything not yet scored.
The rules every run follows
- Every model gets the identical canonical prompt: same text, default parameters, fixed seeds (or repeated samples where an API is seedless). Prompt tuning per model is never allowed.
- Prompts derive from real production use cases (web-to-print, merch, creative automation, template tooling), not from what makes models look good.
- The prompt suite is versioned and append-only; published scores carry their suite and methodology version and are never silently restated.
- Every result ships with its raw measurements: seed, latency, cost, native resolution, measured color distances, and the original file's content hash.
- No model gets special treatment. IMG.LY partner models run under exactly the same conditions, and the AI Gateway promotion stays visually separated from rankings.
- Anything not yet scored is labeled pending: it is never rendered as a zero and never silently dropped from averages.
Criteria and scoring tiers
| Tier | How it is measured | Pilot status | |
|---|---|---|---|
| Resolution | Auto | Native output dimensions per run | Live |
| Latency | Auto | Submit-to-complete wall clock, p50 reported | Live |
| Cost | Auto | Provider price per generation | Live |
| Transparency | Auto | Alpha channel present; stray semi-transparent pixel share | Live |
| Color accuracy | Auto | CIEDE2000 between requested hexes and dominant swatches | Live |
| Text accuracy | Human, then OCR | Rendered text vs required strings, similarity-scored | Human transcription (pilot) |
| Prompt adherence | VLM | Pass/fail checklist per prompt assertion | Claude-VLM (pilot) |
| Composition | VLM | Requested negative space exists and is clean | Claude-VLM (pilot) |
| Design-readiness | Expert panel | Blind rubric 1–5: "ship this asset with ≤5 min cleanup?" | pending |
Worked example: measured color drift
FLUX.2 · seed 1111
ΔE 5.9 visible shift
#ff3b30 → #f90102
ΔE 3.7 close
#1d1d1f → #282827
FLUX.2 · seed 2222
ΔE 6.4 visible shift
#ff3b30 → #f60101
ΔE 3.7 close
#1d1d1f → #14141a
FLUX.2 · seed 3333
ΔE 5.8 visible shift
#ff3b30 → #fa0101
ΔE 6.5 visible shift
#1d1d1f → #050101
ΔE2000 between each requested brand color and the nearest dominant swatch in the output. Under 2 is imperceptible; above 10 is a clearly different color.
Worked example: text accuracy
Expected
Aurelia
Rendered
- FLUX.2 · seed 1111 A u r e l i a 100% match · 5/5
- FLUX.2 [pro] · seed 1111 A u r e l i a 100% match · 5/5
- FLUX.2 [dev] Turbo · seed 1111 A u r e l i a 100% match · 5/5
- Gemini 2.5 Flash Image · sample 1 A U R E L I A 100% match · 5/5
- GPT Image 1.5 · sample 1 A u r e l i a 100% match · 5/5
- Ideogram 3.0 · seed 1111 A U R E L I A 100% match · 5/5
- Luma Photon · seed 1111 A U R E L i A 100% match · 5/5
- Nano Banana 2 · sample 1 A U R e L I A 100% match · 5/5
- Nano Banana 2 Lite · sample 1 A U R E L I A 100% match · 5/5
- Nano Banana Pro · sample 1 A U R E L I A 100% match · 5/5
- Qwen-Image · seed 1111 A u r e l i a 100% match · 5/5
- Recraft V3 · seed 1111 A u r e l i a 100% match · 5/5
- Seedream 4.5 · seed 1111 A u r e l i a 86% match · 4/5
- Seedream 5.0 Lite · seed 1111 A U R E L I A 100% match · 5/5
- Stable Diffusion 1.5 · seed 1111 A u r l l i l a A l r j i r a 71% match · 3/5
Worked example: transparency
Current status and caveats
This is the pilot dataset (pilot-0): a deliberately small run that exercises the full pipeline end to end. The auto tier is live, and the VLM-judged quality criteria (prompt adherence, composition) are scored by an interim Claude-VLM tier, not yet the blind expert panel that arrives with the frozen suite. Even with quality scored, the blended overall is an unweighted mean, so it over-rewards cheap, fast models and the premium flagship still lands last on it. That is why the use-case pages reweight the same measurements by job, and every page that shows a blended score says so.
When the frozen suite replaces the pilot, scores restate once, the suite version changes, and from that point published scores are never restated without a new, clearly labeled methodology version.

