GenAI Benchmarks

Methodology

The benchmark answers one question: which image model should you build on for production design work? Everything below exists so the answer is defensible: measurable criteria, identical conditions, published method, and honest labels on anything not yet scored.

Current dataset: suite pilot-0 · 15 models · 37 prompts · 1662 generations · last run Jul 6, 2026

The rules every run follows

  • Every model gets the identical canonical prompt: same text, default parameters, fixed seeds (or repeated samples where an API is seedless). Prompt tuning per model is never allowed.
  • Prompts derive from real production use cases (web-to-print, merch, creative automation, template tooling), not from what makes models look good.
  • The prompt suite is versioned and append-only; published scores carry their suite and methodology version and are never silently restated.
  • Every result ships with its raw measurements: seed, latency, cost, native resolution, measured color distances, and the original file's content hash.
  • No model gets special treatment. IMG.LY partner models run under exactly the same conditions, and the AI Gateway promotion stays visually separated from rankings.
  • Anything not yet scored is labeled pending: it is never rendered as a zero and never silently dropped from averages.

Criteria and scoring tiers

Each criterion is scored by the cheapest tier that is reliable for it. Community voting arrives later as a separate signal and never overwrites expert scores.
  Tier How it is measured Pilot status
Resolution Auto Native output dimensions per run Live
Latency Auto Submit-to-complete wall clock, p50 reported Live
Cost Auto Provider price per generation Live
Transparency Auto Alpha channel present; stray semi-transparent pixel share Live
Color accuracy Auto CIEDE2000 between requested hexes and dominant swatches Live
Text accuracy Human, then OCR Rendered text vs required strings, similarity-scored Human transcription (pilot)
Prompt adherence VLM Pass/fail checklist per prompt assertion Claude-VLM (pilot)
Composition VLM Requested negative space exists and is clean Claude-VLM (pilot)
Design-readiness Expert panel Blind rubric 1–5: "ship this asset with ≤5 min cleanup?" pending

Worked example: measured color drift

The two-color illustration benchmark requests exact brand hexes. We extract each output's dominant swatches and compute the CIEDE2000 distance to the nearest one. Under 2 is imperceptible; above 10 reads as a different color.

FLUX.2 · seed 1111

ΔE 5.9 visible shift

#ff3b30 → #f90102

ΔE 3.7 close

#1d1d1f → #282827

FLUX.2 · seed 2222

ΔE 6.4 visible shift

#ff3b30 → #f60101

ΔE 3.7 close

#1d1d1f → #14141a

FLUX.2 · seed 3333

ΔE 5.8 visible shift

#ff3b30 → #fa0101

ΔE 6.5 visible shift

#1d1d1f → #050101

ΔE2000 between each requested brand color and the nearest dominant swatch in the output. Under 2 is imperceptible; above 10 is a clearly different color.

Worked example: text accuracy

Typography benchmarks require exact strings. Rendered text is transcribed (by a human reviewer in the pilot; OCR at full-suite scale) and similarity-scored against the requirement, character by character.

Expected

Aurelia

Rendered

  • FLUX.2 · seed 1111 A u r e l i a 100% match · 5/5
  • FLUX.2 [pro] · seed 1111 A u r e l i a 100% match · 5/5
  • FLUX.2 [dev] Turbo · seed 1111 A u r e l i a 100% match · 5/5
  • Gemini 2.5 Flash Image · sample 1 A U R E L I A 100% match · 5/5
  • GPT Image 1.5 · sample 1 A u r e l i a 100% match · 5/5
  • Ideogram 3.0 · seed 1111 A U R E L I A 100% match · 5/5
  • Luma Photon · seed 1111 A U R E L i A 100% match · 5/5
  • Nano Banana 2 · sample 1 A U R e L I A 100% match · 5/5
  • Nano Banana 2 Lite · sample 1 A U R E L I A 100% match · 5/5
  • Nano Banana Pro · sample 1 A U R E L I A 100% match · 5/5
  • Qwen-Image · seed 1111 A u r e l i a 100% match · 5/5
  • Recraft V3 · seed 1111 A u r e l i a 100% match · 5/5
  • Seedream 4.5 · seed 1111 A u r e l i a 86% match · 4/5
  • Seedream 5.0 Lite · seed 1111 A U R E L I A 100% match · 5/5
  • Stable Diffusion 1.5 · seed 1111 A u r l l i l a A l r j i r a 71% match · 3/5
Full comparison: Wordmark benchmark

Worked example: transparency

The die-cut sticker benchmark requests a transparent background. The auto tier checks for a real alpha channel and stray semi-transparent pixels. Swap the backing below: in the pilot, no model shipped real alpha, and some render a fake checkerboard into the pixels.
Background:

Current status and caveats

This is the pilot dataset (pilot-0): a deliberately small run that exercises the full pipeline end to end. The auto tier is live, and the VLM-judged quality criteria (prompt adherence, composition) are scored by an interim Claude-VLM tier, not yet the blind expert panel that arrives with the frozen suite. Even with quality scored, the blended overall is an unweighted mean, so it over-rewards cheap, fast models and the premium flagship still lands last on it. That is why the use-case pages reweight the same measurements by job, and every page that shows a blended score says so.

When the frozen suite replaces the pilot, scores restate once, the suite version changes, and from that point published scores are never restated without a new, clearly labeled methodology version.