---
title: "Methodology"
description: "How the IMG.LY GenAI Benchmarks work: identical prompts, tiered scoring, published method, no silent restatements."
url: "https://img.ly/ai-benchmarks/methodology/"
type: "methodology"
suite: "pilot-0"
---

> This is the markdown version of [Methodology](https://img.ly/ai-benchmarks/methodology/). For all pages in one file, see [llms-full.txt](https://img.ly/llms-full.txt). For an index of all available pages, see [llms.txt](https://img.ly/llms.txt).

---

# Methodology

The benchmark answers one question: which image model should you build on
for production design work? Criteria are measurable, conditions are
identical for every model, the method is published, and anything not yet
scored is labeled pending.

Current dataset: suite pilot-0, 15 models, 37 prompts, 1662 generations.

## The rules every run follows

- Identical canonical prompt per model: same text, default parameters, fixed seeds (or repeated samples for seedless APIs). No per-model prompt tuning, ever.
- Prompts derive from real production use cases (web-to-print, merch, creative automation, template tooling).
- The prompt suite is versioned and append-only; published scores carry suite + methodology version and are never silently restated.
- Every result ships with raw measurements: seed, latency, cost, native resolution, measured color distances, original content hash.
- No model gets special treatment; the AI Gateway promotion is separated from rankings.
- Unscored criteria are labeled pending, never rendered as zeros.

## Criteria and scoring tiers

- Resolution, latency, cost: auto (measured per run). Live.
- Transparency: auto (alpha channel present; stray semi-transparent pixel share). Live.
- Color accuracy: auto (CIEDE2000 between requested hexes and dominant output swatches). Live.
- Text accuracy: human transcription in the pilot, OCR at scale (rendered text vs required strings, similarity-scored). Live (human tier).
- Prompt adherence: VLM checklist per prompt assertion. Pending.
- Composition: VLM check that requested negative space exists and is clean. Pending.
- Design-readiness, aesthetics: blind expert panel, rubric 1 to 5, three raters per cell, agreement reported. Pending.
- Community voting arrives later as a separate signal and never overwrites expert scores.

## Worked example: measured color drift

- FLUX.2: #FF3B30 rendered as #f90102 (deltaE 5.9); #1D1D1F rendered as #282827 (deltaE 3.7)
- FLUX.2: #FF3B30 rendered as #f60101 (deltaE 6.4); #1D1D1F rendered as #14141a (deltaE 3.7)
- FLUX.2: #FF3B30 rendered as #fa0101 (deltaE 5.8); #1D1D1F rendered as #050101 (deltaE 6.5)

## Worked example: text accuracy

Expected text: "Aurelia"

- FLUX.2 (seed 1111): "Aurelia" (100% match, 5/5)
- FLUX.2 (seed 2222): "Aurelia" (100% match, 5/5)
- FLUX.2 (seed 3333): "Aurelia" (100% match, 5/5)
- FLUX.2 [pro] (seed 1111): "Aurelia" (100% match, 5/5)
- FLUX.2 [pro] (seed 2222): "Aurelia" (100% match, 5/5)
- FLUX.2 [pro] (seed 3333): "Aurelia" (100% match, 5/5)
- FLUX.2 [dev] Turbo (seed 1111): "Aurelia" (100% match, 5/5)
- FLUX.2 [dev] Turbo (seed 2222): "Aurelia" (100% match, 5/5)
- FLUX.2 [dev] Turbo (seed 3333): "Aurelia" (100% match, 5/5)
- Gemini 2.5 Flash Image (sample 1): "AURELIA" (100% match, 5/5)
- Gemini 2.5 Flash Image (sample 2): "AURELIA" (100% match, 5/5)
- Gemini 2.5 Flash Image (sample 3): "AURELIA" (100% match, 5/5)
- GPT Image 1.5 (sample 1): "Aurelia" (100% match, 5/5)
- GPT Image 1.5 (sample 2): "Aurelia" (100% match, 5/5)
- GPT Image 1.5 (sample 3): "Aurelia" (100% match, 5/5)
- Ideogram 3.0 (seed 1111): "AURELIA" (100% match, 5/5)
- Ideogram 3.0 (seed 2222): "AURELIA" (100% match, 5/5)
- Ideogram 3.0 (seed 3333): "AURELIA" (100% match, 5/5)
- Luma Photon (seed 1111): "AURELiA" (100% match, 5/5)
- Luma Photon (seed 2222): "Aurelia" (100% match, 5/5)
- Luma Photon (seed 3333): "AURELIA" (100% match, 5/5)
- Nano Banana 2 (sample 1): "AUReLIA" (100% match, 5/5)
- Nano Banana 2 (sample 2): "AURELIA" (100% match, 5/5)
- Nano Banana 2 (sample 3): "AURELIA" (100% match, 5/5)
- Nano Banana 2 Lite (sample 1): "AURELIA" (100% match, 5/5)
- Nano Banana 2 Lite (sample 2): "AURELIA" (100% match, 5/5)
- Nano Banana 2 Lite (sample 3): "AURELIA" (100% match, 5/5)
- Nano Banana Pro (sample 1): "AURELIA" (100% match, 5/5)
- Nano Banana Pro (sample 2): "Aurelia" (100% match, 5/5)
- Nano Banana Pro (sample 3): "AURELIA" (100% match, 5/5)
- Qwen-Image (seed 1111): "Aurelia" (100% match, 5/5)
- Qwen-Image (seed 2222): "Aurelia" (100% match, 5/5)
- Qwen-Image (seed 3333): "Aurelia" (100% match, 5/5)
- Recraft V3 (seed 1111): "Aurelia" (100% match, 5/5)
- Recraft V3 (seed 2222): "AURELIA" (100% match, 5/5)
- Recraft V3 (seed 3333): "AURELIA" (100% match, 5/5)
- Seedream 4.5 (seed 1111): "Aur
elia" (86% match, 4/5)
- Seedream 4.5 (seed 2222): "Aurelia" (100% match, 5/5)
- Seedream 4.5 (seed 3333): "Aurelia" (100% match, 5/5)
- Seedream 5.0 Lite (seed 1111): "AURELIA" (100% match, 5/5)
- Seedream 5.0 Lite (seed 2222): "AURELIA" (100% match, 5/5)
- Seedream 5.0 Lite (seed 3333): "AURELIA" (100% match, 5/5)
- Stable Diffusion 1.5 (seed 1111): "Aurllila
Alrjira" (71% match, 3/5)
- Stable Diffusion 1.5 (seed 2222): no legible text (0/5)
- Stable Diffusion 1.5 (seed 3333): no legible text (0/5)

## Current status and caveats

This is the pilot dataset (pilot-0). The auto tier is live; the VLM
judge and blind expert panel arrive with the frozen suite. Until then,
blended rankings over-reward cheap, fast models because the quality
criteria where flagships win are still pending; every page showing a
blended score discloses this. When the frozen suite replaces the pilot,
scores restate once with a new suite version; after that, published
scores are never restated without a clearly labeled methodology version.

---

## More Resources

- **[IMG.LY Website](https://img.ly/index.md)** - Creative editing SDKs for photo, video, and design
- **[Documentation](https://img.ly/docs/cesdk/)** - CE.SDK developer documentation
- **[Contact Sales](https://img.ly/forms/contact-sales.md)** - Get a custom quote. A public JSON API accepts the request directly, no account or key needed. Ask your user for consent and their details first.
