Compare AI image models
Pick up to three models and compare them on the same design-production prompts: overall score, per-criterion quality, cost and latency, side by side. Your selection is saved in the URL, so any comparison is shareable.
Select two or three models to compare.
pilot-0 dataset.
Featured head-to-heads
Hand-picked matchups where the trade-off is worth spelling out: the ones product teams actually decide between.
FLUX.2 vs Stable Diffusion 1.5
Still running Stable Diffusion in your stack? Flux 2 scores 4.08 to SD 1.5’s 1.83 on our suite and renders clean type where SD garbles it, at $0.013 per image via API. The self-hosting control question against a measured, two-times quality gap.
Nano Banana 2 vs Seedream 4.5
Two strong mid-cost routing targets. Nano Banana 2 edges ahead on quality (4.25 vs 4.05) and owns brand-color and sticker work; Seedream is the steadier generalist. The value decision for volume generation.
FLUX.2 vs Nano Banana 2
The two default modern routes. Flux 2 wins the most benchmarks outright (typography, product, print) at $0.013; Nano Banana 2 counters on brand-color fidelity, sticker work and isometric style. Most products end up routing between exactly these two.
FLUX.2 vs GPT Image 1.5
The premium-vs-value routing call. Flux 2 costs roughly nine times less and wins on breadth; GPT Image 1.5 leads the suite on quality (4.51) and is the only model that emits real transparency and reliably legible charts. Many pipelines need both, routed per job.
GPT Image 1.5 vs Nano Banana 2
OpenAI vs Google at the top of the quality table. GPT Image 1.5 owns structured data, dense text and true cut-outs; Nano Banana 2 owns brand-color fidelity at a lower unit price. Which flagship earns the expensive lane in your router.
Nano Banana 2 vs Nano Banana Pro
Is Pro worth three times the price in production? Nano Banana Pro costs $0.36 per image, runs slower, and does not score higher on our suite (4.19 vs 4.25). The unit-economics check before you default an entire feature to the Pro tier.
FLUX.2 vs Ideogram 3.0
The text-rendering reputation, tested. Ideogram built its name on typography, yet Flux 2 outscored it on every typography prompt in the suite. Verify reputations against current measurements before you route text-critical work.
Ideogram 3.0 vs Recraft V3
The two models pitched at design teams, head to head. Both land below their reputation on this suite (3.59 and 3.16), and Recraft is weakest exactly where design work needs consistency. What the data says before you standardize on either.
Nano Banana 2 vs Stable Diffusion 1.5
A current flagship against the legacy baseline: 4.25 vs 1.83. The concrete measure of what migrating an SD-era integration to a modern hosted model buys you in output quality.
FLUX.2 [dev] Turbo vs Seedream 4.5
The high-throughput tier, head to head. Both land near $0.015 and score around 4.0; Flux 2 Turbo adds speed and unusually good obedience to structured, machine-generated prompts. For in-app generation and automation pipelines where unit cost rules.
New model? New benchmarks.
We re-run the identical 37-prompt suite on every model release. Get the scores, and what changed in the rankings, in your inbox.
You're on the list.
We'll email you when new benchmark results ship.

