Image generation models usually get compared on how good the pictures look. If you’re putting one inside a creative product, that’s only part of the question, because what you need back is a layer someone can then edit. Whether a model can give you one comes down to four things: does it give you a real alpha channel, does it render text you can ship, does it hold a brand color, and what does it cost at volume.

We built an image model benchmark to answer those questions with numbers instead of impressions, running twenty-three models against identical prompts and scoring each one per criterion. This guide is what those results say about choosing a model for a creative application. It’s about still images only. Text-to-video is a separate purchase with different constraints, and we don’t benchmark it, so it isn’t ranked here.

Updated September 2026.

What you are actually choosing between

The models split into three groups, and the group matters more than the individual model, because it decides what you can promise your users.

Hosted APIs. You send a prompt, you get an image, you pay per image. That’s most of the field: GPT Image 1.5, the Nano Banana family, Seedream, Ideogram, Recraft, Luma Photon. They’re fast to integrate and there’s nothing to operate, but you inherit the provider’s terms, rate limits and content policy.

Open weights. Stable Diffusion 1.5, Qwen-Image and FLUX.2 [dev] can be self-hosted, with the weights distributed through Hugging Face, though FLUX.2 [dev]‘s terms separate non-commercial use from hosted commercial use. You take on GPUs and operations, and you gain control over data residency, fine-tuning and per-image cost at scale. Qwen-Image is the one worth looking at. Stable Diffusion 1.5 is from 2022 and scores like it.

Gateway models. Four of the twenty-three are available through the IMG.LY AI gateway: FLUX.2, Seedream 4.5, Nano Banana Pro and Nano Banana 2. That matters for billing and for swapping models without touching your integration code. It doesn’t change output quality.

A self-hosted model and a metered hosted API are very different commitments, so settle the group first.

Three findings that shape the integration

Running the same prompts across twenty-three models turned up three results that decide how much work your integration becomes.

Text rendering is close to solved

In 2025 this was the thing that separated models. It isn’t any more: twenty-two of twenty-three now score 4.0 or better on text accuracy. Nine of them reach a perfect 5.0, FLUX.2 [pro], GPT Image 1.5 and Nano Banana Pro among them, and seven more sit at 4.92 or above. The one exception is Stable Diffusion 1.5 at 1.17, which came out in 2022.

If you ruled out image generation because the text came back garbled, that is no longer a reason to. See the text reliability finding for the per-model numbers.

Transparency narrows the field to three at default settings

Twenty of the twenty-three models score zero on alpha. At default settings, only GPT Image 1.5 (4.22), Recraft V4.1 (2.0) and Recraft V3 (2.0) return any alpha channel at all, and at 2.0 the two Recraft channels are only half usable. The benchmark scores this as whether the returned file carries a genuine alpha channel with clean edges, so a model that paints a white or checkerboard background and calls it transparent scores zero.

If a model can’t produce transparency, it can’t produce a sticker, a logo, a product cutout, or anything else meant to sit over something. The same gap shows up in editing: a model that only hands back a finished rectangle can’t inpaint a region or swap one object and leave the rest alone, which is what most in-product editing asks for. You can bolt on background removal afterwards, and plenty of products do, but that adds another model and another point of failure in the pipeline. If your app composites, this one column narrows the field to three at default settings before you look at anything else. The transparency finding has the detail.

Brand color is unsolved across the board

Color accuracy is the weakest criterion in the entire benchmark, with the best score going to GPT Image 2 at 4.11 and most of the field sitting between 1.67 and 3.22, so nothing here is reliably good at it. The measurement is CIEDE2000 distance between the hex the prompt asks for and the color actually rendered, so the models are being asked for a specific value and missing it.

For a creative app this is a design constraint rather than a model choice. If a brand hex has to be exact, don’t ask the model to hit it in the prompt. Generate the imagery, then apply the brand color as an editable layer afterwards. The brand color finding covers what does and does not survive a prompt.

How we evaluated these

Every model ran the same prompts through the same test rig and was scored on eight criteria: prompt adherence, composition, text accuracy, color accuracy, resolution, alpha, latency and cost. The methodology page documents the prompt set, the scoring scale and the run counts.

The benchmark scores models available through the IMG.LY gateway alongside models that are not. Prices are per image at the time of testing and move often, so treat the ratios as durable and the absolute figures as a snapshot.

We haven’t tested every model on the market. Midjourney is excluded because its own Community Guidelines state that, with a few rare exceptions that are explicitly granted, it does not provide an API and prohibits automating interactions with the service. Adobe Firefly is absent from the benchmark, so this guide does not rank it.

Five more names that come up alongside these models are not in the run either, so here is where each one sits.

NameWhat it isWhy it is not in the run
DALL·E 3OpenAI’s earlier text-to-image modelRemoved from the OpenAI API on May 12, 2026. The GPT Image line that replaced it is in the table
SDXLStability AI’s 2023 open-weights model; Stable Diffusion 3.5 is the current lineNot tested. Stable Diffusion 1.5 is the Stability model that was scored
Hugging FaceA hub for open models with hosted inference endpointsNot a model. It is where the open-weights models above are distributed
Runway Gen-2Text-to-video and animated image generationOut of scope. This run covers text-to-image only
Leonardo.AiStyle-consistent 2D asset generationNot tested

The benchmark also has a boundary that is wider than those exclusions. It covers text-to-image generation and nothing else. It doesn’t measure image-to-image editing, inpainting, upscaling, standalone background removal or video. It also scores a single generation rather than a back-and-forth refinement loop, which is how people actually work in a design tool. A model that scores poorly here on a one-shot prompt can still be the right choice if your product iterates. Treat these numbers as the floor a model starts from rather than a verdict on everything it can do.

The field at a glance

ModelProviderRelative costMedian timeAlphaTextColorNotable
GPT Image 2OpenAIHighest110.1 s05.004.11Best color accuracy, and the costliest
Nano Banana ProGoogleHighest23.3 s05.003.67Strong color, near the top of the range
GPT Image 1.5OpenAIHighest34.0 s4.225.003.44Only model with a strong alpha channel
Seedream 5.0 ProByteDanceHighest65.3 s04.923.78Top resolution, high cost
Nano Banana 2GoogleHigh13.2 s04.963.22Strong prompt adherence and text
Ideogram 3.0IdeogramHigh17.9 s04.712.78
Grok Imagine Image 2.0xAIHigh65.9 s05.002.89
GPT Image 2.5 FlareOpenAILow21.3 s05.003.78Current OpenAI generation, faster of the two
Nano Banana 2 LiteGoogleLow4.0 s04.922.89Strong composition in the low-cost band
Qwen Image 3.0AlibabaLow141.1 s05.002.56Slowest in the set
Seedream 4.5ByteDanceLow12.4 s04.882.89Top resolution
Recraft V3RecraftLow7.4 s24.001.89One of three models with any alpha
Gemini 2.5 Flash ImageGoogleLow7.4 s04.961.67Fast, weakest color, scheduled to shut down October 2, 2026
GPT Image 2.5 SunburstOpenAILow30.4 s04.964.00Highest prompt adherence
Seedream 5.0 LiteByteDanceLow30.7 s04.882.56Top resolution, low cost
Recraft V4.1RecraftLow10.7 s25.002.89Alpha channel, at a fraction of the cost
FLUX.2 [pro]Black Forest LabsLow11.5 s05.002.33Perfect text score
Luma PhotonLuma AILowest16.1 s04.132.11High resolution score
Qwen-ImageAlibabaLowest6.9 s04.832.00Open weights, self-hostable
FLUX.2Black Forest LabsLowest2.0 s04.922.78Low cost and fast
Muse ImageMetaLowest17.8 s05.003.44Low cost, perfect text score
FLUX.2 [dev] TurboBlack Forest LabsLowest2.0 s04.922.67Open weights; near the cheapest, and the fastest
Stable Diffusion 1.5Stability AILowest2.1 s01.172.222022 model, open weights

Scores are out of 5. Full per-criterion results are on the models page.

Which model for which job

We score models against seven jobs rather than one leaderboard, because the right answer changes with what you’re generating.

Stickers, merch and anything that composites. Start with GPT Image 1.5. It’s the only model with a strong alpha channel, and for cutouts that outweighs both its cost and its median 34 seconds per image, seventeen times FLUX.2. That wait rules it out of anything interactive, so generate cutouts in the background rather than in front of a waiting user. Recraft V4.1 and V3 are the fallbacks, both scoring 2.0 on alpha at roughly a quarter to under a third of the cost. One caveat on timing: OpenAI has scheduled GPT Image 1.5 to shut down on December 1, 2026. Transparency carries over to the current GPT Image 2.5 models, which take a background: transparent setting with PNG or WebP output. Both are in the table above, and both score zero on alpha in this run. See apparel, merch and POD.

Text-heavy layouts. Almost anything modern works. Pick on cost and speed instead: FLUX.2 sits in the lowest cost band, scores 4.92 on text and returns in a median 2.0 seconds, among the fastest in the set. See text-heavy layouts.

Print files. Resolution is the gate. Seedream 4.5, Seedream 5.0 Lite and Seedream 5.0 Pro all score 5.0 there, against 2.0 for the FLUX.2 variants. Color management stays your job in the editor. See print products.

High-volume personalization. Cost per image compounds, and the spread across the set is wide enough that the same million-image job differs by more than an order of magnitude depending on which model runs it. Check the current figures on the models page against your own volume before committing. See personalization at scale.

User-generated design inside your app. Prompt adherence and latency matter more than peak quality, because a user waiting on a slow model assumes your app is broken. GPT Image 2.5 Sunburst leads prompt adherence at 4.90 but takes a median 30.4 seconds; FLUX.2 returns in 2.0. Below roughly five seconds a generation feels like part of the app, and above about fifteen it needs its own progress state. See in-app UGC.

What you actually pay

Cost per image spans close to two orders of magnitude across the benchmark, and more than an order of magnitude between the cheapest modern model and the most expensive. At the top of that range the cost starts to shape how you build. Per-image figures live on the models page, which is generated from the benchmark data rather than written by hand, so it stays current as providers reprice.

Below roughly ten thousand images a month the difference is small enough that you should pick on capability. Above a million, it dominates every other consideration, and a model that looks expensive per image has to earn it on something specific. GPT Image 1.5 earns its place in the high band on alpha. GPT Image 2, the costliest in the set, has to earn that on color accuracy alone, which is a narrow case.

Self-hosting flips the shape of the bill from per-image to per-GPU-hour. That’s attractive at sustained high volume and unattractive the moment your traffic is spiky, because idle GPUs bill the same as busy ones.

Where CE.SDK fits

Every model in this benchmark returns a flat image. A creative app needs a layer your users can move, restyle, mask or replace, and nothing above does that on its own.

CE.SDK provides the editing surface that generated images land on, so output arrives as an editable element rather than a finished picture. Its AI plugins connect to the model of your choice, and the AI gateway covers billing and model switching for FLUX.2, Seedream 4.5, Nano Banana Pro and Nano Banana 2. Models outside the gateway go through the same plugin interface with your own API key.

If you only need images generated and delivered, with nobody editing them afterwards, you don’t need an editor at all. Call a model API directly.

FAQ

Which image generation API is best for creative apps in 2026? There is no single answer, because the criteria conflict: GPT Image 1.5 is the only model with a strong alpha channel, FLUX.2 and its Turbo variant are the fastest and sit in the lowest cost band, and Seedream leads on resolution. GPT Image 2 leads on color accuracy and costs 13 times more than FLUX.2. Decide which criterion your product cannot compromise on, then pick from that column.

Can I still use GPT-4o for image generation? GPT-4o image generation shipped as the gpt-image-1 API in April 2025 and was superseded by GPT Image 1.5. OpenAI has scheduled gpt-image-1 to shut down on October 23, 2026 and GPT Image 1.5 on December 1, 2026, naming gpt-image-2 as the replacement for both. The benchmark tests four models on that line.

Do any of these models produce transparent backgrounds? Three of the twenty-three tested, at the default settings the benchmark uses. GPT Image 1.5 scores 4.22 on alpha, and Recraft V4.1 and Recraft V3 score 2.0. The other twenty score zero at those settings, so they need a separate background removal step before the output can be composited. Note that GPT Image 1.5 is scheduled to shut down on December 1, 2026. Transparency is available on the current GPT Image 2.5 models through the background: transparent setting, a setting we have not benchmarked.

Is text rendering still a problem? Not for current models. Twenty-two of twenty-three score 4.0 or better out of 5 on text accuracy, and nine of them reach 5.0. Stable Diffusion 1.5, released in 2022, is the outlier at 1.17.

Which models can I self-host? Qwen-Image is published under Apache-2.0, and Stable Diffusion 1.5 under CreativeML OpenRAIL-M. FLUX.2 [dev] has open weights under terms that separate non-commercial use from hosted commercial use. Everything else in the benchmark is API-only.

Why is Midjourney not in the comparison? Its Community Guidelines state that Midjourney does not provide an API except by explicit grant, and that automating interactions with the service is prohibited. That puts it out of scope for a guide about embedding generation in a product.

How much does generation cost at scale? Per-image cost spans more than an order of magnitude depending on the model, so at high volume the choice of model matters more than almost anything else in the bill. The models page carries the current per-image figures.

How current are these numbers? The benchmark covers model versions released between October 2022 and September 2026. Prices and model versions change frequently; the models page carries the current figures.

Sources

Model scores, costs and version dates come from the IMG.LY image model benchmark, with the scoring approach documented on the methodology page. Licensing summaries come from each provider’s own model listing, linked from the individual model pages. The Midjourney exclusion is sourced from its Community Guidelines in the Midjourney help center.


3,000+ creative professionals gain early access to new features and updates. Don’t miss out, and subscribe to our newsletter.