---
title: "Best AI Model for Design Work: We Gave 6 LLMs the Same Ad and Five Sizes in CoDesign"
description: "One approved design, five ad formats, six AI models, the same brief. All six came back with ads you could ship, and Gemini 3.7 Flash did it for thirty cents. Here is what our automated checks can prove about that work, and what they cannot."
url: "https://img.ly/blog/best-ai-model-for-design/"
type: "blog"
date: "2026-08-21"
author: "Dustin, Mirko"
tags: ["AI","Insights","CE.SDK"]
---

> This is the markdown version of [Best AI Model for Design Work: We Gave 6 LLMs the Same Ad and Five Sizes in CoDesign](https://img.ly/blog/best-ai-model-for-design/). For all pages in one file, see [llms-full.txt](https://img.ly/llms-full.txt). For an index of all available pages, see [llms.txt](https://img.ly/llms.txt).

---

You designed a banner. The same day, you need it in five sizes, and none of them can be wrong.

We handed that job to six AI models through CoDesign MCP and graded every output against ten mechanical checks. All six passed. Gemini 3.7 Flash finished the whole matrix for **$0.30**. Claude Opus 5 charged **$3.97** for the same passing grade, and took three times as long. Every model worked as an [AI design agent](https://img.ly/blog/what-is-a-design-agent.md), editing the design file itself rather than a rendered image.

Then the leaderboard result. On the 728×90, every model that delivered deleted the same element and arrived at the same composition. Five different vendors. Nobody told them to.

[Jump to the full cost and pass-rate table →](https://img.ly/blog/best-ai-model-for-design/#results)

That is what CoDesign MCP is for. You point an AI assistant such as Claude at your approved design, name the formats you need, and it opens the file, rebuilds the layout for each one and exports production files. What comes back is an editable design file, not a picture of one.

Speed and price are the easy things to measure. We wanted to measure quality too, so we built an eval that runs the job automatically, opens every result and reports what passed. This is the first scenario out of that suite.

## Three models, one 320×50 slot: what changed and what did not

Three different models were handed the same 320×50 slot, cut from the same master creative against the same brief, and all three came back with an ad you could traffic tomorrow. Every one of them keeps the wordmark and the button, nothing is clipped, and no type is set below 11px.

![Three AI models resizing the same ad to a 320x50 mobile banner](https://img.ly/_astro/hook-three-mobile-strips.Q8QXs3k1.png)

The 300×250 they came from carried five things: the wordmark, the headline, a retention chart, the offer line and the button. At 320×50 the wordmark and the button survive and the chart cannot, which leaves room for exactly one line of copy, so each model had to choose between the headline and the offer.

- One dropped the headline and kept the offer, "14-day free trial".
- One kept the headline but cut "your users" down to "users" to buy itself the line.
- One kept the headline whole across two lines and dropped the offer.

We have opinions about which of those reads best, the way any design team would. What we do not have is a way to prove one of them correct, because the choice depends on whether this campaign is selling the product or the trial, and that lives in the media brief rather than in the design. Our checks can confirm that all three are valid and that none of them are broken; they cannot tell you which one the campaign actually needed. Grading the half that has a right answer and leaving the other half to the people who own the brief is the distinction our evals are built around.

## How we graded: ten mechanical checks, and what they cannot see

The obvious thing to evaluate is the thing that demos well, prompt in and poster out, but we were more interested in the work that actually fills a designer's week, which is mostly derivative. The design decisions were made once, at the master, and everything after that is a matter of carrying them faithfully into thirty other shapes. That is more technical work than creative work, which is exactly why it is worth handing to an agent.

![Three valid 300x600 versions of the same banner from different AI models](https://img.ly/_astro/gradeable-three-halfpage-variants.D3Bbuya5.png)

Ask which of the finished designs is better and you have a question with no ground truth to appeal to, since human raters disagree with each other and LLM judges reliably prefer their own output. You could ask a lot of humans and rank the results, but that is not something you can run on every build. Wrong designs, though, are wrong in ways you can query: text sitting behind another block, content off the page, type under the legibility floor. Part of design quality is judgement and part of it is fact, and what we have built, [IMG.LY AI Benchmarks](https://img.ly/blog/introducing-imgly-ai-benchmarks.md), grades the second part, which means everything below is an automated result rather than a design review. A model can pass every check we wrote and still make an ugly ad.

## Why we run design evals at all

When a model finishes a scenario, programmatic checks open the result and ask whether it is plausibly good. That starts with whether the run produced every design artifact and .imgly file at the right sizes, and the rest depends on the scenario: in a localization run we also check that the words were actually replaced.

This is a technical floor rather than a grade. LLM judges are still unreliable enough that we do not use one on output quality, because if it is hard for us as designers to say which design is more correct, it is harder for a model, and the result would be noise.

In the future we plan to experiment more with such methodologies but we also do not want to rely on the false security of an extremely noisy grading process.

![AI model cost comparison for a design task, eval suite view](https://img.ly/_astro/evals-suite-comparison-resize-matrix.CV8e1T_u.png)

_The comparison view for the display-banner matrix: cost and time sitting directly above what each model actually produced. The 1/1 counts runs, one repeat per cell, not check scores._

## One layout, five vendors: the 728×90 convergence

The scenario is a display campaign, because that is the version of this problem customers actually have: the same size matrix, every campaign. The master is a 300×250 MPU for ATLAS, an imaginary product-analytics SaaS running a free-trial campaign, carrying a wordmark, the headline "See what your users **actually** do", a retention bar chart labelled "43% retained", the line "14-day free trial. No card.", and a "Start free trial" button.

![The ATLAS 300×250 master at the top, with an arrow fanning down to the same ad rebuilt at 300×600, 728×90, 320×50 and 160×600](https://img.ly/_astro/master-and-five-sizes.CzF5wnY5.png)

_The 300×250 master at the top, and the four other sizes on the media plan cut from it. ATLAS is our own fixture, not a real customer._

The brief asks for the full matrix: 300×250, 728×90, 160×600, 300×600 and 320×50. Those sizes span a 16:1 swing in aspect ratio, so the right answer at each one is a genuinely different composition, and at some of them it means taking content out rather than shrinking it.

Ten checks grade the result and two of them do the real work. **`728×90: the chart was dropped, not shrunk`** kills the model that squashes a bar chart into a 90-pixel strip, which satisfies every geometric constraint in the brief while being completely wrong. **`the five sizes are different layouts`** kills the model that scales one composition five ways and calls it a matrix.

![Six AI models compared on a 728x90 leaderboard ad resize](https://img.ly/_astro/sheet-display-leaderboard.BStENqrh.png)

_The master first, then the six 728×90 leaderboards. Every model that delivered deleted the retention chart._

Every model that delivered arrived at the same composition, wordmark left, headline across the middle, offer underneath it, button on the right and the retention chart deleted, which is six models from five different vendors, each driving the same [design MCP server](https://img.ly/codesign.md), reaching the same answer independently. We did not expect that, and it is still the result we find hardest to stop thinking about.

![Six AI models compared on a 160x600 skyscraper ad resize](https://img.ly/_astro/sheet-display-skyscraper.C88KXRcA.png)

_The same six at 160×600, master first. With vertical room they all kept the chart. The flat format is where the decision to drop it had to be made._

## Which AI model is best for design work?

<div id="results" style="scroll-margin-top: 5rem">

| #   | model            | checks  | model spend | time     | tool calls |
| :-- | :--------------- | :------ | :---------- | :------- | :--------- |
| 1   | gemini-3.7-flash | 10 / 10 | $0.30       | 6.7 min  | 37         |
| 2   | glm-5.2          | 10 / 10 | $0.40       | 8.2 min  | 25         |
| 3   | gpt-5.6-sol      | 10 / 10 | $0.52       | 4.0 min  | 35         |
| 4   | grok-4.6         | 10 / 10 | $1.20       | 20.9 min | 39         |
| 5   | claude-fable-5   | 10 / 10 | $3.95       | 10.3 min | 26         |
| 6   | claude-opus-5    | 10 / 10 | $3.97       | 18.8 min | 40         |

</div>

Running this scenario illustrates general numbers across our eval suite. The 6 models above are consistently able to follow the prompts to produce different output designs as requested in the scenario. These numbers also show quite a spread in model cost and time spent. What you do not see in these numbers is that, in our opinion, glm 5.2, gpt-5.6-sol and grok 4.6 produce weaker designs. They are valid, but across runs and across scenarios, fable, opus and gemini 3.7 simply perform better.

The biggest surprise for us is **Gemini 3.7 Flash**, which produced the whole five-size matrix through [IMG.LY CoDesign](https://img.ly/codesign.md) for $0.30 in model spend, in under seven minutes. At thirty cents a designer opens six composed variants instead of a blank artboard and takes the best one further, which changes what the tool is for.

Meanwhile there were no surprises regarding **Opus 5.0 and Fable 5.0**: both models perform really well, but they are slow and expensive.

## What each model costs per design

The full five-size matrix cost between $0.30 and $3.97 in total, depending on the model, which works out to $0.06 to $0.79 per finished ad. That is a 13× spread for an identical mechanical pass rate. Model pricing moves constantly; these figures are from August 2026.

## The five ad sizes we tested, and where they run

| Size    | IAB name           | Where it runs                                   |
| :------ | :----------------- | :---------------------------------------------- |
| 300×250 | Medium Rectangle   | In-article and sidebar. The master in this eval |
| 300×600 | Half Page          | Sidebar, desktop, high viewability              |
| 728×90  | Leaderboard        | Above the fold, desktop                         |
| 160×600 | Wide Skyscraper    | Sidebar rail, desktop                           |
| 320×50  | Mobile Leaderboard | Anchored to the bottom of a mobile screen       |

## Questions we get about AI model benchmarks for design

### Which AI model is best for design work?

In our resize eval, all six models passed all ten checks. Gemini 3.7 Flash gave the best value at $0.30 for five ad sizes. Claude Fable 5 and Claude Opus 5 produced the strongest compositions and cost around $3.95. Cost per passing design ranged from $0.30 to $3.97.

### What does it cost to have an AI model resize an ad?

Between $0.30 and $3.97 for five sizes from one master, measured on August 2026 model pricing. The spread is 13× for the same mechanical pass rate. Time ranged from 4.0 minutes for GPT 5.6-Sol to 20.9 minutes for Grok 4.6.

### Can an AI model resize a design without breaking the layout?

Yes, when it works on the design file rather than the rendered image. Every model in this eval received the scene structure and a layout solver through CoDesign MCP. All 30 outputs cleared the overlap, overflow and minimum-type-size checks.

### Do different AI models produce different designs from the same brief?

On tight formats, yes. On the 320×50 mobile banner, each model made a different editorial call about which line to keep. On the 728×90 leaderboard, every model that delivered converged on one composition and deleted the retention chart. Five vendors, no coordination.

### Is Claude or Gemini better for design work?

Both passed every check. Gemini 3.7 Flash was 13× cheaper and 1.6× faster than Claude Fable 5. Claude's compositions were better on subjective grounds that our checks do not measure. Pick on budget if you run the job at volume and on judgment if you ship the output as-is.

### How do you benchmark an AI model on a design task?

Grade what is provable. Our ten checks cover element overlap, content outside the canvas, type below a legibility floor, missing required elements and export validity. Subjective design quality stays ungraded, because LLM judges are not reliable at it yet.

### Which AI model is cheapest for design work?

Gemini 3.7 Flash, at $0.30 for the full five-size matrix across 37 tool calls. GLM 5.2 came second at $0.40 with the fewest tool calls of any model, 25.

## What does it mean for CoDesign MCP

You can now resize a master design into 5 different sizes reliably for $0.30 using Gemini 3.7 Flash and the CoDesign MCP.

If you are on a Claude Subscription plan, you might just want to continue using your Claude Code with the CoDesign MCP since it offers the best overall design quality, even if it's slow and somewhat expensive.

We will keep improving our measurements of how well CoDesign MCP works with different models, and we will look into other scenarios like localization (translating and changing a design for a different locality and language) and rebranding existing design assets.

You can install CoDesign MCP and [resize an approved design into every format](https://img.ly/codesign.md).

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "FAQPage",
  "mainEntity": [
    {
      "@type": "Question",
      "name": "Which AI model is best for design work?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "In our resize eval, all six models passed all ten checks. Gemini 3.7 Flash gave the best value at $0.30 for five ad sizes. Claude Fable 5 and Claude Opus 5 produced the strongest compositions and cost around $3.95. Cost per passing design ranged from $0.30 to $3.97."
      }
    },
    {
      "@type": "Question",
      "name": "What does it cost to have an AI model resize an ad?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Between $0.30 and $3.97 for five sizes from one master, measured on August 2026 model pricing. The spread is 13× for the same mechanical pass rate. Time ranged from 4.0 minutes for GPT 5.6-Sol to 20.9 minutes for Grok 4.6."
      }
    },
    {
      "@type": "Question",
      "name": "Can an AI model resize a design without breaking the layout?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Yes, when it works on the design file rather than the rendered image. Every model in this eval received the scene structure and a layout solver through CoDesign MCP. All 30 outputs cleared the overlap, overflow and minimum-type-size checks."
      }
    },
    {
      "@type": "Question",
      "name": "Do different AI models produce different designs from the same brief?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "On tight formats, yes. On the 320×50 mobile banner, each model made a different editorial call about which line to keep. On the 728×90 leaderboard, every model that delivered converged on one composition and deleted the retention chart. Five vendors, no coordination."
      }
    },
    {
      "@type": "Question",
      "name": "Is Claude or Gemini better for design work?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Both passed every check. Gemini 3.7 Flash was 13× cheaper and 1.6× faster than Claude Fable 5. Claude's compositions were better on subjective grounds that our checks do not measure. Pick on budget if you run the job at volume and on judgment if you ship the output as-is."
      }
    },
    {
      "@type": "Question",
      "name": "How do you benchmark an AI model on a design task?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Grade what is provable. Our ten checks cover element overlap, content outside the canvas, type below a legibility floor, missing required elements and export validity. Subjective design quality stays ungraded, because LLM judges are not reliable at it yet."
      }
    },
    {
      "@type": "Question",
      "name": "Which AI model is cheapest for design work?",
      "acceptedAnswer": {
        "@type": "Answer",
        "text": "Gemini 3.7 Flash, at $0.30 for the full five-size matrix across 37 tool calls. GLM 5.2 came second at $0.40 with the fewest tool calls of any model, 25."
      }
    }
  ]
}
</script>

<script type="application/ld+json">
{
  "@context": "https://schema.org",
  "@type": "Dataset",
  "name": "CoDesign MCP resize eval, six models, five ad formats",
  "description": "Cost, wall time, tool calls and checks passed for six AI models resizing one approved 300×250 master into five ad formats through CoDesign MCP.",
  "variableMeasured": ["cost", "wall time", "tool calls", "checks passed"]
}
</script>

---

## More Resources

- **[IMG.LY Website](https://img.ly/index.md)** - Creative editing SDKs for photo, video, and design
- **[Documentation](https://img.ly/docs/cesdk/)** - CE.SDK developer documentation
- **[Contact Sales](https://img.ly/forms/contact-sales.md)** - Get a custom quote. A public JSON API accepts the request directly, no account or key needed. Ask your user for consent and their details first.
