Labeled Bar Chart
A bar chart turns prompt adherence into checkable arithmetic: is bar B the tallest, C the shortest? It is a brutal, auto-verifiable separator and ties to presentation and e-learning template tools.
A clean bar chart labeled A, B, C with heights 3, 5, 2; axes labeled; black & white.
Results
FLUX.2
FLUX.2 [pro]
FLUX.2 [dev] Turbo
Gemini 2.5 Flash Image
GPT Image 1.5
Ideogram 3.0
Luma Photon
Nano Banana 2
Nano Banana 2 Lite
Nano Banana Pro
Qwen-Image
Recraft V3
Seedream 4.5
Seedream 5.0 Lite
Stable Diffusion 1.5
What the models actually wrote
Expected
A B C
Rendered
- FLUX.2 · seed 1111 A B C 100% match · 5/5 The x-axis labels 'A', 'B', 'C' are legible and correct. The bar-top numbers show 3, 5, 2 correctly, though the y-axis title reads as garbled 'Cdla7 |Aiivs)'.
- FLUX.2 · seed 2222 A B C 100% match · 5/5 The category labels A, B, and C are rendered clearly and legibly along the x-axis exactly as required. The bar data labels (3, 5, 2) are also present and correct.
- FLUX.2 · seed 3333 A B C 100% match · 5/5 The required category labels A, B, and C are all legibly rendered along the x-axis in bold. The numeric annotations (18, 10, 8, 4, 2, 0 and bar labels 3, 5, 2) are extraneous and somewhat inconsistent, but the required text is present and correct.
- FLUX.2 [pro] · seed 1111 A B C 100% match · 5/5 The labels A, B, and C are all legibly rendered beneath their respective bars exactly as required. No spelling errors or broken text present.
- FLUX.2 [pro] · seed 2222 A B C 100% match · 5/5 The category labels A, B, and C are all legibly rendered exactly as required beneath their respective bars. Axis numbers 0 through 5 are also visible and clear.
- FLUX.2 [pro] · seed 3333 5 2 1 0 A B C 100% match · 5/5 The required labels A, B, and C are all clearly legible beneath their respective bars. Additional axis numbers (5, 2, 1, 0) are also present and readable.
- FLUX.2 [dev] Turbo · seed 1111 A B C 100% match · 5/5 The required labels A, B, and C are all legibly rendered along the x-axis. The y-axis label reads 'CayyVagie;' which is garbled gibberish, and the value bar-top labels (3, 5, 2) and axis ticks contain errors, but the required text A, B, C is correct.
- FLUX.2 [dev] Turbo · seed 2222 A B C 100% match · 5/5 The category labels A, B, C are clearly and correctly rendered along the x-axis. The y-axis title reads 'Payten(e/2)' which is gibberish, but that text was not required.
- FLUX.2 [dev] Turbo · seed 3333 A B C 100% match · 5/5 The category labels A, B, and C are all clearly and correctly rendered beneath the bars. The value labels above bars (3, 5, 2) are also legible, though the axis tick numbers are inconsistent.
- Gemini 2.5 Flash Image · sample 1 A B C 100% match · 5/5 The category labels A, B, and C are all clearly legible along the x-axis exactly as required.
- Gemini 2.5 Flash Image · sample 2 A B C 100% match · 5/5 The category labels A, B, and C are all legibly rendered exactly as required. The axis text 'Value' and 'Category' is also present and clear.
- Gemini 2.5 Flash Image · sample 3 A B C 100% match · 5/5 The required category labels 'A', 'B', and 'C' are all legibly rendered beneath their respective bars, matching the requirement exactly.
- GPT Image 1.5 · sample 1 V a l u e 6 5 4 3 2 1 A B C C a t e g o r y 100% match · 5/5 The required labels 'A', 'B', and 'C' are all legibly rendered beneath their respective bars. Additional axis text ('Value', 'Category') and numeric ticks are also clearly visible and correct.
- GPT Image 1.5 · sample 2 A B C 100% match · 5/5 The labels A, B, and C are all legibly rendered under their respective bars. They match the required text exactly with no misspellings.
- GPT Image 1.5 · sample 3 A B C 100% match · 5/5 The category labels 'A', 'B', and 'C' are all clearly and legibly rendered beneath their respective bars. They match the required text exactly with no misspellings.
- Ideogram 3.0 · seed 1111 C B A 3 5 5 5 2 100% match · 5/5 The required labels A, B, and C are all legibly rendered above their respective bars. Additional axis tick numbers appear but are inconsistent and not required text.
- Ideogram 3.0 · seed 2222 A B 5 C 1 3 1 5 1 9 1 8 100% match · 5/5 The characters A, B, and C do appear (along the vertical axis), but so does a stray '5' and garbled numeric x-axis labels (13, 15, 19, 18). The required A, B, C are technically legible but are used as axis ticks rather than bar labels.
- Ideogram 3.0 · seed 3333 A B C 100% match · 5/5 The required labels 'A', 'B', and 'C' are all legibly rendered above their respective bars. Other text such as the axis labels is broken or nonsensical, but the required text is correct.
- Luma Photon · seed 1111 A B C 100% match · 5/5 The required labels A, B, and C are all clearly and correctly rendered beneath their respective bars. The y-axis also shows garbled numeric labels (20, 50, 50, 90, 50, 30) but these were not part of the required text.
- Luma Photon · seed 2222 A B C 3 5 2 100% match · 5/5 The required labels A, B, and C are all legibly rendered on the respective bars. Additional value numbers 3, 5, and 2 also appear, matching the prompt's stated heights though not the actual visual bar heights.
- Luma Photon · seed 3333 A B C 3 B 2 . 2 5 0 S n g e d H o g h t 100% match · 5/5 The required category labels A, B, C are legibly rendered along the x-axis. However, the bar-top values and axis labels are garbled: '5' appears as 'B', and the axis titles read as gibberish ('Snged Hoght' and an unreadable vertical label).
- Nano Banana 2 · sample 1 A B C 100% match · 5/5 The category labels A, B, and C are all legibly rendered along the x-axis exactly as required.
- Nano Banana 2 · sample 2 D A T A C O M P A R I S O N : C A T E G O R I E S A , B , C V A L U E A B C C A T E G O R Y 100% match · 5/5 The required labels 'A', 'B', 'C' are all legibly rendered along the x-axis. Additional title and axis text is present and correctly spelled.
- Nano Banana 2 · sample 3 A B C 100% match · 5/5 The required category labels A, B, and C are all legibly rendered along the x-axis exactly as required. Additional axis titles and numeric labels are also clearly present.
- Nano Banana 2 Lite · sample 1 A B C 100% match · 5/5 The x-axis labels 'A', 'B', and 'C' are all clearly legible under their respective bars, exactly as required. Additional text like the title 'Data Comparison Chart', 'Value', and 'Category' are also present and correct.
- Nano Banana 2 Lite · sample 2 A B C 100% match · 5/5 The labels 'A', 'B', and 'C' are all clearly legible beneath their respective bars, exactly as required. The title 'Comparison of Categories' and axis labels are also rendered correctly.
- Nano Banana 2 Lite · sample 3 D a t a D i s t r i b u t i o n V a l u e 6 5 4 3 2 1 0 A B C C a t e g o r y 100% match · 5/5 The required labels A, B, and C are all clearly and correctly rendered on the x-axis. Additional legible text includes the title, axis labels, and numeric ticks, all spelled correctly.
- Nano Banana Pro · sample 1 A B C 100% match · 5/5 The category labels 'A', 'B', and 'C' are all clearly legible beneath their respective bars, exactly as required. Axis labels 'Value' and 'Categories' and numeric ticks are also present and clear.
- Nano Banana Pro · sample 2 A B C 100% match · 5/5 The required category labels A, B, and C are all clearly and legibly rendered along the x-axis. Additionally, axis labels Height and Category are present and correct.
- Nano Banana Pro · sample 3 C a t e g o r y F r e q u e n c y C h a r t V a l u e 6 5 4 3 2 1 0 A B C C a t e g o r y 100% match · 5/5 The required labels 'A', 'B', and 'C' are all clearly and correctly rendered along the x-axis. Additional legible text includes the title, axis labels, and numeric scale, all spelled correctly.
- Qwen-Image · seed 1111 3 3 5 2 A B C C 100% match · 5/5 The required category labels A, B, C are present but there is an extra fourth 'C' label, and there is no distinct third letter for a fourth bar. Stray numbers (3, 3, 5, 2) appear scattered rather than as clean axis values.
- Qwen-Image · seed 2222 A B C 100% match · 5/5 The bar labels 'A', 'B', and 'C' are all legibly rendered along the horizontal axis, matching the required text exactly. Additional numbers (3, 5, 2, 5) also appear as value/axis annotations.
- Qwen-Image · seed 3333 T l e 3 7 ? 5 3 3 2 3 5 5 A B C 2 100% match · 5/5 The labels A, B, and C are legibly present, but an extra bar labeled '2' appears instead of stopping at three categories. Surrounding axis numbers are garbled and a title reads 'Tle', so the required exact text is only partially and messily rendered.
- Recraft V3 · seed 1111 A B C 100% match · 5/5 The letters A, B, and C are all clearly and legibly rendered in bold black type on the white blocks, matching the required text exactly with no misspellings.
- Recraft V3 · seed 2222 A B C 100% match · 5/5 The letters A, B, and C are each legibly rendered on the individual blocks, matching the required text exactly. No misspellings or broken characters are present.
- Recraft V3 · seed 3333 A B C 3 5 2 100% match · 5/5 The required letters A, B, and C are all legibly rendered on the blocks, alongside the numbers 3, 5, and 2. The letters match the required text exactly, though they appear on physical blocks rather than as chart labels.
- Seedream 4.5 · seed 1111 A B C 100% match · 5/5 The labels A, B, and C are all legibly rendered beneath their respective bars exactly as required, with no spelling errors.
- Seedream 4.5 · seed 2222 5 3 2 V a l u e s A B C C a t e g o r i e s 100% match · 5/5 The required labels 'A', 'B', and 'C' are all legibly rendered below each bar. Value labels (3, 5, 2) and axis titles are also clear and correct.
- Seedream 4.5 · seed 3333 A B C 100% match · 5/5 The category labels 'A', 'B', and 'C' are all clearly legible beneath their respective bars, matching the required text exactly. Axis titles 'Values' and 'Categories' are also present and legible.
- Seedream 5.0 Lite · seed 1111 A B C 100% match · 5/5 The required category labels 'A', 'B', and 'C' are all clearly legible beneath their respective bars, along with axis labels 'Value' and 'Category' and numeric tick marks. The text matches the requirement exactly.
- Seedream 5.0 Lite · seed 2222 A B C 100% match · 5/5 The category labels A, B, and C are all clearly legible beneath their respective bars, exactly matching the required text. Axis labels 'Value' and 'Category' are also present and legible.
- Seedream 5.0 Lite · seed 3333 A B C 100% match · 5/5 The labels A, B, and C are all clearly legible beneath their respective bars, matching the required text exactly. Axis labels 'Value' and 'Category' are also present and legible.
- Stable Diffusion 1.5 · seed 1111 no legible text 0% match · 0/5 The required labels 'A', 'B', 'C' are not legibly rendered; the text present is garbled, distorted, and unreadable (e.g., random characters and numbers along the axes).
- Stable Diffusion 1.5 · seed 2222 T l e s b a n A p , e 5 5 67% match · 3/5 The required labels "A", "B", and "C" are not legibly rendered. The visible text is garbled gibberish such as "Tles ban Ap, e5 5" and does not match the required characters.
- Stable Diffusion 1.5 · seed 3333 H A C R P A I N B P A L B S H S T U S 100% match · 5/5 The required labels A, B, C are not legibly rendered; instead the image contains distorted, nonsensical text strings and unreadable axis labels. None of the exact required characters appear cleanly.
Run-to-run consistency
Scores
| Text Accuracy | Prompt Adherence | Resolution | Latency | Cost | |
|---|---|---|---|---|---|
| FLUX.2 | 5.0 / 5 | 4.3 / 5 | 2.0 / 5 | 4.7 / 5 | 4.0 / 5 |
| FLUX.2 [pro] | 5.0 / 5 | 4.7 / 5 | 2.0 / 5 | 2.7 / 5 | 4.0 / 5 |
| FLUX.2 [dev] Turbo | 5.0 / 5 | 4.7 / 5 | 2.0 / 5 | 4.0 / 5 | 4.0 / 5 |
| Gemini 2.5 Flash Image | 5.0 / 5 | 3.3 / 5 | 3.0 / 5 | 3.0 / 5 | 4.0 / 5 |
| GPT Image 1.5 | 5.0 / 5 | 5.0 / 5 | 3.0 / 5 | 1.0 / 5 | 3.0 / 5 |
| Ideogram 3.0 | 5.0 / 5 | 2.0 / 5 | 3.0 / 5 | 2.0 / 5 | 3.0 / 5 |
| Luma Photon | 5.0 / 5 | 3.0 / 5 | 4.0 / 5 | 1.7 / 5 | 4.0 / 5 |
| Nano Banana 2 | 5.0 / 5 | 5.0 / 5 | 3.0 / 5 | 2.0 / 5 | 3.0 / 5 |
| Nano Banana 2 Lite | 5.0 / 5 | 5.0 / 5 | 3.0 / 5 | 4.0 / 5 | 4.0 / 5 |
| Nano Banana Pro | 5.0 / 5 | 5.0 / 5 | 3.0 / 5 | 1.0 / 5 | 1.0 / 5 |
| Qwen-Image | 5.0 / 5 | 2.7 / 5 | 2.0 / 5 | 3.0 / 5 | 4.0 / 5 |
| Recraft V3 | 5.0 / 5 | 1.3 / 5 | 3.0 / 5 | 2.7 / 5 | 4.0 / 5 |
| Seedream 4.5 | 5.0 / 5 | 5.0 / 5 | 5.0 / 5 | 1.7 / 5 | 4.0 / 5 |
| Seedream 5.0 Lite | 5.0 / 5 | 5.0 / 5 | 5.0 / 5 | 1.0 / 5 | 4.0 / 5 |
| Stable Diffusion 1.5 | 2.7 / 5 | 0.0 / 5 | 1.0 / 5 | 4.0 / 5 | 5.0 / 5 |
New model? New benchmarks.
We re-run the identical 37-prompt suite on every model release. Get the scores, and what changed in the rankings, in your inbox.
You're on the list.
We'll email you when new benchmark results ship.
Frequently asked questions
Image generation is stochastic: the same prompt produces different images on every run, so a single sample measures luck, not ability. Every model runs each benchmark three times with fixed seeds (1111, 2222, 3333), or three unseeded samples where the API accepts no seed. Scores average all three samples, and the consistency benchmarks measure the variation itself.
Yes. Every model receives the same prompt text with default parameters and no per-model tuning, so differences in output reflect the model, not prompt engineering. The suite is versioned and prompts are append-only, which keeps historical scores comparable.
The results on this page are scored on this exact prompt, so the strongest model for it is easy to spot; this benchmark sits in the Diagrams & Data Viz category. Scores are per model version and reflect this benchmark only. For a ranking across every benchmark see the model rankings, and to weight the numbers by a specific job see the use-case pages.
Each image is scored 0 to 5 per criterion. Measured criteria (resolution, latency, cost, transparency, color accuracy) are computed automatically; quality criteria (text accuracy, prompt adherence, composition) are judged by an automated Claude vision tier against each prompt’s checklist. Page scores are unweighted means over all of a model’s runs in that scope. Blind expert-panel review has not run yet; the dataset is pilot-0.
Model output is a starting point, not a finished asset. Production work usually needs background removal, exact brand colors or editable text, none of which generation guarantees on every run. The IMG.LY AI Editor gives your users those controls to refine any model’s output to production quality.
The suite re-runs on notable model releases so the rankings stay current. The results shown are the pilot-0 dataset, scored by measured criteria plus an automated Claude vision tier, with blind expert review planned. Prompts are versioned and append-only, so scores stay comparable across runs.

