GenAI Benchmarks

Benchmark prompts

Every model runs the identical canonical prompt: same text, default parameters, a fixed seed where the API takes one, no per-model tuning. Each prompt stresses criteria that matter for production design work.

Suite pilot-0 · 37 prompts · 23 models
The benchmark avocado mascot in a lab coat, holding a checklist in front of a wall of framed test images

Showing 37 of 37 prompts