Head to head
Esc

Models · slot A

to navigate to pick
Your vote decides the leaderboard Pick the better image in blind matchups. Results update rankings in real time.
Start Voting

Prompt Adherence · Text-to-Image

Elo rankings from blind votes across 9 challenges in this category.

Highlights

10 images
Price · Performance

Best models, by

Every ranked model plotted by price against its Elo. The upper-left frontier is the best value: top quality at the lowest cost.

Best Text-to-Image Models by Price

Best AI Models for Prompt Adherence

61 models ranked · Last update: July 28, 2026 1:54 PM

Google’s Nano Banana Pro leads the category with a 1283 Elo and 77.7% win rate, while GPT Image 2 maintains the highest overall success rate at 84.8% for a lower 1276 Elo. Though current-generation models like FLUX.2 [max] and Nano Banana 2 dominate the top five, there is a significant performance gap compared to the much lower win rates of older models like DALL-E 3 (27.7%) and Stable Diffusion 3.5 Large (41.5%).

Rivalries

Aggregate head-to-head across the arena

Highlighted challenges

Best
GPT Image 2 is first to make it GPT Image 2
Mid
DALL-E 3
Worst
Stable Diffusion 3.5 Medium
Text-to-Image Art Photorealism

The Reversed Rodeo

This competition tests how well AI image models truly understand language versus how much they rely on visual habits from their training data. The prompt is deliberately simple on the surface but devilishly hard in practice. Most models default to the familiar trope of an astronaut riding a horse. By forcing the reversal, we measure three critical capabilities that separate good models from great ones: Strict instruction following (including negations) Accurate subject-object relationships and spatial hierarchy Resistance to strong dataset biases

Best
ImagineArt 1.5 (Preview)
Mid
Grok Imagine Image Pro
Worst
DALL-E 2
Text-to-Image 3D Imaging & Modeling Photorealism

Geometric Composition

A spatial-reasoning test. Each object has a precise relationship (inside, on top, behind, seen through the glass), so it measures whether a model follows explicit placement instructions and handles transparency and refraction rather than just approximating the scene.

Keep the arena honest

Cast your vote

Pick winners in blind matchups. Every vote nudges the Elo and shapes these rankings.

Cast Your Vote

Suggest a prompt

Got an idea worth testing? Submit a prompt and watch the models battle it out.