Head to head
Esc

Models · slot A

to navigate to pick
Arena / Challenges

Text-to-Image challenges

Every text-to-image challenge in the arena, scored with TrueSkill as the votes come in. Filter by skill to narrow it down.

3D Imaging & Modeling Photorealism

Geometric Composition

A spatial-reasoning test. Each object has a precise relationship (inside, on top, behind, seen through the glass), so it measures whether a model follows explicit placement instructions and handles transparency and refraction rather than just approximating the scene.

Best
ImagineArt 1.5 (Preview)
Mid
FLUX.1 Kontext [dev]
Worst
Wan 2.7 Pro
Portraits Cartoon, Anime & Fantasy

Fantasy Warrior

A character-portrait test focused on fine materials. It checks how well a model renders engraved metal under warm torchlight, layered textures (leather, cloth, braided hair), and lifelike eyes and skin without drifting into a plastic, over-smoothed look.

Best
ImagineArt 1.5 (Preview)
Mid
FLUX.2 [max]
Worst
LongCat-Image
3D Imaging & Modeling Text Rendering

Isometric Miniature Diorama Scenes

Combines a strict isometric viewpoint, clean PBR-style materials, and embedded text in one tightly framed scene. It tests whether a model holds a consistent 45-degree perspective, renders the labels accurately, and keeps the composition centered and uncluttered.

Best
Wan 2.5 (Preview)
Mid
Wan 2.7 Pro
Worst
FLUX.1 [schnell] FP8
Text Rendering Product, Branding & Commercial

Vintage Cafe Logo

A logo and branding test. It checks accurate text (the accented Caffè and Est. 1720), a clean vector-emblem look rather than a photographic one, and a cohesive vintage style. Small text and crisp flat shapes are common weak spots for image models.

Best
GPT Image 1.5
Mid
GPT Image 1
Worst
Vidu Q2
Text Rendering Product, Branding & Commercial

Modern Clean Menu

A layout and typography test. The model has to produce a clean, professional menu with legible text, a clear three-section grid, and food imagery that looks intentional, which separates real design sense from a default template look.

Best
Grok Imagine Image
Mid
Qwen Image 2.0
Worst
Imagen 3.0 Generate 002
Text Rendering Product, Branding & Commercial

Apollo 11: Journey to Tranquility

A structured infographic test. The model must place six labeled steps in order, hold a consistent flat-vector icon style, and render several short labels (step names, crew names, Tranquility) accurately. It rewards layout discipline and punishes garbled text.

Best
FLUX.2 [pro]
Mid
FLUX.1 [schnell]
Worst
Wan 2.7 Pro
Photorealism Aesthetics

Adorable Baby Animals in Sunny Meadow

A multi-subject photorealism test. Four distinct animals must each keep correct anatomy and detailed fur while interacting naturally in one scene under warm directional light. The usual failure points are blended or malformed animals and fur that turns mushy.

Best
Imagen 4.0 Fast Generate 001
Mid
Imagen 4.0 Generate 001
Worst
Nano Banana 2 Lite
Photorealism Aesthetics

Candid Street Photography

Tests whether a model can produce a believable documentary photo instead of a polished render. The difficulty is in natural skin and lighting, wet-pavement reflections, and the imperfect framing and motion blur that real street photography has but models tend to smooth away.

Best
GPT Image 1.5
Mid
Imagen 4.0 Ultra Generate 001
Worst
Recraft V4.1 Pro
Text Rendering Photorealism

Magic Burger Explosion: Fiery Photorealism Challenge

This prompt forces models to simultaneously nail a highly specific, multi-layered commercial scene: dynamic exploded composition with multiple flying food elements, photorealistic textures, dramatic fiery lighting with embers, and precisely integrated glowing text, all while keeping strong visual impact. It is a perfect stress test that quickly separates models with true prompt mastery and creative control from those that miss details, break physics, or produce generic results.

Best
GPT Image 2
Mid
Nano Banana
Worst
FLUX.1 [schnell] FP8
Photorealism Prompt Adherence

The Capybara Taxi Driver

This challenge seems to be difficult for models because it mixes reality with fiction. Most models struggle to keep the taxi realistic or loose instructions like placing the passenger not in the backseat.

Best
Z-Image Turbo
Mid
Recraft V4
Worst
Qwen Image
Text Rendering Photorealism

Chalkboard Menu

This challenge forces models to use one consistent handwritten style across an entire dense menu instead of defaulting to clean printed text for the smaller details, a very common failure that reveals how well they actually understand and maintain stylistic coherence.

Best
Grok Imagine Image
Mid
Nano Banana 2
Worst
Stable Diffusion 3.5 Large Turbo
Art Photorealism

The Reversed Rodeo

This competition tests how well AI image models truly understand language versus how much they rely on visual habits from their training data. The prompt is deliberately simple on the surface but devilishly hard in practice. Most models default to the familiar trope of an astronaut riding a horse. By forcing the reversal, we measure three critical capabilities that separate good models from great ones: Strict instruction following (including negations) Accurate subject-object relationships and spatial hierarchy Resistance to strong dataset biases

Best
Imagen 3.0 Generate 002
Mid
FLUX.2 [pro]
Worst
Nano Banana 2