Head to head
Esc

Models · slot A

to navigate to pick
Arena / Challenges

Text-to-Image · Creativity challenges

Every text-to-image challenge testing Creativity, scored with TrueSkill as the votes come in.

Portraits Cartoon, Anime & Fantasy

Fantasy Warrior

A character-portrait test focused on fine materials. It checks how well a model renders engraved metal under warm torchlight, layered textures (leather, cloth, braided hair), and lifelike eyes and skin without drifting into a plastic, over-smoothed look.

Best
ImagineArt 1.5 (Preview)
Mid
FLUX.2 [max]
Worst
LongCat-Image
3D Imaging & Modeling Text Rendering

Isometric Miniature Diorama Scenes

Combines a strict isometric viewpoint, clean PBR-style materials, and embedded text in one tightly framed scene. It tests whether a model holds a consistent 45-degree perspective, renders the labels accurately, and keeps the composition centered and uncluttered.

Best
Wan 2.5 (Preview)
Mid
Wan 2.7 Pro
Worst
FLUX.1 [schnell] FP8
Text Rendering Product, Branding & Commercial

Vintage Cafe Logo

A logo and branding test. It checks accurate text (the accented Caffè and Est. 1720), a clean vector-emblem look rather than a photographic one, and a cohesive vintage style. Small text and crisp flat shapes are common weak spots for image models.

Best
GPT Image 1.5
Mid
GPT Image 1
Worst
Vidu Q2
Text Rendering Product, Branding & Commercial

Apollo 11: Journey to Tranquility

A structured infographic test. The model must place six labeled steps in order, hold a consistent flat-vector icon style, and render several short labels (step names, crew names, Tranquility) accurately. It rewards layout discipline and punishes garbled text.

Best
FLUX.2 [pro]
Mid
FLUX.1 [schnell]
Worst
Wan 2.7 Pro
Art Photorealism

The Reversed Rodeo

This competition tests how well AI image models truly understand language versus how much they rely on visual habits from their training data. The prompt is deliberately simple on the surface but devilishly hard in practice. Most models default to the familiar trope of an astronaut riding a horse. By forcing the reversal, we measure three critical capabilities that separate good models from great ones: Strict instruction following (including negations) Accurate subject-object relationships and spatial hierarchy Resistance to strong dataset biases

Best
Imagen 3.0 Generate 002
Mid
FLUX.2 [pro]
Worst
Nano Banana 2
Text Rendering Art

The Halloween Invitation

While rendering simple text is no longer a real challenge for today’s SOTA models, this test shows something more interesting: how much visual taste a model has, and whether it can create a layout that feels like it came from a professional designer instead of a basic Canva template (no offense to Canva).

Best
Nano Banana 2
Mid
Vidu Q2
Worst
Recraft V4 Pro