Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 2 OpenAI Qwen Image Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

GPT Image 2

28.1 arena score

#3 of 62 in Text-to-Image

Top 3 in Text-to-Image
Skill signature · Text-to-Image

Qwen Image

21.4 arena score

#35 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 2

0%

win rate

Ties

0%

Qwen Image

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 2
Qwen Image

AI Judge Analysis

GPT Image 2

  • + Excellent photographic texture on the book cover and wooden table.
  • + Highly realistic refraction and transparency in the glass cube.
  • + Accurate handling of the plant's visibility through the glass.
  • The glass cube has internal structural seams that make it look a bit like a frame rather than a solid glass object.

Qwen Image

  • + Successfully includes all requested elements in a clean composition.
  • + Good reflection of the blue sphere on the bottom surface of the cube.
  • + Soft window lighting is accurately portrayed.
  • Lower overall resolution and sharpness compared to Image A.
  • The plant's appearance through the glass is slightly inconsistent with the leaves behind it.
  • The wooden texture on the table is less detailed.

Verdict: GPT Image 2 is the superior image due to its exceptional level of detail, particularly in the textures of the red book and the wooden table. While both models followed the prompt perfectly, Qwen (Image B) has a softer, slightly more synthetic look, whereas GPT Image 2 achieves a convincing photographic quality with better light interaction.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 2
Qwen Image

AI Judge Analysis

GPT Image 2

  • + Excellent skin texture and hyper-realistic facial details
  • + Accurate representation of 1990s-style Japanese city streets and signage
  • + Natural 'imperfect' framing that feels like a real candid photo
  • The rain effect is very subtle, almost difficult to see
  • The motion blur on the car in the background is slightly inconsistent with the depth of field

Qwen Image

  • + Stronger visual representation of 'light rain' and atmospheric mood
  • + Good use of reflections on the wet pavement
  • + Solid composition with clear depth of field
  • The facial features and hair look slightly artificial and smoothed out
  • The bicycle has structural issues, such as a kickstand floating in mid-air
  • Lower image resolution and loss of fine texture compared to Model A

Verdict: GPT Image 2 (Model A) is the superior image due to its incredible attention to detail, realistic skin textures, and authentic Japanese street setting. While Qwen (Model B) captures the rainy atmosphere better, it suffers from typical AI artifacts like floating objects and a less realistic, 'cleaner' aesthetic that clashes with the 'no stylization' requirement.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 2
Qwen Image

AI Judge Analysis

GPT Image 2

  • + Exceptional photorealistic skin texture and eye detail
  • + Incredibly intricate engraving and weathered patina on the armor
  • + Natural-looking braids with subtle, believable bead integration
  • The 'bokeh sparks' are extremely subtle, almost unnoticeable
  • Background is slightly generic

Qwen Image

  • + Strong adherence to the 'beads' and 'bokeh sparks' prompt elements
  • + Good depiction of scarred skin and varied textures
  • + High contrast lighting captures the torch effect well
  • The sparks look like artificial icons rather than natural lens bokeh
  • Skin and hair textures are slightly smoothed and 'CG' in appearance
  • Beads look more like modern plastic carnival beads than historical/fantasy accessories

Verdict: GPT Image 2 produces a significantly more cinematic and realistic result with superior texture work on the skin and metal. While Qwen Image followed the specific prompt for beads and sparks more literally, it resulted in a less coherent and more artificial aesthetic compared to the lifelike quality of GPT Image 2.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 2
Qwen Image

AI Judge Analysis

GPT Image 2

  • + Exceptional text rendering with coherent item names, descriptions, and prices.
  • + Highly professional layout that fulfills all prompt requirements for specific sections.
  • + High-quality, distinct food photography for every menu item.
  • The layout is very dense, which may slightly push the boundaries of 'minimalist' compared to a simpler grid.

Qwen Image

  • + Strong minimalist aesthetic with a clear grid-based design.
  • + Interesting use of vibrant background colors behind food photography.
  • Text is largely gibberish and poorly rendered.
  • Failed to create three distinct sections, merging pizza and mains into one header.
  • Lower resolution food images with repeating visual elements.

Verdict: GPT Image 2 is the clear winner as it produces a fully functional, professional-grade menu with legible text and high-quality photography. Qwen Image follows the minimalist grid prompt well but fails significantly on text legibility and logical sectioning.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 2
Qwen Image

AI Judge Analysis

GPT Image 2

  • + Excellent adherence to the fiery/glowing text prompt across all three elements.
  • + High-quality textures and realistic lighting on food components.
  • + Strong dynamic composition with a clear 'exploded' view.
  • The text takes up a very large portion of the frame compared to the product.

Qwen Image

  • + Clean, readable text layout that feels more like a commercial poster.
  • + Bright, appetizing colors on the burger components.
  • Failed to render the 'LIMITED TIME ONLY' and '€6.99' text with the requested fiery, glowing effect.
  • The burger is less 'exploded' and more of a whole burger with a few extra pieces floating around it.
  • The price tag lacks the fiery glow requested in the prompt.

Verdict: GPT Image 2 followed the complex styling instructions much more accurately, successfully applying the fiery, glowing effect to all three requested text elements and the starburst. While Qwen Image produced a cleaner layout, it failed to style the secondary text and price according to the prompt, and the 'exploded' effect on the burger was less dramatic. GPT Image 2 is the clear winner for its superior prompt adherence and photorealistic detail.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 2
Qwen Image

AI Judge Analysis

GPT Image 2

  • + Excellent chalk texture with realistic dusty edges and pressure variations.
  • + Perfect adherence to the requested text and date.
  • + Highly realistic cursive handwriting that truly looks manually written.
  • The lighting in the upper left corner slightly washes out a small portion of the text.

Qwen Image

  • + Clear, legible text layout.
  • + Good use of space on the board.
  • Failed the date requirement by writing '20026' instead of '2026'.
  • The text looks like a smooth digital font with a slight glow rather than realistic chalk.
  • Text lacks the authentic chalk texture and physical grit requested in the prompt.

Verdict: GPT Image 2 is the clear winner as it masterfully captured the requested chalk texture, specific handwriting style, and accurate text. Qwen Image failed on technical details, including a typo in the year and a failure to render realistic chalk physics, resulting in a look that feels more like a digital overlay.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 2
Qwen Image

AI Judge Analysis

GPT Image 2

  • + Excellent adherence to the 'horse on top' spatial instruction
  • + High textural detail in the spacesuit and lunar surface
  • + Clever integration of riding tack including stirrups for the horse's hooves
  • The astronaut's hands/gloves have a somewhat unnatural, finger-heavy appearance

Qwen Image

  • + Good cinematic lighting and planet design
  • + Clean rendering of the astronaut and horse components
  • Completely failed the negative constraint to have the horse on top
  • Anatomical issues with the horse's legs appearing disjointed

Verdict: GPT Image 2 followed the specific and unusual instruction to place the horse on top of the astronaut perfectly, creating a truly surreal image. Qwen Image produced a generic 'astronaut on a horse' image, failing the primary challenge of the prompt while also exhibiting several anatomical errors in the horse's legs.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 2
Qwen Image

AI Judge Analysis

GPT Image 2

  • + Excellent fur texture and lighting integration for the capybara
  • + High photographic realism in the interior and background bokeh
  • + Accurate depiction of a professional 'NY taxi driver' expression
  • The passenger's hands and phone interaction are slightly blurry and less defined

Qwen Image

  • + Clear composition showing the full scene including the taxi's exterior and interior
  • + Good adherence to the 'bored' expression requested for the passenger
  • + Includes the exterior taxi roof light for additional context
  • The capybara's hands look more like monkey hands/paws than capybara paws
  • The capybara's head is fused strangely with the car seat headrest
  • Internal lighting is a bit flat compared to the realism of Model A

Verdict: GPT Image 2 is the superior image due to its exceptional photorealism and better integration of the capybara into the scene. While Qwen Image provides a wider view, it suffers from several anatomical and perspective errors, such as the incorrectly shaped paws and the headrest merging into the capybara's fur.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 2
Qwen Image

AI Judge Analysis

GPT Image 2

  • + Excellent typography with perfect spelling for all requested text.
  • + Highly detailed and cohesive vintage aesthetic with intricate borders and cinematic lighting.
  • + Smart inclusion of a bridge background to represent 'The Arches' location.
  • The parchment background is very dark, which might make some side details harder to see.

Qwen Image

  • + Strong contrast between the parchment edge and the dark central scene.
  • + Follows the general composition requested, including the pumpkin and twisted trees.
  • Significant spelling errors in the main title ('Halle Party Invitation') and extra gibberish text at the top.
  • Lower level of artistic detail compared to Model A.
  • The 'scroll' banner is broken into two pieces, reducing the polished look.

Verdict: GPT Image 2 is the clear winner as it successfully rendered all the requested text with zero spelling errors and maintained a high-quality, professional aesthetic. Qwen Image failed on the primary title text and lacked the intricate, cinematic depth found in the first image.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 2
Qwen Image

AI Judge Analysis

GPT Image 2

  • + Excellent PBR textures and lighting that create a high-end 3D render look.
  • + Perfect text rendering and placement as requested.
  • + Highly detailed and realistic sushi models within the diorama context.
  • The diorama content is quite crowded, pushing the definition of 'minimal garnish'.

Qwen Image

  • + Successfully captures the requested 'cartoon miniature' style with a soft, clean aesthetic.
  • + Good adherence to the diorama base and minimal garnish instructions.
  • + Correct 45° isometric perspective.
  • Text rendering is slightly messy with overlapping and uneven spacing.
  • The small flag icon in the text section is distorted and lacks the clean execution of Model A.

Verdict: GPT Image 2 is the superior output because it perfectly executes the complex text requirements while delivering high-quality PBR textures that still maintain a stylized miniature feel. Qwen Image adheres better to the 'minimal' aspect of the prompt but fails on the technical execution of the typography and icon clarity.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 2
Qwen Image

AI Judge Analysis

GPT Image 2

  • + Excellent depiction of dynamic movement with realistic paws and poses
  • + Very realistic fur texture and lighting integration on all animals
  • + Includes all requested animals with distinct, accurate features
  • The fox's limb positioning looks a bit awkward in the mid-leap
  • The background butterflies are a bit blurry compared to the foreground

Qwen Image

  • + Beautiful dew sparkles and soft bokeh effect in the grass
  • + Cute, expressive eyes on all animals
  • + Great realization of the sunset god rays
  • The fox has a cat-like paw structure and slightly stylized features
  • The golden retriever puppy looks static compared to the others
  • Some minor blending artifacts where the rabbit meets the dog

Verdict: GPT Image 2 (Model A) is the winner because it better captures the 'playfully chasing' and 'tumbling together' aspect of the prompt with more realistic anatomy and physics. While Qwen Image (Model B) has a lovely atmosphere with its dew sparkles, GPT Image 2 feels more like a cohesive, high-energy scene with superior fur detailing and more accurate animal characteristics.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 2
Qwen Image

AI Judge Analysis

GPT Image 2

  • + Excellent typography with perfect spelling of 'Caffè Florian'.
  • + Sophisticated vintage engraving style with high-quality cross-hatching.
  • + Balanced composition with a professional emblem frame and banner.
  • Slightly more ornate than 'minimalist', though it fits the 'vintage' prompt perfectly.

Qwen Image

  • + Stronger adherence to the 'minimalist' aspect of the prompt.
  • + Correct color palette of warm browns and cream.
  • Severe typography errors with cluttered, overlapping text ('FLoraIAN').
  • The graphic style is a bit too simplified and lacks the requested 'classic' feel.

Verdict: GPT Image 2 is the clear winner as it produces a professional-grade logo with perfect text rendering and beautiful vintage detailing. Qwen Image fails significantly on the typography, creating an unreadable mess of letters for the restaurant name, despite capturing a more minimalist aesthetic.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 2
Qwen Image

AI Judge Analysis

GPT Image 2

  • + Excellent typography and legible text rendering throughout the infographic.
  • + Highly detailed and accurate visual representations of the Saturn V and Lunar Module.
  • + Comprehensive layout that follows all six requested steps in a logical sequence.
  • The illustration style leans more toward 3D rendering than the 'flat-vector' style requested.
  • Slightly more complex gradients than the 'subtle' ones requested.

Qwen Image

  • + Accurately captures the requested flat-vector aesthetic with simple shapes.
  • + Adheres strictly to the NASA-inspired color palette requested in the prompt.
  • Severe spelling and text rendering errors ('ApolL', 'Sarth Orbit', 'Arnnstrong').
  • The composition is cluttered and fails to logically follow the requested six-step sequence.
  • Included meta-instructions in the final image text like '(Stop at landing)'.

Verdict: GPT Image 2 is the clear winner as it produces a professional, high-quality infographic with perfect spelling and logical flow. While Qwen followed the 'flat-vector' style more closely, its failure in basic text rendering and inability to organize the six-step sequence makes it unusable as an infographic.

Next steps

Explore each model