Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 1.5 OpenAI Qwen Image Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

GPT Image 1.5

27.1 arena score

#7 of 62 in Text-to-Image

Top 3 in Image Editing
Skill signature · Text-to-Image

Qwen Image

21.4 arena score

#35 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 1.5

0%

win rate

Ties

0%

Qwen Image

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 1.5
Qwen Image

AI Judge Analysis

GPT Image 1.5

  • + Excellent adherence to lighting instructions with a clear left-side window source
  • + Realistic glass thickness and internal reflections
  • + High-quality textures on the red book and wooden table
  • The sphere appears slightly large relative to the prompt 'small'

Qwen Image

  • + Clean, minimalist composition
  • + Correct placement of all objects according to the prompt
  • The glass cube appears to have a mirrored base which wasn't requested
  • The green plant is less visible through the glass compared to Model A
  • The sphere appears to be floating rather than resting on the surface

Verdict: GPT Image 1.5 is the winner as it handles the physics of light and glass much more realistically, particularly in how the plant is visible through the cube. Qwen Image produces a cleaner look but includes a mirrored bottom to the cube and lacks the depth of texture seen in the first image.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 1.5
Qwen Image

AI Judge Analysis

GPT Image 1.5

  • + Excellent adherence to the 'imperfect framing' and '50mm lens' look
  • + Highly realistic skin textures and wet surfaces
  • + Includes specific details like a tool box and realistic bicycle mechanics
  • The motion blur on the car is slightly static compared to the rain effect

Qwen Image

  • + Good inclusion of background motion blur on the white car
  • + Correctly captures the 'red bicycle' and 'elderly Japanese man'
  • The bicycle rendering is physically incoherent with a floating pedal and missing chain
  • The overall image quality feels somewhat AI-smoothed and less like a professional 50mm shot

Verdict: GPT Image 1.5 significantly outperforms the competitor by delivering a photographically believable image with realistic textures, lighting, and mechanical components. Qwen Image fails on technical details, specifically with the bicycle's anatomy which has floating parts and missing drivetrain components, making it look much less 'realistic'.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 1.5
Qwen Image

AI Judge Analysis

GPT Image 1.5

  • + Exceptional photographic realism with lifelike skin texture and eyes.
  • + Beautiful integration of warm torchlight and glowing bokeh embers.
  • + Intricate detail on the engraved armor, worn leather straps, and textured cloth.
  • The hair beads are metal rings rather than the traditionally expected decorative beads.

Qwen Image

  • + Accurately depicts specific colorful beads in the hair braids.
  • + Clear representation of a torch as the light source.
  • + Strong engraving details on the shoulder pauldron.
  • The facial scars look painted on rather than realistic skin texture.
  • The bokeh sparks appear as artificial geometric star shapes.
  • Lower overall visual fidelity compared to the rival image.

Verdict: GPT Image 1.5 provides a much more convincing and cinematic interpretation of the prompt, featuring superior skin textures, realistic lighting, and tactile material details. While Qwen Image followed the bead instruction with more literal color, its overall quality is hampered by artificial-looking scars and stylized spark effects that clash with the 'lifelike' requirement.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 1.5
Qwen Image

AI Judge Analysis

GPT Image 1.5

  • + Excellent text rendering with no spelling errors or hallucinations.
  • + Photorealistic food imagery that matches the specific menu items listed.
  • + Logical and professional hierarchy with clear pricing and descriptions.
  • The layout is a bit crowded with narrow margins.
  • The 'vibrant accents' are limited to simple colored underlines.

Qwen Image

  • + Strong minimalist aesthetic with a clear grid system for photos.
  • + Creative use of color blocking behind some grid items.
  • + Bold typography fits the modern sans-serif requirement.
  • Illegible, nonsensical text throughout the menu.
  • Food images are repetitive and less realistic than the competitor.
  • Failed to create three distinct sections for appetizers, pizza, and mains as requested.

Verdict: GPT Image 1.5 produces a functional, professional-grade menu with perfect text and accurate food representations. While Qwen Image has an interesting graphic design layout, it fails completely on text legibility and logical content organization, making it unusable for a real-world application.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 1.5
Qwen Image

AI Judge Analysis

GPT Image 1.5

  • + Excellent adherence to the 'exploded' request with detailed vertical separation of ingredients.
  • + Superior photorealistic textures on the meat patty, bun, and vegetables.
  • + Integrated the fiery, glowing effect into the typography and starburst perfectly as requested.
  • The overall composition is a bit crowded with many small floating particles.
  • The 'MAGIC BURGER' text is slightly clipped at the top edges.

Qwen Image

  • + Clean, professional graphic design layout with good breathing room.
  • + Highly legible secondary text ('LIMITED TIME ONLY').
  • + Effective use of depth of field in the background flames.
  • Failed the architectural part of the prompt; the burger is mostly assembled rather than 'exploded'.
  • The glowing effect on the text feels more like neon tubing than the requested fiery effect.
  • Less photorealistic detail in the food textures compared to the competitor.

Verdict: GPT Image 1.5 is the clear winner as it successfully rendered the 'exploded' burger effect with incredible photorealistic detail, whereas Qwen Image provided a mostly assembled burger. GPT Image 1.5 also followed the stylistic instructions for the text better, creating a cohesive fiery aesthetic that matches the background perfectly.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 1.5
Qwen Image

AI Judge Analysis

GPT Image 1.5

  • + Excellent text rendering with perfect spelling for all complex menu items.
  • + Authentic chalk texture with realistic dusty smudges and varying stroke opacity.
  • + Completed the third menu item correctly despite the truncated prompt.
  • The composition is a very tight close-up, lacking the 'cozy café' environmental context.

Qwen Image

  • + Good inclusion of the café environment in the background.
  • + Clean, legible handwriting with a pleasant slant.
  • Incorrect year rendered as '20026' instead of '2026'.
  • Spelling errors in menu items such as 'Risoto' instead of 'Risotto'.
  • Text appears more like a digital brush than realistic chalk on board.

Verdict: GPT Image 1.5 performed exceptionally well, accurately rendering all text and correctly inferring the full names of items even where the prompt was cut off. While Qwen Image provided a better sense of a café environment, it failed on basic spelling and accuracy, notably rendering the year incorrectly as 20026.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 1.5
Qwen Image

AI Judge Analysis

GPT Image 1.5

  • + Excellent surface textures with high-frequency detail in the space suit and horse's fur.
  • + Dynamic composition with a cinematic sense of motion and scale.
  • + Rich, complex background featuring celestial bodies, nebulae, and lunar hardware.
  • Completely failed the negative constraint to have the horse on top of the astronaut.

Qwen Image

  • + Clean, clear lighting and a minimalist, surreal aesthetic.
  • + Realistic anatomy for the horse and astronaut gear.
  • + Smooth rendering with minimal digital artifacts.
  • Completely failed the negative constraint to have the horse on top of the astronaut.
  • Less detailed background compared to the competitor.

Verdict: Both GPT Image 1.5 and Qwen Image failed the specific spatial logic instruction to place the 'horse on top', instead defaulting to the common trope of an astronaut riding a horse. GPT Image 1.5 is the superior image due to its significantly higher level of detail, complex lighting, and cinematic atmosphere, whereas Qwen Image feels somewhat generic by comparison.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 1.5
Qwen Image

AI Judge Analysis

GPT Image 1.5

  • + Excellent photorealistic fur texture and lighting integration.
  • + Correct taxi hat and steering wheel placement for a capybara.
  • + The businesswoman's bored expression perfectly captures the requested irony.
  • The paws on the steering wheel look slightly like human fingers in gloves rather than natural capybara feet.

Qwen Image

  • + Clean, modern composition with a clear view of the city background.
  • + Accurate jacket and name tag detail on the driver.
  • The paws are anatomically incorrect, looking more like primate hands.
  • Text rendering on the roof sign says 'YOXI' instead of 'TAXI'.
  • The capybara's head is not looking toward the road, creating an unnatural driving posture.

Verdict: GPT Image 1.5 is the clear winner due to its superior photorealistic quality and adherence to the 'bored' expression of the passenger. While Qwen Image provides a nice composition, it fails on text rendering and the capybara's anatomy, whereas GPT Image 1.5 feels like a real, gritier movie still from a New York street.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 1.5
Qwen Image

AI Judge Analysis

GPT Image 1.5

  • + Perfect text rendering for all requested strings including long sentences and dates.
  • + Highly detailed and consistent gothic aesthetic with a beautiful thorn and web border.
  • + Excellent cinematic lighting and composition that feels like a professional poster.
  • The parchment texture is very dark, which may slightly impact readability of the bottom text.

Qwen Image

  • + Successfully incorporates all requested elements like twisted trees, bats, and thorns.
  • + Clean layout with clear separation between the illustration and the document edges.
  • Significant typos in the main title, rendering it as 'Halloforty Invitation'.
  • The thorn border is a flat graphic rather than an integrated part of the cinematic scene.
  • Visual style is slightly more cartoonish and less 'vintage gothic' than requested.

Verdict: GPT Image 1.5 is the clear winner as it followed all text instructions perfectly, whereas Qwen Image failed significantly on the main title text. GPT Image 1.5 also exhibited a much more sophisticated artistic style that truly captured the 'vintage gothic' and 'cinematic' requirements of the prompt.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 1.5
Qwen Image

AI Judge Analysis

GPT Image 1.5

  • + Excellent PBR material rendering with realistic wood, ceramic, and food textures.
  • + Perfect text rendering and alignment for 'JAPAN' and 'SUSHI'.
  • + High level of detail in the miniature diorama elements like the teapot and soy sauce bottle.
  • The scene is a bit more realistic than the requested '3D cartoon' style.
  • The diorama base is quite busy compared to the 'minimal' request.

Qwen Image

  • + Perfectly captures the '3D cartoon' and 'soft refined texture' aesthetic requested.
  • + Better adherence to the 'minimal garnish and plate' instruction.
  • + Clean isometric composition with a clear miniature toy-like feel.
  • Text rendering is slightly less crisp and has a strange orange artifact under the flag icon.
  • Lower level of intricate detail compared to Model A.

Verdict: Both models followed the prompt very well, but they interpreted the style differently. GPT Image 1.5 leaned into a high-fidelity realistic miniature look with complex textures, while Qwen Image perfectly captured the soft, clean, 3D cartoon aesthetic requested, making it feel more like a stylized game asset.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 1.5
Qwen Image

AI Judge Analysis

GPT Image 1.5

  • + Excellent fur texture and fine detail rendering on all four animals.
  • + Superior lighting and atmosphere with convincing god rays and dew sparkles.
  • + Dynamic and joyful composition that perfectly captures the 'tumbling together' prompt.
  • The kitten has an anatomical anomaly with an extra-looking limb or paw placement under its chin.
  • The fox kit's face looks slightly more like a dog than a wild fox.

Qwen Image

  • + Clean composition with clearly defined subjects and a playful arrangement.
  • + Accurate representation of the specific species requested with distinct features.
  • + Good use of bokeh and sparkles in the meadow to create a magical atmosphere.
  • The golden retriever puppy appears much flatter and less detailed in its fur compared to the other animals.
  • The kitten's front right leg is anatomically distorted and looks disconnected from the body.
  • Lighting is a bit generic and lacks the 'hyper-photorealistic' depth requested.

Verdict: GPT Image 1.5 is the clear winner due to its superior texture rendering, atmospheric lighting, and overall image quality. While it has a minor anatomical quirk with the kitten, it far surpasses Qwen Image in terms of fur detail and the 'masterpiece' aesthetic requested. Qwen Image feels more like a standard stock photo and suffers from significant anatomical errors in the kitten's limbs and a lack of detail on the puppy.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 1.5
Qwen Image

AI Judge Analysis

GPT Image 1.5

  • + Excellent typography with perfect spelling and professional layout.
  • + High-quality vector aesthetic with sophisticated shading and texture.
  • + Correctly follows the vintage cloche and banner concept.

Qwen Image

  • + Applies a nice subtle paper texture to the background.
  • + Good alignment with the requested warm brown and cream color scheme.
  • Serious spelling and typography errors, overlapping text characters.
  • Loss of 'minimalist' feel due to messy letter placement.
  • The cloche illustration is very basic compared to its counterpart.

Verdict: GPT Image 1.5 is the clear winner as it produces a professional-grade logo with perfect spelling and elegant typography. While it used a black background instead of the light background requested, Qwen Image failed significantly on the text rendering, creating illegible and overlapping letters.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 1.5
Qwen Image

AI Judge Analysis

GPT Image 1.5

  • + Excellent text rendering with no spelling errors.
  • + Clear, logical panel-based layout following all six requested steps.
  • + Consistently applied flat-vector style with high-quality icons.
  • The rocket in the first panel is missing its top nose cone section.
  • The horizon in the Launch panel looks more like Mars than a NASA-inspired Earth red.

Qwen Image

  • + Successfully captured the requested NASA-inspired color palette.
  • + Good use of negative space in a single-page infographic style.
  • Numerous spelling errors in text (e.g., 'Sarth Orbit', 'Apoll', 'Aldlin').
  • Failed to include all 6 distinct steps, skipping some and mislabeling others.
  • Included prompt instructions like '(Stop at landing)' as actual text on the image.

Verdict: GPT Image 1.5 is significantly superior in this task, providing a professional-grade infographic with correct spelling, all requested narrative steps, and a clean layout. Qwen Image struggled with basic text rendering, hallucinated prompt instructions into the visual text, and failed to follow the logical sequence of the mission steps.

Next steps

Explore each model