Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 1.5 OpenAI Qwen Image 2.0 Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

GPT Image 1.5

27.1 arena score

#5 of 62 in Text-to-Image

Top 3 in Image Editing
Skill signature · Text-to-Image

Qwen Image 2.0

21.7 arena score

#34 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 1.5

0%

win rate

Ties

0%

Qwen Image 2.0

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 1.5
Qwen Image 2.0

AI Judge Analysis

GPT Image 1.5

  • + Perfectly captures the physics of the scene with the sphere resting on the bottom surface.
  • + Refractions through the glass cube and reflections on the table are highly realistic.
  • + The lighting accurately comes from the left as requested, creating soft shadows.
  • The plant in the background is quite dense, which slightly obscures the clarity of the glass cube concept.

Qwen Image 2.0

  • + Excellent texture work on the book cover and the pages of the book.
  • + The plant is well-positioned and clearly visible through the glass panels.
  • The sphere is floating in the center of the cube without any support, which feels unnatural for the requested scene.
  • The reflection/refraction of the sphere appears duplicated and physically inconsistent on the side panels.

Verdict: GPT Image 1.5 is the superior image because it follows the laws of physics, showing the blue sphere resting naturally inside the cube, whereas Qwen Image 2.0 has the sphere floating inexplicably. GPT Image 1.5 also demonstrates more realistic light behavior and glass refractions, creating a more cohesive and believable scene.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 1.5
Qwen Image 2.0

AI Judge Analysis

GPT Image 1.5

  • + Excellent photographic atmosphere with wet textures and rain visibility.
  • + Stronger cinematic lighting and reflection quality on the pavement.
  • + Anatomy and tool interactions are highly realistic.
  • The car in the background is quite sharp, lacking the requested motion blur.
  • The framing feels a bit too balanced for 'imperfect framing'.

Qwen Image 2.0

  • + Better adherence to 'imperfect framing' with a more candid, snap-shot composition.
  • + Skin textures are highly detailed and naturalistic.
  • + Includes motion blur on the passing vehicle as requested.
  • Physical logic errors with the bicycle, such as the chain and pedal placement.
  • The man's hands have structural inconsistencies (too many fingers/blended joints).

Verdict: GPT Image 1.5 produces a much more coherent and visually stunning image with superior lighting and realistic mechanical details. While Qwen Image 2.0 followed prompts for motion blur and framing more literally, it suffered from significant AI artifacts in the hands and bicycle structure.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 1.5
Qwen Image 2.0

AI Judge Analysis

GPT Image 1.5

  • + Exceptional photographic realism in skin texture and eyes
  • + Masterful use of warm lighting and bokeh to create atmosphere
  • + Incredibly detailed engraving on the armor and texture on the fabric/leather
  • The hair beads are somewhat simple and look slightly fused with the hair in places

Qwen Image 2.0

  • + Strong character design with distinct beaded braids
  • + Good representation of ornate armor and battle scars
  • + Vibrant colors and clear background story elements
  • The lighting looks artificial and does not realistically reflect off the skin
  • Anatomical issues with the hand holding the sword pommel
  • Eyes have an unnatural, glowing quality that detracts from realism

Verdict: GPT Image 1.5 delivers a stunningly realistic portrait with cinematic lighting and intricate textures that align perfectly with the prompt's request for lifelike eyes and detailed armor. Qwen Image 2.0 provides an interesting character interpretation but suffers from lighting inconsistencies and anatomical errors in the hand. GPT Image 1.5 is the clear winner for its superior technical execution and believable presence.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 1.5
Qwen Image 2.0

AI Judge Analysis

GPT Image 1.5

  • + Excellent text readability and logical content
  • + High-quality, appetizing food photography
  • + Professional UI/UX layout that follows standard menu design principles
  • The grid of images is slightly asymmetrical compared to the text blocks

Qwen Image 2.0

  • + Strong minimalist aesthetic with a consistent image grid
  • + Appearing more like a modern digital app menu
  • Gibberish text making the menu unusable
  • Incorrect categorization, such as showing pizza under the 'Mains' column
  • Repetitive pricing and lack of descriptive text

Verdict: GPT Image 1.5 produces a fully functional and professional menu with clear text, realistic pricing, and relevant food photography. Qwen Image 2.0 has a cleaner grid layout but fails significantly on text generation and logical organization, placing pizzas under the 'Mains' heading and using illegible characters.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 1.5
Qwen Image 2.0

AI Judge Analysis

GPT Image 1.5

  • + Excellent photorealistic texture on the meat and bun.
  • + Vibrant and dynamic composition with sparks and embers filling the space.
  • + Clear and professional typography that integrates well with the theme.
  • The 'exploded' effect is slightly more clumped than floating individual layers.

Qwen Image 2.0

  • + Strong 'exploded' burger effect with clear separation between components.
  • + Effective use of literal fire within the text design.
  • + Good adherence to the starburst pricing requirement.
  • The burger ingredients look slightly more synthetic and less appetizing than Model A.
  • The background feels more like a flat black backdrop with some smoke rather than a fiery environment.

Verdict: GPT Image 1.5 produced a much more appetizing and professional advertisement with superior lighting and textures. While Qwen Image 2.0 did a great job separating the layers of the burger, the overall visual quality and fiery atmosphere of GPT Image 1.5 feel more cohesive and high-end.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 1.5
Qwen Image 2.0

AI Judge Analysis

GPT Image 1.5

  • + Perfect text accuracy for all requested items and prices.
  • + Highly realistic chalk texture with dusty smudges and varying pressure.
  • + Excellent handwriting consistency that feels human rather than a font.
  • Lacks the environmental context of the 'cozy café' mentioned in the prompt, focusing only on the board.
  • Slightly cramped composition toward the bottom.

Qwen Image 2.0

  • + Successfully includes the 'cozy café' background with depth of field.
  • + Strong composition making use of the board's physical presence in a space.
  • + Accurate spelling and character rendering.
  • The text layout is slightly awkward with large gaps before the prices.
  • The handwriting looks a bit more like a digital brush than physical chalk compared to the other model.

Verdict: GPT Image 1.5 wins on technical execution of the chalkboard itself, providing a texture and handwriting style that is virtually indistinguishable from a real photo. While Qwen Image 2.0 did a better job with the environmental setting by showing the café interior, GPT Image 1.5's superior chalk realism and cleaner alignment of the text make it the better choice for this specific prompt.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 1.5
Qwen Image 2.0

AI Judge Analysis

GPT Image 1.5

  • + Excellent cinematic lighting and texture on the spacesuit and horse fur.
  • + High level of detail in the lunar surface and background celestial bodies.
  • + Dynamic composition with a sense of motion and grit.
  • The astronaut is on top of the horse, failing the specific spatial instruction 'horse on top, not vice versa'.

Qwen Image 2.0

  • + Beautiful surreal atmosphere with shimmering horse skin and floating droplets.
  • + Clean, sharp rendering of the astronaut and planetary background.
  • + Better usage of negative space to create a 'space' feel.
  • Fails the specific prompt instruction 'horse on top, not vice versa'.
  • Anatomical issues with the horse's front legs.

Verdict: Both models completely failed the negative constraint and specific spatial instruction to place the horse on top of the astronaut, instead providing the more common 'astronaut riding a horse' trope. GPT Image 1.5 is the preferred image because its visual quality is significantly higher, featuring cinematic lighting and complex textures that better fit the 'highly detailed' requirement compared to the slightly more plastic look of Qwen Image 2.0.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 1.5
Qwen Image 2.0

AI Judge Analysis

GPT Image 1.5

  • + Excellent photorealistic lighting and textures in the capybara's fur and jacket.
  • + Stronger cinematic composition that captures both characters clearly.
  • + Superior interior detailing, including the dashboard meter and realistic upholstery.
  • The passenger's eyes and facial features are slightly distorted under zoom.

Qwen Image 2.0

  • + Clearer view of the capybara's paws on the steering wheel.
  • + Good background bokeh effect with Manhattan city lights.
  • The passenger is seated too far forward, appearing to be in the front seat or mid-air rather than the back seat.
  • The blue car seats and interior feel less like a traditional New York taxi cabinet.
  • Noticeable anatomy artifact on the passenger's left hand.

Verdict: GPT Image 1.5 is the superior image due to its much more realistic and detailed rendering of the capybara and the authentic New York taxi interior. Qwen Image 2.0 fails on the spatial layout, placing the passenger awkwardly in the front of the vehicle and having more technical artifacts in the character rendering.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 1.5
Qwen Image 2.0

AI Judge Analysis

GPT Image 1.5

  • + Excellent typography style that feels integrated into the vintage aesthetic.
  • + Beautiful cinematic lighting and atmospheric painterly quality.
  • + Perfect adherence to all requested text and visual elements.
  • The 'Date' text at the bottom is slightly smaller and more compressed than other text elements.

Qwen Image 2.0

  • + Clear, legible text rendering across all sections.
  • + Strong central composition with a bold jack-o-lantern.
  • + Includes nice stylistic gothic fonts for the heading.
  • The background and foreground elements feel somewhat disjointed, like a digital collage.
  • The border appears slightly more generic compared to the intricate thorns in Model A.
  • Lighting is flatter and less 'cinematic' than the other version.

Verdict: GPT Image 1.5 produced a much more cohesive and atmospheric invitation that truly captures the 'vintage' and 'dark parchment' feel requested. While Qwen Image 2.0 followed all instructions perfectly, its visual style felt more like a modern digital graphic than a vintage poster, whereas GPT Image 1.5 achieved a sophisticated artistic finish.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 1.5
Qwen Image 2.0

AI Judge Analysis

GPT Image 1.5

  • + Expertly captures the isometric miniature diorama aesthetic with a layered base.
  • + Excellent text rendering with clean, stylized typography centered as requested.
  • + Rich, high-quality PBR-like textures on the wooden board and ceramic teapot.
  • Includes more elements (teapot, soy sauce) than the 'minimal' request asked for.
  • The flag icon is slightly simple compared to the high-detail 3D scene.

Qwen Image 2.0

  • + Follows the 'minimal garnish and plate' instruction closely.
  • + Accurate 45-degree angle and solid blue background.
  • + Clear, bold text centered at the top.
  • Lacks the 'miniature 3D cartoon' style, looking more like a standard product photo.
  • The 'diorama base' is just a simple cutting board rather than a stylized scene.
  • The text layout is less integrated and looks like a basic overlay.

Verdict: GPT Image 1.5 follows the artistic intent of the prompt significantly better, delivering a true 3D isometric diorama with stylized cartoon textures. While Qwen Image 2.0 is clean, it lacks the 'miniature' aesthetic and creativity found in the first image, appearing as a more generic food photograph. GPT Image 1.5 also showcases much stronger graphic design in the typography and flag placement.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 1.5
Qwen Image 2.0

AI Judge Analysis

GPT Image 1.5

  • + Features very high-quality lighting with distinct god rays and warm golden tones.
  • + Captures the 'big expressive eyes' and 'tumbling' aspects of the prompt very effectively.
  • + Includes all requested animals in a harmonious, intimate composition.
  • The fox kit has the same facial lighting and structure as the puppy, making them look slightly too similar.
  • The butterfly appears slightly flat against the background compared to the 3D quality of the animals.

Qwen Image 2.0

  • + Excellent action-oriented composition that looks more like a natural play session.
  • + Realistic fur textures and distinct anatomical features for each of the four animals.
  • + Better depth of field and variety in the wildflower meadow.
  • The fox kit is poorly positioned and appears to have an anatomical issue where its head meets its body.
  • The background lighting is a bit hazy and lacks the crisp 'masterpiece' feel of the animals in the foreground.

Verdict: GPT Image 1.5 produced a much more cohesive and aesthetically pleasing 'masterpiece' that perfectly captures the requested lighting and expressions. While Qwen Image 2.0 has great energy in its composition, the fox kit's anatomy is distorted, and it lacks the magical, high-clarity finishing seen in the GPT Image 1.5 version.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 1.5
Qwen Image 2.0

AI Judge Analysis

GPT Image 1.5

  • + Excellent typography with a sophisticated, professional font style.
  • + Accurate use of a banner and a secondary banner-like accent.
  • + Good vector-style shading and texture application.
  • Ignored the request for a light background, providing a black one instead.
  • The steam element is somewhat detached and small.

Qwen Image 2.0

  • + Correctly used a light background as specifically requested.
  • + The cloche is central and well-integrated into the design.
  • + Text is clear and highly legible.
  • The steam looks more like fire and is placed awkwardly inside the cloche handle area.
  • The banner on the bottom has an illogical fold/curl on the right side.
  • Typography is very basic and lacks the 'vintage' feel of a high-end cafe.

Verdict: GPT Image 1.5 produced a much more professional and aesthetically pleasing logo with superior typography, though it failed the background color instruction. Qwen Image 2.0 followed the background color prompt but struggled with a logical banner composition and had less sophisticated font choices. GPT Image 1.5 is preferred for its design quality despite the background error.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 1.5
Qwen Image 2.0

AI Judge Analysis

GPT Image 1.5

  • + Excellent layout with distinct, sequential panels for each step.
  • + Highly legible text with perfect spelling for all required stages.
  • + Consistent and professional vector illustration style with a clear color palette.
  • Slightly repetitive background elements in the space panels.
  • Translunar panel shows both a tiny moon and a large moon which is a bit confusing.

Qwen Image 2.0

  • + Elegant vertical layout that mimics a traditional poster.
  • + Includes additional thematic elements like the names of the astronauts in the corner.
  • + Clean, minimalistic design that adheres well to a modern NASA-inspired palette.
  • Spelling error in 'Translunjar'.
  • The Descent and Landing labels overlap the corresponding illustrations awkwardly.
  • Visual hierarchy is slightly weaker with elements floating on a single background.

Verdict: GPT Image 1.5 is the superior choice because it functions perfectly as an infographic, using clearly defined panels and flawless text rendering to tell the story of the mission. While Qwen Image 2.0 has a pleasing vertical composition, the spelling error 'Translunjar' and the cluttered overlap of text and icons make it less effective as a professional graphic.

Next steps

Explore each model