Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 1 OpenAI Qwen Image Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

GPT Image 1

23.2 arena score

#29 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image

21.4 arena score

#35 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 1

100.0%

win rate

Ties

0.0%

Qwen Image

0.0%

win rate

100.0% 0.0% ties 0.0%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 1
Qwen Image

AI Judge Analysis

GPT Image 1

  • + Excellent depiction of the plant through the glass, showing distortion and color shifts.
  • + High detail on the textures of the book pages and the blue sphere.
  • + Accurate lighting that matches the 'soft window light from the left' instruction.
  • The glass cube appears to have a metal or reflective base rather than being fully glass.
  • The sphere seems to be floating rather than resting on the bottom surface.

Qwen Image

  • + Clean, minimalist composition that accurately places all requested items.
  • + The glass cube has realistic internal reflections and light refraction on the tabletop.
  • + Consistent depth of field with the plant correctly positioned in the background.
  • The plant is not visible *through* the glass cube as requested, but rather entirely behind its silhouette.
  • The blue sphere's reflection on the bottom surface suggests a mirrored base rather than clear glass.

Verdict: GPT Image 1 followed the complex instruction of showing the green plant through the glass much more effectively than Qwen Image, which obscured the plant behind the solid-looking book and cube frame. GPT Image 1 also featured superior texture detail and lighting, though Qwen Image produced a cleaner, more photographic look overall.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 1
Qwen Image

AI Judge Analysis

GPT Image 1

  • + Excellent skin texture and realistic facial details
  • + Strong cinematic atmosphere with effective shallow depth of field
  • + Natural-looking rain and wet surface reflections
  • The bike anatomy is slightly compressed to fit the frame
  • Motion blur on the background cars is subtle rather than pronounced

Qwen Image

  • + Full body composition shows the entire scene and bicycle
  • + Clear rain streaks and strong pavement reflections
  • + Follows the bicycle color prompt accurately
  • Faces and hands lack the fine detail and realism of the competitor
  • The white car in the background looks somewhat static and digital
  • Less successful at capturing the 'imperfect framing' requested, appearing more like a standard stock photo

Verdict: GPT Image 1 captures the emotional and technical requirements of the prompt far better, offering superior skin textures, realistic lighting, and a truly cinematic feel. Qwen Image provides a broader view, but it lacks the detail in the subjects' features and the 'candid' photographic quality requested.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 1
Qwen Image
100% wins 0% ties 0% wins

AI Judge Analysis

GPT Image 1

  • + Exceptional photographic realism in skin texture and eyes
  • + Natural integration of warm torchlight and subtle bokeh sparks
  • + Consistent engraving details on the armor
  • The beads in the hair are very subtle and blend in with the hair color
  • Missing visible leather straps mentioned in the prompt

Qwen Image

  • + Explicitly includes colorful beads, leather straps, and chainmail layers
  • + Higher contrast and more dynamic action-oriented lighting
  • + Clear depiction of the torch source
  • Face scars look like painted-on makeup rather than actual skin trauma
  • Sparks look like digital clipart rather than light artifacts
  • Anatomical issues with how the braid connects to the head

Verdict: GPT Image 1 produces a far more lifelike and cinematic portrait with superior skin rendering and realistic lighting, though it misses some specific textural elements like leather straps. Qwen Image includes more of the requested props (beads, straps) but suffers from artificial-looking scars and poorly integrated background effects. GPT Image 1 is the clear winner for its professional photographic quality and convincing 'battle-worn' aesthetic.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 1
Qwen Image

AI Judge Analysis

GPT Image 1

  • + Excellent high-resolution food photography that looks appetizing and professional.
  • + Clean, readable typography with clear section headers.
  • + Professional minimalist layout that frames the food photos perfectly.
  • Nonsense filler text for descriptions ('Apperoiation descrigion').
  • Partial crop of the bottom images.

Qwen Image

  • + Successful implementation of the grid layout for photos as requested.
  • + Captures the full 'page' layout of a menu with a brand header.
  • + Vibrant color accents used in the background of photo tiles.
  • Extremely poor text rendering with garbled characters and illegible fonts.
  • Low-quality, AI-distorted food images that look artificial.
  • Lacks the crisp, high-end feel of a professional dining menu.

Verdict: GPT Image 1 produces high-quality, professional food photography and a clean font that, despite some filler text, feels like a real menu. Qwen follows the grid layout prompt more literally but fails significantly on text legibility and image quality, resulting in a cluttered and unprofessional appearance. GPT Image 1 is the clear winner for its superior visual appeal and professional execution.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 1
Qwen Image

AI Judge Analysis

GPT Image 1

  • + Excellent photorealistic texture on the food components
  • + Perfect character spacing and glowing effects on all requested text
  • + Balanced composition with high-quality embers
  • The price tag contains a typo showing '.99' instead of '6.99'

Qwen Image

  • + Dynamic sense of motion with many small flying ingredients
  • + Correctly rendered the specific price and currency
  • + Vibrant and exciting fiery background
  • Text lacks the requested glowing effect for the secondary phrases
  • The burger itself is less an 'exploded view' and more a solid burger with side debris

Verdict: GPT Image 1 produces a superior 'exploded' view where the layers of the burger are clearly separated and the textures are highly photorealistic, but it fails the specific price text prompt. Qwen Image creates a more chaotic and energetic scene with perfect text accuracy, though the food components and overall lighting feel slightly more digital and less cohesive than its competitor.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 1
Qwen Image

AI Judge Analysis

GPT Image 1

  • + Excellent chalk texture on the individual letters
  • + Perfect spelling and date accuracy
  • + Consistent stylistic variations that feel truly handwritten
  • Failed to provide the 'elegant cursive' requested for the title
  • The framing is a tight crop rather than showing the 'cozy café' atmosphere

Qwen Image

  • + Successfully captured the cozy café background and full chalkboard context
  • + Dynamic layouts for the price points and menu items
  • + Good cursive-leaning title style as requested
  • Text error in the year rendering it as '20026'
  • The text looks more like a digital brush than realistic chalk texture
  • Noticeable artifacts and smudged text at the bottom of the board

Verdict: GPT Image 1 followed the complex text instructions perfectly, providing incredibly realistic chalk texture and accurate spelling, though it missed the environmental context of the café. Qwen Image provides a better overall composition with the requested cursive title, but it fails on technical details by hallucinating the year 20026 and lacks the authentic grainy texture of real chalk.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 1
Qwen Image

AI Judge Analysis

GPT Image 1

  • + High visual quality with a rich, cinematic color palette.
  • + Excellent textures on both the space suit and the horse's coat.
  • + Strong composition with a dynamic pose.
  • Completely failed the negative constraint: the astronaut is riding the horse.

Qwen Image

  • + Clear, high-resolution imagery with bright lighting.
  • + Accurate rendering of space suit details and horse anatomy.
  • Completely failed the negative constraint: the astronaut is riding the horse.
  • The composition feels slightly more generic and less 'surreal' than requested.

Verdict: Both GPT Image 1 and Qwen Image completely failed the unique prompt requirement of having the horse on top of the astronaut, instead providing standard interpretations of an astronaut riding a horse. GPT Image 1 is the preferred choice because it better captures the requested 'cinematic' and 'surreal' mood through its lighting and atmosphere, whereas Qwen Image looks like a standard digital composite.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 1
Qwen Image

AI Judge Analysis

GPT Image 1

  • + Excellent texture on the capybara's fur and the taxi cap.
  • + Stronger cinematic lighting that captures a moody New York night vibe.
  • + More accurate capybara anatomy compared to the other model.
  • The 'paws' on the steering wheel look slightly anthropomorphized and fuzzy.
  • The background passenger's face is quite blurry.

Qwen Image

  • + Clearer, more detailed background passenger with a perfect bored expression.
  • + Includes a taxi sign on top of the car for added context.
  • + Good composition with more of the vehicle interior visible.
  • The hands on the wheel are unsettlingly humanoid and do not resemble capybara paws.
  • The text on the roof sign is gibberish ('YOXI').
  • The capybara's neck and head connection to the jacket looks slightly less natural.

Verdict: GPT Image 1 produces a more convincing, cinematic photograph with superior animal textures and a more professional-looking capybara driver. While Qwen Image succeeds in rendering a clearer human passenger and more of the car's exterior, it is undermined by the highly distorted, humanoid hands on the steering wheel and misspelled text.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 1
Qwen Image

AI Judge Analysis

GPT Image 1

  • + Excellent typography with a cohesive gothic font choice
  • + All text requested is legible and correctly spelled
  • + Highly polished, cinematic lighting with a moody atmosphere
  • Merged the 'Time' and 'Location' lines incorrectly at the bottom
  • Composition is a bit top-heavy with the large title

Qwen Image

  • + Included the thorn border explicitly as requested
  • + All event details at the bottom are typed correctly on separate lines
  • + Used a clear parchment paper background effect
  • Title text contains spelling errors and overlapping letters
  • The extra line of decorative text at the top is gibberish
  • The scroll banner text is split awkwardly between the scroll and the background

Verdict: GPT Image 1 produces a much more professional and aesthetically pleasing invitation with superior typography and lighting. Although it made a small error in the layout of the bottom details, Qwen's multiple spelling errors in the main title and messy banner text make it less functional as a real invitation.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 1
Qwen Image

AI Judge Analysis

GPT Image 1

  • + Excellent text rendering with clean, centered typography.
  • + Higher fidelity PBR materials, especially visible in the realistic texture of the salmon and rice grains.
  • + Strong adherence to the '45° top-down isometric' and 'clean' prompt requirements.
  • The diorama base is a bit plain and bulky compared to the elegant garnish.

Qwen Image

  • + Vibrant, appealing color palette with a nicely layered diorama base.
  • + Includes more thematic elements like the physical flag and chopstick rest.
  • + Good 'cartoon' interpretation of the sushi shapes.
  • The word 'SUSHI' is slightly off-center and the flag icon is poorly integrated next to it.
  • The perspective on the chopsticks and the physical flag pole feels slightly inconsistent with the isometric view.
  • Lower overall clarity and texture detail compared to Model A.

Verdict: GPT Image 1 is the superior output because it perfectly adheres to the layout and typography requirements, delivering ultra-clean text and high-quality PBR textures that feel modern and polished. Qwen Image provides a charming diorama, but its text placement is less balanced and the overall image lacks the crisp resolution and material sophistication found in GPT Image 1.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 1
Qwen Image

AI Judge Analysis

GPT Image 1

  • + Excellent sense of motion and 'tumbling' as requested in the prompt.
  • + Superior integration of lighting and 'god rays' for a cinematic feel.
  • + Highly expressive and realistic animal anatomy.
  • The fox's front right paw is a bit blurry/poorly defined.

Qwen Image

  • + Features very clear, distinct butterflies with high-contrast patterns.
  • + Good use of bokeh and foreground dew sparkles in the meadow.
  • The composition feels static and posed rather than 'tumbling together'.
  • The kitten has an anatomical error with a third front leg/paw visible underneath it.
  • The lighting on the animals feels a bit flat compared to the intense background sun.

Verdict: GPT Image 1 successfully captures the kinetic energy and warmth requested by the prompt, showing the animals in motion with beautiful lighting. Qwen Image feels more like a static portrait and suffers from a significant anatomical glitch on the kitten, making GPT Image 1 the clear winner for quality and adherence.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 1
Qwen Image

AI Judge Analysis

GPT Image 1

  • + Excellent typography with correct spelling and accents.
  • + Clean minimalist vector aesthetic.
  • + Accurate adherence to the cloche and banner request.
  • Ignored the request for a light background, providing a black one instead.
  • Subtle texture is barely visible.

Qwen Image

  • + Followed the color scheme and light background request perfectly.
  • + Included the requested subtle texture on the background.
  • + Nice steam illustration above the cloche.
  • Severe spelling and typography errors, mangling the name into 'FLOraiAN'.
  • The banner construction is slightly clunky for a minimalist logo.

Verdict: GPT Image 1 is the superior choice for a logo because it maintains professional typography and spelling, which are critical for branding. While Qwen Image followed the background color instructions better, its failure to correctly render the text 'Caffè Florian' makes it unusable.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 1
Qwen Image

AI Judge Analysis

GPT Image 1

  • + Excellent vector iconography with a high-end graphic design feel.
  • + Precise adherence to the specified NASA-inspired color palette.
  • + Includes names of the astronauts with clean, professional typography.
  • Confused the prompt instructions, creating a misspelled label 'EARLLUNAR' and misaligning labels to icons.
  • Layout is a bit disjointed due to the placement of text boxes.

Qwen Image

  • + Better flow and narrative structure, following the sequence from Earth to the lunar surface.
  • + Includes a fairly accurate NASA logo recreation.
  • + Incorporates a clear numbering system for the mission steps.
  • Several spelling errors like 'Tranar Orbit' and 'Aldin'.
  • Poorly rendered astronaut icons and text at the bottom.
  • Vector quality is lower with some blurred edges around the rocket exhaust.

Verdict: GPT Image 1 produces much higher quality vector assets and a more sophisticated aesthetic, though it struggles with the layout and introduces a significant typo. Qwen Image follows the 'storyboard' flow of the prompt more logically but suffers from lower visual fidelity and numerous spelling mistakes. GPT Image 1 is the better graphic design piece despite the labeling alignment issues.

Next steps

Explore each model