Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 1 OpenAI Qwen Image Max Alibaba

Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.

GPT Image 1

22.6 arena score

#32 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image Max

21.3 arena score

#35 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 1

100.0%

win rate

Ties

0.0%

Qwen Image Max

0.0%

win rate

100.0% 0.0% ties 0.0%
Shared challenges 8

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 1
Qwen Image Max

AI Judge Analysis

GPT Image 1

  • + Excellent adherence to lighting instructions with natural shadows and highlights.
  • + High clarity and realistic textures on the glass edges and book cover.
  • + Clean, logical composition with a well-defined plant in the background.
  • The sphere appears to be floating without a clear point of contact or support inside the cube.

Qwen Image Max

  • + Features realistic light refractions and reflections within the glass panels.
  • + Includes a complex shadow pattern on the table denoting window light.
  • The sphere is unrealistically levitating in the center of the cube.
  • The book has strange black graphical artifacts/marks on the cover not mentioned in the prompt.

Verdict: GPT Image 1 is the superior image as it renders a much cleaner and more professional-looking scene. While both models struggled with gravity by making the sphere float, GPT Image 1 has better material textures and follows the 'red book' instruction without adding the unnecessary black markings seen in Qwen Image Max.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 1
Qwen Image Max

AI Judge Analysis

GPT Image 1

  • + Exquisite engraving detail on the plate armor with natural-looking wear.
  • + The soft lighting and skin texture appear highly photorealistic.
  • + Subtle, realistic interpretation of battle-worn features like dirt and faint scars.
  • The braid beads are very simple and less prominent than requested.
  • Leather and cloth details are mostly obscured by the armor design.

Qwen Image Max

  • + Excellent adherence to all prompt elements, including prominent beads and leather straps.
  • + Strong emotional expression and clear story-telling through the 'battle-worn' character design.
  • + Includes the actual torch source to justify the warm lighting and bokeh.
  • Some skin textures and scars look slightly over-sharpened or artificial.
  • The bokeh sparks are a bit uniform and busy, distracting from the face.

Verdict: Both models performed exceptionally well, but Qwen Image Max is the winner for its superior prompt adherence, successfully incorporating the leather straps, cloth underlayers, and specific braided beads that GPT Image 1 mostly overlooked. While GPT Image 1 has a more naturalistic, cinematic lighting quality, Qwen Image Max creates a more complete and detailed visual narrative of a paladin.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 1
Qwen Image Max

AI Judge Analysis

GPT Image 1

  • + Excellent typography with clean, consistent glowing effects
  • + Perfectly centered and organized 'exploded' layout
  • + High photorealistic texture on the patty and vegetables
  • Missed the '6' in the price starburst, displaying '.99' instead
  • The background is somewhat static compared to the requested motion

Qwen Image Max

  • + Successfully included all requested text and the correct price
  • + Dynamic sense of motion with flying embers and smoke effects
  • + Vibrant, high-contrast colors that fit the 'fiery' theme
  • The 'exploded' view is less clear, as components overlap heavily
  • The 'starburst' for the price is a bit messy and over-saturated

Verdict: GPT Image 1 has superior typography and a cleaner exploded burger composition, but it fails the prompt by missing a digit in the price. Qwen Image Max follows all text instructions perfectly and captures a stronger sense of explosive motion, making it the more effective advertisement overall despite a slightly more cluttered layout.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 1
Qwen Image Max

AI Judge Analysis

GPT Image 1

  • + Excellent fur texture and lighting integration on the capybara.
  • + The capybara's expression and posture perfectly match the 'calm, professional' prompt.
  • + High cinematic quality with realistic depth of field and bokeh windows.
  • The paws on the steering wheel look slightly anthropomorphized and thick.
  • The lady in the back is a bit out of focus compared to the rest of the scene.

Qwen Image Max

  • + Detailed taxi interior including dashboard, sun visor, and seatbelt.
  • + Fuller view of the Manhattan street scene through the front windshield.
  • + Capture of the passenger's bored expression is very accurate to the prompt.
  • The capybara's hands look more like human/monkey hands with fur rather than capybara paws.
  • The capybara's head has a strange anatomical transition where it meets the neck.
  • The textures on the interior ceiling look cracked and damaged, which wasn't requested.

Verdict: GPT Image 1 is the winner due to its superior photographic realism and the character design of the capybara, which looks much more natural and expressive than in the competing image. While Qwen Image Max does a great job with the taxi interior details, the anatomical issues with the driver's hands and neck make it less convincing overall.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 1
Qwen Image Max

AI Judge Analysis

GPT Image 1

  • + Excellent atmospheric lighting and moody color palette.
  • + Perfectly rendered gothic title font.
  • + Precise execution of the requested dark parchment texture.
  • Confused the event details, putting the location where the time should be.
  • The background elements like trees and moon are very faint and lack detail.

Qwen Image Max

  • + Highly accurate adherence to all text instructions, including date, time, and location.
  • + Richly detailed Illustration with prominent thorns and spiderwebs in the border.
  • + Clearly defined central glowing jack-o-lantern and twisted trees.
  • The parchment texture is a bit clean, leaning more toward a modern digital illustration than 'vintage'.
  • The scroll banner is slightly warped at the ends.

Verdict: While GPT Image 1 captures a more authentic vintage mood and superior cinematic lighting, it fails on the specific text requirements by omitting the time and mislabeling the location. Qwen Image Max successfully includes all requested text elements and decorative features like the thorns and webs, making it the more functional invitation despite a slightly less 'aged' appearance.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 1
Qwen Image Max

AI Judge Analysis

GPT Image 1

  • + Perfectly centered text layout according to the prompt hierarchy.
  • + Excellent clay-like 'cartoon' textured finish that looks high-quality.
  • + Very clean, minimal garnish and aesthetic balance.
  • The 'JAPAN' text feels a bit flat compared to the 3D scene.
  • The rice grains look like uniform bumps rather than realistic grains.

Qwen Image Max

  • + Impressive realistic PBR materials, especially the glass base and fish shine.
  • + Diverse selection of sushi types with high detail.
  • + Text has a nice 3D pop with drop shadows.
  • The flag icon is placed to the side rather than top-center as requested.
  • The 'JAPAN' text is slightly clipped at the top edge of the frame.
  • The garnish (spring onions) looks a bit cluttered compared to the 'minimal' request.

Verdict: GPT Image 1 followed the layout instructions much better, keeping the text and flag perfectly centered and within the frame, whereas Qwen Image Max clipped the top text and misplaced the flag. Qwen Image Max had superior material rendering, particularly the glass base and the realistic look of the fish, but GPT Image 1's cleaner composition and stylized cartoon aesthetic fit the 'miniature' prompt more cohesively.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 1
Qwen Image Max

AI Judge Analysis

GPT Image 1

  • + Excellent adherence to the list of animals including the rabbit.
  • + Very persuasive lighting with soft god rays and realistic bokeh.
  • + Anatomy is relatively consistent for all featured animals.
  • The fox has slightly blurred, dark paws that look a bit like blobs.
  • The golden retriever's back leg/tail area is a bit ambiguous in the background.

Qwen Image Max

  • + Successfully captures the 'tumbling' and playful interaction between animals.
  • + Lush, colorful wildflower meadow with high detail in the flora.
  • + Vibrant butterfly colors.
  • Missing the baby bunny entirely.
  • Anatomical issues including two separate golden retriever-like dogs (or one with a detached body) and a fox with an extra-long, awkwardly placed limb.
  • The fox's face looks slightly more like an adult than a kit.

Verdict: GPT Image 1 followed the prompt perfectly, including all four requested animals and maintaining a high level of photorealism. Qwen Image Max failed to include the baby bunny and suffers from significant anatomical errors where the animals' bodies overlap and blend together incorrectly. GPT Image 1's lighting and composition feel much more cohesive and professional.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 1
Qwen Image Max
100% wins 0% ties 0% wins

AI Judge Analysis

GPT Image 1

  • + Clean vector aesthetic with perfect spelling.
  • + Maintains high contrast suitable for a real logo emblem.
  • Ignores the request for a light background, opting for black.
  • The steam element is overly simplistic and looks a bit isolated.

Qwen Image Max

  • + Perfect adherence to the light background and subtle texture requirements.
  • + Excellent classic typography and detailed steam illustration.
  • + Accurate color palette using the requested warm brown and cream tones.
  • Slightly less 'minimalist' than Model A due to the shading lines.

Verdict: Qwen Image Max followed every specific instruction in the prompt, including the light background and parchment texture which GPT Image 1 ignored. While both models handled the text and main cloche elements well, Qwen Image Max's composition feels more like a complete vintage logo design.

Next steps

Explore each model

The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.