Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 2 OpenAI Qwen Image Max Alibaba

Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.

GPT Image 2

27.7 arena score

#4 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image Max

21.3 arena score

#35 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 2

0%

win rate

Ties

0%

Qwen Image Max

0%

win rate

Shared challenges 8

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 2
Qwen Image Max

AI Judge Analysis

GPT Image 2

  • + Excellent adherence to lighting instructions with a clear source from the left.
  • + Realistic rendering of the glass edges and reflections.
  • + High-quality textures on the book and table surface.
  • The sphere is resting on the bottom rather than being suspended, which is a neutral but safe choice.

Qwen Image Max

  • + Creative interpretation with a levitating sphere.
  • + Strong visual balance and clear prompt adherence.
  • The glass refractive logic is slightly inconsistent on the left side of the cube.
  • The book has black markings that were not requested in the prompt.

Verdict: GPT Image 2 is the preferred output due to its superior realism and lighting consistency. While Qwen Image Max offers a more creative floating sphere, it suffers from minor optical distortions in the glass and added details on the book that weren't requested.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 2
Qwen Image Max

AI Judge Analysis

GPT Image 2

  • + Excellent photorealistic skin texture and eye detail
  • + Subtle and realistic integration of braided hair with small beads
  • + Beautifully rendered engraving on the plate armor with natural-looking wear
  • The warmth of the torchlight is slightly muted compared to the prompt's potential
  • Lower number of visible sparks for the requested bokeh effect

Qwen Image Max

  • + Strong prompt adherence regarding the bokeh sparks and warm torchlight
  • + Distinct and detailed leather straps and cloth underlayers
  • + Character captures a very intense 'battle-worn' expression with deep scars
  • Slightly more 'digital' feel to the skin texture compared to Model A
  • The torch in the background is a bit distracting in its proximity to the head

Verdict: GPT Image 2 provides a more nuanced and photorealistic portrait with subtle textures and high-quality armor engraving. Qwen Image Max follows the atmospheric prompts like bokeh sparks and cloth details more literally, but the overall composition and realism in GPT Image 2 are superior for a character study.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 2
Qwen Image Max

AI Judge Analysis

GPT Image 2

  • + Excellent exploded view showing all individual ingredients suspended in mid-air
  • + Perfect text rendering with high adherence to the fiery/glowing requested style
  • + High photorealistic detail on the patty texture and sauce droplets
  • Composition is slightly crowded on the left side with the text overlays

Qwen Image Max

  • + Strong composition with a centralized focus and good depth of field
  • + Clean and legible text placement
  • + Good motion blur effects on the flying embers
  • Failed the 'exploded burger' requirement, as the buns and patty are mostly touching
  • Text missing the requested fiery glow effect on the price starburst
  • Some ingredients (like the tomato slice) look less realistic and more like 3D renders

Verdict: GPT Image 2 followed the prompt much more accurately, particularly regarding the 'exploded' nature of the burger and the specific fiery styling of the various text elements. Qwen Image Max produced a high-quality advertisement, but it failed to separate the burger components and used standard gradients for some of the text instead of the requested glowing fire effect.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 2
Qwen Image Max

AI Judge Analysis

GPT Image 2

  • + Follows the hat description perfectly with a formal taxi driver cap featuring an 'NYC' logo.
  • + Excellent cinematic lighting and realistic bokeh in the background streets.
  • + The capybara's professional and calm expression is very well executed.
  • The capybara's front paws look slightly more like gloved human hands than realistic paws.
  • The passenger in the back is quite blurry due to the shallow depth of field.

Qwen Image Max

  • + Shows a more comprehensive view of the taxi interior and the dashboard.
  • + The passenger's 'bored' expression is very clearly rendered and matches the prompt well.
  • + Good texture detail on the capybara's fur and the jacket fabric.
  • The hat is a standard baseball cap rather than a 'taxi driver cap'.
  • The capybara has distinct primate-like fingers/hands rather than paws.
  • The roof of the car has strange, distracting cracks that look more like a digital artifact than realistic wear.

Verdict: GPT Image 2 is the preferred output because it captures the specific aesthetic of a New York taxi driver more accurately, specifically regarding the iconic cap. While Qwen Image Max provides a wider interior view, its transformation of the capybara's paws into human-like fingers is unsettling and less faithful to the animal's anatomy than the approach in GPT Image 2.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 2
Qwen Image Max

AI Judge Analysis

GPT Image 2

  • + Excellent typography style that perfectly fits the gothic aesthetic.
  • + Superior texture work on the dark parchment and intricate border.
  • + Cleverly includes a subtle NYC skyline in the background to match the location.
  • The parchment is very dark, which may slightly reduce readability for the bottom text.

Qwen Image Max

  • + High contrast between the jack-o-lantern and the background creates a strong focal point.
  • + Clean and legible layout for all text elements.
  • + Includes all requested elements like the scroll banner and twisted trees clearly.
  • The thorns in the border look a bit more like generic barbed wire than organic gothic brambles.
  • The overall composition feels a bit more like a modern digital illustration than a 'vintage' poster.

Verdict: GPT Image 2 is the winner because it captures the 'vintage gothic' aesthetic much more effectively with its weathered textures, intricate filigree border, and sophisticated typography. While Qwen Image Max is clean and literal, GPT Image 2 adds creative atmospheric touches like the NYC skyline and lanterns that make the invitation feel premium and immersive.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 2
Qwen Image Max

AI Judge Analysis

GPT Image 2

  • + Excellent text rendering with clean outlines
  • + Rich detail in the sushi textures especially the fish and ikura
  • + Superior diorama base design with miniature lantern and realistic stone.
  • The scene is slightly less minimalist than requested due to the extra environmental props.

Qwen Image Max

  • + Clean aesthetic with a very clear, centered composition
  • + Accurately represents the 45-degree isometric top-down angle
  • + Faithful adherence to the minimal diorama base requirement.
  • The flag icon is placed to the side of the text rather than below it
  • Slightly less realistic PBR textures on the sushi ingredients compared to the competitor.

Verdict: GPT Image 2 provides a more detailed and visually appealing diorama with professional-grade PBR textures and perfect text placement. While Qwen Image Max captures the minimalist requirement well, its flag placement is incorrect and the overall lighting and texture quality lack the polish seen in GPT Image 2.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 2
Qwen Image Max

AI Judge Analysis

GPT Image 2

  • + Perfect adherence to the animal list (puppy, kitten, bunny, fox).
  • + Beautiful rim lighting and realistic god rays that enhance the morning atmosphere.
  • + Dynamic composition with a great sense of motion and 'tumbling' as requested.
  • The fox kit's anatomy is slightly elongated in the hind area.

Qwen Image Max

  • + High level of fur detail and clear textures.
  • + Strong adherence to the 'wholesome tumbling' vibe.
  • Failed to include the 'baby bunny' requested in the prompt.
  • Included two golden retrievers instead of just one.
  • Anatomical issues with the fox's paws touching the puppy's head.

Verdict: GPT Image 2 is the clear winner as it correctly rendered all four specific animal types requested, whereas Qwen Image Max missed the bunny entirely and duplicated the dog. GPT Image 2 also features superior lighting and a more natural integration of the animals within the meadow environment.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 2
Qwen Image Max

AI Judge Analysis

GPT Image 2

  • + Excellent typography with clean, professionally spaced letters
  • + Successful inclusion of all requested elements including the banner and steam
  • + High-quality vector engraving style with a sophisticated frame
  • The complexity of the frame might lean more towards 'ornate' than 'minimalist'

Qwen Image Max

  • + Strong iconography with the cloche dome featured prominently
  • + Good adherence to the requested color palette
  • + Accurate text rendering for both name and date
  • The paper texture looks more like a stained parchment than a subtle professional background
  • The layout feels slightly bottom-heavy compared to Model A
  • The 'Est. 1720' is in a simple box rather than a classic flowing banner

Verdict: GPT Image 2 (Model A) is the superior logo as it demonstrates a much more refined grasp of classic typography and emblem composition. While Qwen Image Max (Model B) followed all instructions, its execution feels like a modern illustration on a faux-old background, whereas Model A feels like a genuine, high-end vintage brand identity.

Next steps

Explore each model

The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.