Head to head
Esc

Models · slot A

to navigate to pick

GPT Image 1 Mini OpenAI Qwen Image 2.0 Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

GPT Image 1 Mini

25.0 arena score

#13 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image 2.0

21.7 arena score

#34 of 62 in Text-to-Image

Vote tally

Where the votes landed

GPT Image 1 Mini

0%

win rate

Ties

0%

Qwen Image 2.0

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

GPT Image 1 Mini
Qwen Image 2.0

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent photographic quality with realistic soft lighting.
  • + Accurate glass physics and simple, clean composition.
  • + Perfect adherence to the 'small blue sphere' description.
  • The plant is slightly out of focus and less visible through the glass than it could be.

Qwen Image 2.0

  • + Strong adherence to all prompt elements, including the window light source.
  • + Great plant visibility through the glass facets.
  • + High-quality texture on the red book cover.
  • The blue sphere appears floating and duplicated in reflections in an unrealistic way.
  • Perspective of the cube feels slightly distorted.

Verdict: GPT Image 1 Mini produced a more realistically grounded image with superior lighting and physics, correctly depicting a single small sphere. While Qwen Image 2.0 followed the prompt well, the floating sphere and confusing internal reflections make it feel less coherent as a photograph.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

GPT Image 1 Mini
Qwen Image 2.0

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent shallow depth of field and bokeh realism
  • + Sophisticated cinematic color grading and lighting
  • + Realistic skin textures and weather effects
  • Anatomical issues with the hands appearing merged or jumbled
  • Mechanical inaccuracies in how the bicycle chain and frame connect

Qwen Image 2.0

  • + Natural street photography composition and imperfect framing
  • + Accurate motion blur on passing vehicles
  • + Better physical interaction between the person and the bicycle mechanics
  • Slightly softer facial details compared to Model A
  • Rain effect is less visible than requested

Verdict: GPT Image 1 Mini produces a more aesthetically pleasing, cinematic image with superior textures, but it suffers from significant anatomical and structural glitches in the hands and bike. Qwen Image 2.0 captures the 'candid street photo' prompt more accurately with a realistic lens feel and better mechanical logic, despite having slightly less detail in the skin.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

GPT Image 1 Mini
Qwen Image 2.0

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent fine engraving detail on the armor plate.
  • + Superior lighting integration with warm, realistic torchlight tones.
  • + Highly detailed facial texture including pores and fine hair.
  • Missed the request for beads in the hair braids.
  • The background is quite dark, bordering on mudiness.

Qwen Image 2.0

  • + Followed the specific instruction for beads in the hair.
  • + Captures the 'battle-worn' scars and facial dirt more prominently.
  • + Includes the requested leather straps and cloth underlayer in clear detail.
  • The hand and sword hilt rendering is anatomically awkward and messy.
  • The fire effect in the background looks a bit flat and artificial compared to the subject.
  • The armor engraving is less ornate and simpler than Model A.

Verdict: GPT Image 1 Mini produces a much more cinematic and high-quality image with superior lighting and texture, though it ignores the request for beads. Qwen Image 2.0 is better at capturing all the specific prompt elements (beads, scars, cloth layers) but suffers from significant anatomical errors in the hand and less realistic lighting.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

GPT Image 1 Mini
Qwen Image 2.0

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent typography with perfectly legible text.
  • + True minimalist design with clean placeholders for menu items.
  • + Consistent, high-quality food photography in a structured grid.

Qwen Image 2.0

  • + High image count provides a fuller sense of a diverse menu.
  • + Features pricing elements alongside the items.
  • + Modern aesthetic with rounded corners and vibrant photography.
  • Nonsensical garbled text for food item descriptions.
  • Menu structure is confusing as 'Appetizers' and 'Mains' are at the top of columns but contain Pizza images.
  • Several pizza images are repeated or very similar.

Verdict: GPT Image 1 Mini followed the instructions for specific sections (Appetizers, Pizza, Mains) much better than Qwen Image 2.0, which placed pizza images under the wrong headers. While Qwen Image 2.0 has a more visually dense layout, its text is completely illegible and nonsensical compared to the clean, professional result from GPT Image 1 Mini.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

GPT Image 1 Mini
Qwen Image 2.0

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent typography rendering with consistent fiery glow effect.
  • + Clean, professional composition following all prompt layout requirements.
  • + Strong photorealistic textures on the bun and meat patty.
  • The 'exploded' effect is a bit static compared to the second image.
  • The starburst shape is slightly irregular on the right side.

Qwen Image 2.0

  • + Dynamic sense of motion with flying crumbs, smoke, and embers.
  • + Highly detailed ingredients with juicy, appetizing textures.
  • + Vibrant flame effects on the main title.
  • The secondary 'LIMITED TIME ONLY' text lacks the requested fiery/glowing effect.
  • The price starburst is cluttered and less legible than the first image.
  • The background is slightly more chaotic, detracting from the food.

Verdict: GPT Image 1 Mini followed the layout and text styling instructions perfectly, creating a clean and professional advertisement. While Qwen Image 2.0 offered more dynamic motion and slightly better food textures, it failed to apply the requested styling to all text elements. GPT Image 1 Mini is the winner for its superior prompt adherence and typography consistency.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

GPT Image 1 Mini
Qwen Image 2.0

AI Judge Analysis

GPT Image 1 Mini

  • + Perfect text accuracy for all requested menu items and numbers.
  • + Clean and balanced composition within the wooden frame.
  • The font looks too consistent and digital rather than truly handwritten.
  • Failed to follow the instruction for 'elegant cursive' for the title.

Qwen Image 2.0

  • + Excellent chalk texture throughout the board, including smudge marks.
  • + Stronger adherence to the 'handwritten' feel with varying slant and cursive elements in the title.
  • + Better environmental context for a 'cozy café'.
  • Minor duplicate dash after 'Chip Cookies'.
  • Total text character count slightly noisier than Model A.

Verdict: GPT Image 1 Mini followed the text instructions perfectly but produced a result that looked too much like a digital font overlay. Qwen Image 2.0 captured the requested aesthetic much better, providing realistic chalk textures, elegant cursive flourishes in the title, and a natural cafe atmosphere despite a minor punctuation error.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

GPT Image 1 Mini
Qwen Image 2.0

AI Judge Analysis

GPT Image 1 Mini

  • + High cinematic visual quality with moody lighting and excellent textures.
  • + Excellent composition with the moon adding depth to the scene.
  • Completely failed the negative constraint/spatial instruction of 'horse on top'.
  • The pose is a standard cliché, missing the requested surreal inversion.

Qwen Image 2.0

  • + Bright, high-contrast colors and sharp details on the suit and horse.
  • + Interesting scale-like texture on the horse's neck adds to the surreal theme.
  • Failed the specific spatial instruction of 'horse on top, not vice versa'.
  • Anatomical issues with the horse's legs, particularly the front right and rear right legs.

Verdict: Both GPT Image 1 Mini and Qwen Image 2.0 failed the specific spatial logic prompt to have the horse on top of the astronaut, both delivering the standard image of an astronaut riding a horse. GPT Image 1 Mini is the better image overall due to its cinematic lighting and superior anatomical correctness of the horse compared to the distorted limbs in the Qwen Image 2.0 output.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

GPT Image 1 Mini
Qwen Image 2.0

AI judge analysis unavailable for this challenge.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

GPT Image 1 Mini
Qwen Image 2.0

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent layout with centered, symmetrical typography and a strong central focal point.
  • + Very clean text rendering with a consistent vintage font for the details.
  • + Atmospheric lighting that blends the jack-o-lantern glow realistically with the dark environment.
  • The parchment texture is very dark, making the border details like thorns and webs harder to see.
  • The gothic title font is relatively simple compared to the 'elegant gothic' request.

Qwen Image 2.0

  • + Beautiful gothic calligraphy for the main title that perfectly matches the 'elegant' prompt requirement.
  • + Distinct parchment color that provides better contrast for the spooky trees and border elements.
  • + Includes a visible moon in the moody night sky, adding to the composition.
  • Minor text artifact in the banner with a comma instead of a space ('Your, are').
  • The jack-o-lantern feels slightly less integrated into the background compared to Model A.
  • The typography at the bottom is a bit plain compared to the top half of the poster.

Verdict: Both models followed the complex prompt very well, but GPT Image 1 Mini produced a more cohesive and professional-looking layout with perfect spelling. Qwen Image 2.0 had superior gothic typography and better contrast on the parchment, but was held back by a small punctuation error in the banner and a slightly less unified composition. GPT Image 1 Mini is the winner for its flawless text execution and atmospheric lighting.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

GPT Image 1 Mini
Qwen Image 2.0

AI Judge Analysis

GPT Image 1 Mini

  • + Perfectly captures the requested 3D cartoon style with soft, stylized textures.
  • + Excellent typography layout that feels integrated into the design.
  • + Clean isometric composition with high-clarity materials.
  • The sushi toppings look slightly plastic rather than 'realistic PBR' fish.

Qwen Image 2.0

  • + More realistic textures on the sushi fish and wooden base.
  • + Includes a wider variety of sushi pieces.
  • + Accurate text rendering and flag inclusion.
  • The mix of realistic food on a flat background creates a slight compositing clash.
  • Text feels like a simple overlay rather than part of a miniature scene.
  • The lighting is a bit harsh compared to the 'gentle lighting' request.

Verdict: GPT Image 1 Mini followed the stylistic instructions for a '3D cartoon scene' much better than Qwen Image 2.0, which opted for more photorealistic elements. While Qwen produced more realistic textures, GPT Image 1 Mini's composition, typography, and clean isometric execution feel more like a cohesive miniature diorama.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

GPT Image 1 Mini
Qwen Image 2.0

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent depiction of motion with all four animals 'running' or 'tumbling' as requested.
  • + Very soft, high-quality fur textures across all subjects.
  • + Beautifully implemented god rays and golden hour lighting.
  • The fox has black 'socks' that are slightly more characteristic of an adult than a kit.
  • The composition is a bit more 'posed' in a line rather than a chaotic tumble.

Qwen Image 2.0

  • + Successfully captures the 'tumbling together' aspect of the prompt with physical interaction between the animals.
  • + Rich variety and detail in the wildflower meadow.
  • + Strong adherence to the specific 'sun sparkles' and 'dew' mentioned in the prompt.
  • The fox kit's face looks slightly distorted and its eye color is an unnatural blue/white.
  • The butterfly on the right is disproportionately large and appears flat against the background.
  • The lighting on the animals feels a bit harsh and inconsistent with the background sun.

Verdict: GPT Image 1 Mini creates a more aesthetically pleasing and high-quality image with consistent lighting and superior fur rendering. While Qwen Image 2.0 did a better job of showing the animals actually tumbling and playing together, it suffered from minor anatomical distortions and less realistic butterfly integration.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

GPT Image 1 Mini
Qwen Image 2.0

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent typography with correct accent usage on 'Caffè'
  • + Clear and professional banner implementation for the establishment date
  • + Strong vector emblem feel suitable for branding
  • Failed to follow the background color instruction by using a black background
  • The cloche dome is stylized effectively but the steam is very minimal

Qwen Image 2.0

  • + Successfully used a light, textured background as requested
  • + Rendered a more detailed and visually interesting cloche dome
  • + Accurate text rendering for both the name and the establishment date
  • The placement of steam inside the cloche looks slightly like a flame
  • The typography is less 'classic' and feels a bit more modern compared to the requested vintage style

Verdict: Qwen Image 2.0 is the overall winner as it followed all prompt instructions, including the specific request for a light background which GPT Image 1 Mini ignored. While GPT Image 1 Mini produced a very professional logo with superior typography, Qwen Image 2.0 captured the 'vintage' aesthetic and color palette more accurately.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

GPT Image 1 Mini
Qwen Image 2.0

AI Judge Analysis

GPT Image 1 Mini

  • + Excellent typography with clean, legible sans-serif fonts.
  • + Consistent vector illustration style across all icons.
  • + Strictly follows the requested NASA-inspired color palette.
  • The 'Translunar' trajectory line is nonsensical and messy.
  • Layout is a bit crowded and cut off at the bottom.

Qwen Image 2.0

  • + Strong vertical composition that feels like a professional poster.
  • + Includes additional accurate details like the silhouette names of the crew.
  • + Effective use of the dark navy background to create atmospheric depth.
  • Spelling error in 'Translunjar'.
  • Inconsistent icon style, particularly the 'Descent' wireframe compared to the 'Landing' 3D model.

Verdict: Both models followed the complex prompt instructions well. GPT Image 1 Mini produced more consistent icons and better typography but struggled with the logic of the trajectory line. Qwen Image 2.0 created a more compelling poster layout and captured the 'NASA' feel better, though it suffered from a typo and architectural inconsistency in its spacecraft illustrations.

Next steps

Explore each model