OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#29 of 62 in Text-to-Image
Qwen Image 2.0
#34 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
0%
win rate
Ties
0%
Qwen Image 2.0
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1
- + Excellent photographic quality with realistic textures on the book and sphere.
- + Accurately represents the blue sphere resting naturally on the bottom of the cube.
- + High level of visual clarity and clean composition.
- − The glass cube looks more like a frame with thick seams rather than a solid glass object.
Qwen Image 2.0
- + Captures the 'glass' material well with realistic reflections and refractions.
- + Accurate lighting direction from the window on the left.
- − The blue sphere is levitating inside the cube, which feels physically unnatural.
- − Confusing reflections make it look like there are multiple spheres inside the cube.
Verdict: GPT Image 1 is the superior choice because it presents a high-quality, physically grounded scene where the sphere sits naturally on the base of the cube. Qwen Image 2.0 struggles with spatial logic, showing the sphere floating in the air and creating confusing artifacts in the glass reflections.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1
- + Excellent shallow depth of field and bokeh realism
- + Highly realistic skin textures and fine details on the face
- + Captures a very cinematic mood and lighting
- − Physical logic of the bicycle frame and chain is somewhat distorted
- − Lacks the motion blur of passing cars requested in the prompt
Qwen Image 2.0
- + Naturalistic framing that feels more like a candid street photo
- + Better integration of the rain effect and wet pavement reflections
- + More realistic bicycle geometry compared to Image A
- − Misses the shallow depth of field and 50mm lens look, appearing too sharp in the background
- − Fails to include motion blur on the passing car
- − Contains a distracting cropped person on the right edge
Verdict: GPT Image 1 excels in visual quality and cinematic appeal, providing incredibly realistic skin textures and a beautiful shallow depth of field that matches the 50mm lens requirement. While Qwen Image 2.0 captures a more grounded, candid street composition and more accurate bicycle details, it fails to deliver the specific photographic depth of field and motion blur requested. GPT Image 1 is the winner for its superior technical execution and atmosphere, despite minor anatomical issues with the bike.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1
- + Exceptional photographic realism on the skin textures and eyes.
- + Highly intricate engraving detail on the plate armor that matches the 'worn' aesthetic.
- + Perfect execution of the shallow depth of field and warm lighting.
- − The 'small beads' in the hair are a bit subtle and few compared to the prompt.
Qwen Image 2.0
- + Excellent depiction of multiple braids with colorful beads.
- + Good visible layering of cloth, leather, and metal.
- + Clear portrayal of 'battle-worn' through distinct facial scars.
- − Anatomical issues with the hand and fingers on the sword hilt.
- − Lighting is a bit harsh and lacks the subtler 'warm torchlight' atmosphere of Image A.
- − The background fire is less 'bokeh' and more distracting than the prompt suggests.
Verdict: GPT Image 1 is the superior image due to its incredible photographic fidelity and sophisticated lighting. While Qwen Image 2.0 followed the braiding/beading instructions more literally, it suffered from noticeable anatomical defects in the hand and a less cohesive, more artificial lighting style.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1
- + Excellent font legibility with clean sans-serif typography.
- + Realistic, high-quality food photography with vibrant colors.
- + Professional layout that balances text and imagery effectively.
- − Minor spelling errors in placeholder text like 'descrigion'.
- − Includes a 'Main' item under the Pizza column, slightly breaking the grid logic.
Qwen Image 2.0
- + Stronger adherence to the 'grid' request with a 3x3 layout.
- + Consistent rounded-corner photo styling fits the modern minimalist theme.
- − Text is largely unintelligible gibberish.
- − Pricing and text alignment are messy and overlap some images.
- − Lower quality food renderings with some AI artifacts in the pizza toppings.
Verdict: GPT Image 1 is the superior choice because it produces a functional, professional-looking menu with legible text and high-quality photography. While Qwen Image 2.0 followed the grid layout request more strictly, its output is marred by gibberish text and poor typographic execution.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with a consistent glowing-ember style across all text elements.
- + High-quality, photorealistic textures on the burger bun and patty.
- + Perfect composition with a centered exploded view that feels balanced and professional.
- − The price in the starburst is missing the '6', rendering as '€.99'.
- − The background is slightly less dynamic than requested, focusing more on sparks than fiery atmosphere.
Qwen Image 2.0
- + Dynamic use of steam and falling droplets to create a strong sense of motion.
- + Accurate rendering of the price '€6.99' in a sharp, vibrant starburst.
- + Intense fiery background that matches the 'glancing embers' requirement perfectly.
- − The 'exploded' effect is less pronounced, with several ingredients still clumped together.
- − The 'LIMITED TIME ONLY' text is small and lacks the 'fiery, glowing effect' specified.
Verdict: GPT Image 1 offers a much cleaner and more professional layout with superior food photography aesthetics, though it fails on the specific price numerals. Qwen Image 2.0 captures the motion and fiery atmosphere more effectively, but the typography is less integrated and the 'exploded' burger effect is not as clearly executed as in the first image. GPT Image 1 is preferred for its overall polish and adherence to the visual style of the text elements.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Excellent text legibility and spelling accuracy
- + Consistent chalk texture across all letters
- + Followed the specific pricing and menu items accurately
- − Failed to provide the requested 'elegant cursive' for the title
- − The handwriting looks somewhat digital and uniform rather than having natural variations
- − Missed the 'cozy café' background, showing only the board
Qwen Image 2.0
- + Successfully rendered the title in a cursive-influenced style
- + Beautiful 'cozy café' background with realistic lighting and smudged chalk effects
- + Handwriting has very natural variations in stroke weight and slant
- − Minor layout issues with the pricing on the third item
- − Does not follow the line-by-line formatting as strictly as the other model
Verdict: GPT (Image 1) provides much cleaner and more accurate text rendering with perfect spelling, but it feels more like a digital font and lacks the requested cafe environment. Qwen (Image 2.0) better captures the artistic requirements of the prompt, including the elegant cursive title, the atmospheric background, and a more authentic chalk-on-blackboard aesthetic, making it the more visually convincing result despite minor alignment issues.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1
- + High cinematic visual quality with realistic lighting and textures.
- + Clear, detailed composition and excellent rendering of the astronaut suit.
- − Failed the primary prompt constraint of having the horse on top of the astronaut.
Qwen Image 2.0
- + Creative interpretation with scale-like textures on the horse skin.
- + Vibrant colors and interesting floating water droplets add to the surrealism.
- − Failed the primary prompt constraint of having the horse on top of the astronaut.
- − Anatomical issues with the horse's front legs and hoof alignment.
Verdict: Both GPT Image 1 and Qwen Image 2.0 failed to follow the specific instruction to place the horse on top of the astronaut, instead defaulting to the standard image of an astronaut riding a horse. GPT Image 1 is technically superior in terms of lighting, texture, and anatomical realism, while Qwen Image 2.0 offers a more surreal artistic style but suffers from minor anatomical artifacts.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1
- + Excellent photorealistic lighting and depth of field
- + Clear, high-quality text on the taxi cap
- + The businesswoman’s expression and pose perfectly match the 'bored' requirement
- − The capybara's paws are somewhat indistinctly merged with the steering wheel
Qwen Image 2.0
- + Dynamic composition and vibrant colors
- + Captures both characters clearly within the frame
- + Detailed texture on the capybara's fur and jacket
- − The passenger is sitting in the front seat instead of the back seat as requested
- − Minor anatomical artifacts on the capybara's paws
- − The passenger's hair and the taxi exterior have slightly surreal lighting
Verdict: GPT Image 1 followed the spatial instructions much better, correctly placing the human passenger in the back seat and capturing the requested 'bored' expression perfectly. Qwen Image 2.0 produced a vibrant and sharp image, but failed the prompt by placing the passenger in the front seat and missing the specific taxi text on the hat.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the 'dark parchment' and 'cinematic lighting' requirements.
- + Superior typography choice that feels authentically gothic and well-integrated into the art.
- + Better implementation of the spooky border with subtle webs and thorns.
- − Merged the Time and Location text incorrectly, omitting '7pm' and listing the location next to the 'Time' header.
- − The jack-o-lantern's glow is less vibrant than the competitor's.
Qwen Image 2.0
- + Perfect text accuracy for all event details including date, time, and location.
- + High visual clarity and sharp details on the trees and jack-o-lantern.
- + Dynamic composition with the scroll banner and twisted trees framing the pumpkin.
- − The parchment is a bit too bright, losing some of the 'dark parchment' moody aesthetic requested.
- − The typography for the bottom details is a standard serif font which clashes slightly with the gothic theme.
Verdict: GPT Image 1 captures the 'vintage gothic' mood and aesthetic perfectly with superior atmospheric lighting and integrated font styles, though it fails on the specific text data at the bottom. Qwen Image 2.0 provides a much more functional invitation with perfect text accuracy and clear, high-quality illustrations, but it feels slightly less 'cinematic' and moody. Qwen Image 2.0 is the overall winner for successfully including all requested information while maintaining a high standard of visual design.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent 3D cartoon miniature aesthetic with soft, clay-like textures.
- + Perfect alignment of requested text and icon elements.
- + Clean, professional lighting and shadows on a consistent diorama base.
- − The rice texture is slightly simplified into spheres rather than realistic grains.
Qwen Image 2.0
- + Good realistic textures on the fish and wood grain of the diorama base.
- + Varied selection of sushi types including nigiri, maki, and unagi.
- − Missed the 'cartoon' aspect of the prompt, opting for a photorealistic style instead.
- − The text is less integrated into the 3D scene compared to the other model.
- − The 45-degree isometric angle is slightly shallow.
Verdict: GPT Image 1 followed the aesthetic requirements of the prompt much better, delivering a beautiful 3D cartoon diorama with perfectly integrated text. Qwen Image 2.0 provided a more realistic interpretation that ignored the 'cartoon' and 'soft texture' instructions, resulting in a standard product shot look.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Excellent depiction of dynamic movement through the animals' poses
- + Highly effective use of god rays and warm sunrise lighting
- + Very clean anatomical details for all four animals
- − The fox looks a bit too similar in facial structure to the puppy
- − Bokeh is quite heavy, losing some of the 'lush meadow' detail in the background
Qwen Image 2.0
- + Perfectly captures the 'tumbling together' part of the prompt with physical interaction
- + Great variety of colorful wildflowers enhancing the meadow scene
- + Sharp focus on the fur textures and butterfly details
- − The fox's eyes and face look a bit distorted while it is on its back
- − Shadows on the kitten's face are a bit harsh compared to the 'warm sunrise' request
Verdict: Both models followed the prompt exceptionally well, including all four requested animals. GPT Image 1 (Model A) creates a more ethereal and cinematic atmosphere with its lighting, while Qwen Image 2.0 (Model B) better captures the specific 'tumbling' interaction and provides a richer, more detailed floral environment. GPT Image 1 is slightly more polished in terms of animal anatomy, making it the preferred choice for a masterpiece-quality image.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1
- + Strong minimalist vector aesthetic that perfectly matches the 'emblem style' instruction.
- + Excellent typography and formatting of the banner.
- + Accurate text rendering including the 'è' accent.
- − Ignored the 'light background' instruction, providing a black background instead.
- − The steam element is very simple and somewhat disconnected from the cloche.
Qwen Image 2.0
- + Successfully applied the warm brown and cream tones on a light, textured background.
- + Creative integration of the steam element inside the cloche.
- + Good aesthetic appeal with a clear vintage feel.
- − Included a strange scroll-like artifact or tail on the banner that breaks symmetry.
- − Typography is a bit generic compared to the requested 'classic' style.
- − The cloche dome rendering is less 'minimalist vector' and more illustrative/shaded.
Verdict: GPT Image 1 followed the technical vector style and typography requirements more precisely, but failed to provide the light background requested. Qwen Image 2.0 captured the color palette and background texture much better but suffered from a messy banner design and a less professional logo layout. GPT Image 1 is the likely winner for its superior font handling and cleaner geometric design, which fits a professional 'logo' brief better.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with very clean, modern font choices
- + Strong vector aesthetics that perfectly match the flat-design request
- + Creative addition of astronaut silhouettes that enhance the infographic feel
- − Several spelling errors including 'EARLLUNAR' and missing icons for specific steps
- − Layout is a bit cluttered and lacks a clear chronological flow
Qwen Image 2.0
- + Successfully includes all 6 requested steps in a clear vertical hierarchy
- + Excellent adherence to the NASA-inspired color palette
- + Accurate technical drawings of the lunar module for the descent and landing stages
- − Spelling error in 'Translunjar'
- − Some text elements overlap with icons, slightly reducing legibility
Verdict: Qwen Image 2.0 is the winner because it successfully captured all six requested mission steps in a logical, vertical flow that represents the actual progression of the flight. While GPT Image 1 has a more 'professional' graphic design aesthetic and better font rendering, it failed the core instruction of displaying the six sequential steps and included several confusing spelling and iconography errors.
Explore each model
Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request