OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 9 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#28 of 62 in Text-to-Image
Qwen Image 2.0 Pro
#32 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
0%
win rate
Ties
0%
Qwen Image 2.0 Pro
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1
- + Excellent depiction of the plant's distortion and visibility through the glass.
- + The glass cube has realistic thickness and greenish edge tinting.
- + Accurate soft window lighting consistent with the left-side source.
- − The sphere is resting on the bottom rather than being centered, though this wasn't explicitly specified in the prompt.
Qwen Image 2.0 Pro
- + Successfully placed the sphere inside the cube as requested.
- + The wooden table has a rich, realistic texture.
- − The glass rendering is inconsistent, with the front face appearing missing or having impossible reflections.
- − The sphere is levitating mysteriously in the center, which looks unnatural.
- − The plant visibility through the glass is less clear than in Model A.
Verdict: GPT Image 1 is the superior image because it renders the physics and optical properties of glass much more realistically, specifically how the plant and light interact with the cube's surfaces. Qwen Image 2.0 Pro struggles with the structural integrity of the cube, making the front face appear non-existent and having the sphere levitate without context.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Excellent photorealistic texture on the bun and patty.
- + Highly consistent typography with the requested fiery, glowing effect.
- + Better vertical composition and spacing of the exploding ingredients.
- − Missed the '6' in the price tag, rendering it as '€.99'.
- − Less 'fire' in the background compared to model B, focusing more on embers.
Qwen Image 2.0 Pro
- + Accurately rendered all text including the correct price '€6.99'.
- + Dynamic smoke and flame effects enhance the 'fiery' atmosphere.
- + Great textural detail on the seared patty and fresh lettuce.
- − The starburst price tag looks like a flat 2D sticker rather than being integrated into the 3D scene.
- − The 'MAGIC BURGER' text has slight irregularities in the flame mask.
Verdict: Both models followed the prompt well, but Qwen Image 2.0 Pro is the winner for its superior text accuracy, correctly rendering the price while GPT Image 1 missed a digit. While GPT Image 1 has more cohesive glowing typography, Qwen Image 2.0 Pro captures a better sense of heat and motion with smoke and realistic fire effects.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Excellent chalk texture on the individual letters.
- + Perfect spelling and full completion of the prompt's text including the bottom disclaimer.
- − The handwriting style is too uniform, looking more like a digital chalk font than natural handwriting.
- − Failed to use 'elegant cursive' for the title as requested.
Qwen Image 2.0 Pro
- + Captures a much more realistic and 'cozy café' atmosphere with the background elements.
- + Handwriting has natural variations in slant, size, and weight that feel authentic.
- + Includes specific requested chalk smudges and imperfect board texture.
- − Failed to provide the 'elegant cursive' title requested in the prompt.
- − The perspective makes the text slightly harder to read compared to a flat shot.
Verdict: GPT Image 1 followed the text instructions more literally by including every word requested, but the output feels like a digital graphic rather than a real chalkboard. Qwen Image 2.0 Pro provided a much more believable and artistic interpretation of a café environment with authentic-looking handwriting, though it missed the cursive requirement for the title. Qwen is the preferred choice for its superior realism and natural variation in chalk strokes.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1
- + Excellent fur texture rendering on the capybara.
- + Sophisticated cinematic lighting that feels very 'New York at night'.
- + Bored expression of the passenger perfectly captures the requested irony.
- − The capybara's paws look somewhat glove-like or humanized rather than anatomically correct paws.
- − The 'TAXI' text on the hat is a bit generic.
Qwen Image 2.0 Pro
- + Stronger sense of realism in the taxi interior details, such as the TLC license and dashboard lights.
- + The capybara's paws are more realistically integrated onto the steering wheel.
- + The passenger is well-framed and clearly engaged with her phone.
- − The hat badge and text on the license ('Licesed') contain spelling and rendering errors.
- − The composition feels slightly more cluttered compared to the cinematic focus of the other model.
Verdict: GPT Image 1 offers a more artistically cohesive and cinematic shot with superior lighting and texture, making the surreal prompt feel like a still from a high-end film. Qwen Image 2.0 Pro provides more environmental detail (like the TLC plate) and more realistic paw placement, but suffers from minor text artifacts and a less polished aesthetic.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent typography style that matches the gothic aesthetic perfectly.
- + Sophisticated, moody lighting with a cinematic glow from the jack-o-lantern.
- + The composition feels cohesive and professionally designed.
- − The 'Time' and 'Location' lines of text are merged together incorrectly.
- − The border elements are very subtle, almost fading into the dark background.
Qwen Image 2.0 Pro
- + Successfully included and separated all text lines, including Date, Time, and Location.
- + Strong adherence to the border requirement with highly visible thorns and webs.
- + Crisp details on the bats and central pumpkin character.
- − The green lighting feels a bit cartoonish compared to the 'vintage gothic' request.
- − The banner text uses a modern script that clashes slightly with the gothic header.
Verdict: Qwen Image 2.0 Pro is the overall winner because it successfully rendered all specific event details (Date, Time, and Location) as separate lines, whereas GPT Image 1 merged the time and location. While GPT Image 1 had a more authentic vintage gothic atmosphere, Qwen Image 2.0 Pro followed the complex layout and border instructions more accurately.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent soft 3D textures that accurately reflect the 3D cartoon request
- + Clean and stylistically integrated text rendering
- + Great use of lighting and shadows to create depth on the diorama base
- − The sushi rice looks a bit like bumpy bubbles rather than distinct grains
- − The flag icon is slightly simplified and lacks a border
Qwen Image 2.0 Pro
- + High clarity on individual rice grains
- + Perfect adherence to the 45° isometric angle
- + Includes more variety of sushi pieces as implied by the dish
- − Lighting is a bit flat compared to the other model
- − The text looks like a simple 2D overlay rather than being integrated into the 3D scene
- − The shiso leaf feels a bit lower quality than the rest of the elements
Verdict: GPT Image 1 captures the 'soft refined textures' and 'miniature 3D cartoon' aesthetic much more effectively with its clay-like rendering and gentle lighting. While Qwen Image 2.0 Pro offers better detail on the rice grains, its composition and text integration feel less cohesive as a stylized 3D scene. GPT Image 1 is the preferred choice for its superior artistic polish and execution of the requested PBR materials.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Excellent depiction of motion and playful 'tumbling' as requested.
- + Fantastic fur texture and high-quality lighting that creates a warm, atmospheric glow.
- + Unified expressions of joy across all four animals.
- − The fox has black paws that look slightly like blurred blobs.
- − The scale of the rabbit relative to the kitten is slightly off.
Qwen Image 2.0 Pro
- + Beautifully detailed wildflower meadow with variety in flora.
- + Excellent composition with clear 'god rays' and a defined background landscape.
- + Included all four animals with distinct, realistic features.
- − The kitten's pose is a bit stiff and static compared to the others.
- − The animals aren't really 'chasing' or 'tumbling' as much as they are just sitting together.
Verdict: GPT Image 1 captures the dynamic movement and playful spirit of the prompt much better, showing the animals in active 'tumbling' poses. Qwen Image 2.0 Pro has a superior background with more varied flowers and atmospheric depth, but the animals feel more like they are posing for a photo than playing. GPT Image 1 is the preferred winner for its better adherence to the action and cohesive emotional energy.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1
- + Perfect text rendering with correct spelling and accents
- + Clean minimalist vector style
- + Accurately follows the banner and steam request
- − Failed to provide a light background as requested in the prompt
- − The steam effect is very basic compared to the rest of the logo
Qwen Image 2.0 Pro
- + Accurately followed the light background and warm brown/cream tone request
- + Pleasant subtle paper texture
- + Elegant cloche design with well-rendered steam
- − Spelling error in the main text ('Floriian' instead of 'Florian')
- − Cluttered composition where the banner overlaps the cloche awkwardly
- − The date is floating rather than being inside the banner
Verdict: GPT Image 1 followed the technical typography requirements perfectly, including the specific accent in 'Caffè', although it completely ignored the request for a light background. Qwen Image 2.0 Pro captured the aesthetics and color palette of the prompt much better, but included a noticeable spelling error ('Floriian') and failed to place the date inside the banner as requested. GPT Image 1 is the likely winner for its professional-grade utility despite the background color error.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1
- + Excellent typography rendering with almost perfect spelling of names.
- + Strong adherence to the 'flat vector' style with clean, bold icons.
- + Uses the requested NASA-inspired muted color palette effectively.
- − The layout is a bit cluttered and the sequence of steps is difficult to follow.
- − Contains a spelling error in a key label ('EARLLUNAR').
- − The Saturn V rocket is cut off by the frame.
Qwen Image 2.0 Pro
- + Logical vertical layout that clearly illustrates the mission progression from top to bottom.
- + Perfect adherence to all 6 requested steps in the correct order.
- + Better composition as a 'poster' with a clear path and focal points.
- − The lunar module illustration is too complex and deviates from the requested 'flat vector' style.
- − The astronaut icons at the bottom are inconsistent (one is a helmet, two are silhouettes).
- − Text for the names is very small and difficult to read.
Verdict: Qwen Image 2.0 Pro is the winner because it successfully followed the structural 'Steps' requirement of the prompt, creating a coherent 1-6 sequence that tells the story of the mission. While GPT Image 1 has superior typography and a more consistent flat-vector aesthetic, its layout is chaotic and fails to present the mission phases in a logical order.
Explore each model
Alibaba's Qwen Image 2.0 Pro model offering higher quality image generation with enhanced detail and accuracy