OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 2
#4 of 62 in Text-to-Image
Qwen Image 2.0
#34 of 62 in Text-to-Image
Where the votes landed
GPT Image 2
80.0%
win rate
Ties
0.0%
Qwen Image 2.0
20.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 2
- + Excellent physical realism and materials, especially the leather texture of the book.
- + Flawless glass rendering with accurate refraction of the plant behind it.
- + Very high resolution and natural lighting that captures the soft window light perfectly.
- − The plant is slightly more to the side than directly 'behind' the cube.
Qwen Image 2.0
- + Correctly places the plant directly behind the cube as requested.
- + Meets all prompt requirements including the blue sphere and red book placement.
- − The blue sphere appears to be floating unnaturally in the center of the air.
- − The glass cube has internal vertical seams that make it look like multiple panes rather than a single cube.
- − Lower overall texture quality on the table and book compared to the competitor.
Verdict: GPT Image 2 is the superior generation due to its high level of photorealism and sophisticated handling of glass reflections and material textures. While Qwen Image 2.0 followed the spatial instruction of putting the plant behind the cube well, it failed on physics by showing a floating sphere and had awkward geometry in the glass construction.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 2
- + Excellent shallow depth of field and bokeh realism.
- + Superior skin texture and natural lighting on the subject.
- + Captures the motion blur from passing cars perfectly as requested.
- − The bike's proportions and rear wheel structure are slightly warped.
- − The toolbox in the foreground looks a bit cluttered.
Qwen Image 2.0
- + Strong reflection on the wet pavement.
- + Clearer focus on the mechanical action of repairing the pedal.
- + Good adherence to the red bicycle prompt.
- − The skin texture appears somewhat oversaturated and gritty rather than natural.
- − Lacks the motion blur from cars requested in the prompt.
- − The composition feels a bit more staged than personal/candid.
Verdict: GPT Image 2 is the clear winner as it successfully incorporated all technical aspects of the prompt, including the difficult motion blur and the specific 50mm lens look with shallow depth of field. While Qwen Image 2.0 captures a nice reflection and clear action, it fails to deliver the requested motion blur and has a less realistic skin rendering compared to GPT Image 2.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 2
- + Excellent skin texture with realistic dirt and freckles
- + Beautifully detailed engraving on the plate armor
- + Subtle and sophisticated lighting that feels like a real photograph
- − The hair beads are somewhat small and less distinct than in Model B
- − The close-crop composition cut off the top of the character's head
Qwen Image 2.0
- + Strong adherence to the 'beads in hair' and 'braids' prompt requirements
- + Effective use of bokeh sparks and warm lighting colors
- + Clearly visible scars as requested
- − The eyes appear unnatural with a glowing yellow effect not specified in the prompt
- − Noticeable anatomical issues with the hand on the sword hilt
- − Visual quality feels slightly more like a digital painting than a 'lifelike' photograph
Verdict: GPT Image 2 provides a much higher level of photographic realism and intricate texture on the armor, though it is a bit conservative with the 'beads' and 'sparks' elements. Qwen Image 2.0 interprets the prompt elements more literally but suffers from anatomical errors in the hand and less realistic lighting. GPT Image 2 is the better overall image due to its superior composition and detail fidelity.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 2
- + Excellent text rendering with clear, legible names and descriptions.
- + Perfect logical organization matching the requested categories of appetizers, pizza, and mains.
- + Professional graphic design elements including icons, pricing, and social media badges.
- − The layout is very dense, which slightly pushes the boundaries of 'minimalist'.
Qwen Image 2.0
- + Strong minimalist aesthetic with large, clean image tiles.
- + High-quality photographic realism for the food items.
- + Simple grid structure that is easy to scan.
- − Text is largely unintelligible gibberish.
- − The grid does not logically follow the requested categories (pizzas are mixed into other columns).
- − Missing detailed descriptions requested by the prompt.
Verdict: GPT Image 2 (Model A) is the clear winner as it produces a fully functional, professional-grade menu with legible text and a logical hierarchy. Qwen Image 2.0 (Model B) has high-quality visuals but fails significantly on text rendering and the logical organization of the requested sections.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 2
- + Excellent text integration with consistent fiery effects across all labels
- + Perfect 'exploded' view with clearly separated and recognizable components
- + Superior sense of motion through sauce splashes and embers
- − The composition is a bit crowded with large text elements
Qwen Image 2.0
- + Clean, professional font choice for the 'MAGIC BURGER' title
- + Good food texture and lighting on the bun and patty
- − Fails the 'exploded' prompt as most components are still stacked together
- − The '€6.99' text lacks the requested fiery/glowing effect, appearing metallic instead
- − Low contrast in the 'LIMITED TIME ONLY' text makes it hard to read
Verdict: GPT Image 2 followed the prompt precisely, delivering a dynamic exploded view where every ingredient is suspended in mid-air and all text follows the fiery styling. Qwen Image 2.0 failed to properly explode the burger and struggled with the specific styling requested for the price tag, resulting in a more static and less cohesive advertisement.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the prompt's text requirements with zero spelling errors.
- + Very realistic chalk texture with dusty, semi-transparent strokes that mimic real handwriting.
- + Consistent and elegant cursive style for the title as requested.
- − The spacing between lines is slightly tighter at the bottom, though this is common in real chalkboards.
Qwen Image 2.0
- + Dynamic lighting and realistic smudging on the chalkboard surface give it a lived-in feel.
- + Good handwriting style that varies naturally in size and slant.
- − The title lacks the specific 'elegant cursive' style requested, appearing more like standard print.
- − The text layout is slightly messy with unnecessary line breaks for the last two items.
Verdict: GPT Image 2 followed the prompt's text requirements perfectly, including the specific cursive title and exact pricing/wording. While Qwen Image 2.0 captured a beautiful aesthetic with realistic chalk smudges, it failed to provide the cursive title and had awkward line breaks. GPT Image 2 is the clear winner for its superior text rendering and prompt adherence.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the specific 'horse on top' prompt instruction
- + Realistic textures on the space suit and horse fur
- + Coherent surrealist composition with the horse holding the reins
- − The astronaut's hands/gloves are shaped more like feet than human hands
Qwen Image 2.0
- + High visual quality with vibrant colors and lighting
- + Dynamic sense of motion with the horse running through space
- + Good level of detail on the space suit and horse's mane
- − Failed the primary prompt instruction to have the horse on top of the astronaut
- − Included strange scaly skin artifacts on the horse's neck
Verdict: GPT Image 2 successfully followed the difficult logical constraint of the prompt ('horse on top, not vice versa'), creating a surreal and humorous image. Qwen Image 2.0 ignored the specific positioning instructions and produced a standard 'astronaut riding a horse' image. Despite some minor anatomical issues with the astronaut's hands, GPT Image 2 is the winner for its superior prompt adherence.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 2
- + Excellent cinematic lighting and textures that feel very photorealistic.
- + Accurate composition with the passenger clearly in the back seat as requested.
- + Detailed taxi driver cap with a logically placed logo.
- − The capybara's hands/paws on the steering wheel look slightly like human gloves.
- − The camera angle is slightly tight, obscuring more of the cab interior.
Qwen Image 2.0
- + High clarity and vibrant colors in the Manhattan background.
- + Good rendering of the capybara's fur and individual claws.
- + The passenger's expression perfectly matches the 'bored' requirement.
- − The passenger is sitting in the front passenger seat rather than the back seat.
- − The capybara's paw placement on the wheel is anatomically awkward.
- − The blue interior seats feel less like a standard New York taxi than the neutral tones in A.
Verdict: GPT Image 2 followed the spatial instructions more accurately by placing the businesswoman in the back seat, whereas Qwen Image 2.0 placed her in the front. GPT Image 2 also achieved a more convincing photorealistic look with cinematic lighting that fits a night-time New York setting.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 2
- + Excellent typography with a cohesive gothic font for all sections.
- + Stunning visual complexity with intricate borders, background silhouettes, and atmospheric lighting.
- + Perfect adherence to all prompt details, including the specific date and location.
- − The parchment texture is quite dark, making parts of the thorns less distinct.
Qwen Image 2.0
- + Clean, readable layout with a clear focal point on the jack-o-lantern.
- + Accurately represents all requested text and objects.
- − The background and trees look somewhat generic compared to the 'cinematic' request.
- − Text rendering on the scroll is slightly shaky and less integrated than Model A.
- − Lighting feels flat and lacks the 'moody' atmosphere requested.
Verdict: GPT Image 2 is the superior choice as it fully captures the 'cinematic' and 'vintage gothic' aesthetic with professional-grade typography and rich textures. While Qwen Image 2.0 followed the prompt instructions accurately, its execution is much simpler and looks more like a standard clip-art flyer than a polished invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 2
- + Perfectly adheres to the 3D cartoon style with high-quality PBR materials.
- + Flawless text rendering and icon placement that matches the requested layout.
- + Exceptional detail in the diorama base, including stone textures and miniature pagoda.
- − The composition is slightly crowded for a 'minimal' request, though it fits the diorama theme.
Qwen Image 2.0
- + Clean, minimalistic layout that focuses on the food.
- + Accurate colors and solid background as requested.
- − Fails to capture the 'miniature 3D cartoon' aesthetic, appearing more like a standard photo.
- − Text and flag icon are plain and lack the stylized visual appeal requested.
- − The diorama base is just a simple wooden block rather than a detailed scene.
Verdict: GPT Image 2 is the clear winner as it perfectly captures the '3D cartoon' and 'isometric miniature' aesthetic with high-quality stylized textures. Qwen Image 2.0 provides a more literal, photographic interpretation that lacks the creative depth and specific PBR material look requested in the prompt.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 2
- + Excellent depiction of god rays and warm golden light.
- + All four animals are clearly visible and well-structured anatomically.
- + Dynamic sense of movement with the cat reaching for a butterfly.
- − The fox's eyes appear slightly more stylized/cartoonish compared to the others.
- − The depth of field makes the foreground flowers a bit distracting.
Qwen Image 2.0
- + Successfully captures the 'tumbling together' part of the prompt with animals interacting physically.
- + High-quality fur texture and realistic anatomy for the fox and rabbit.
- + Very detailed wildflower environment with a soft, misty background.
- − The cat's front right paw is anatomically muddled where it meets the other animals.
- − The dog's left eye looks slightly unfocused or misaligned with its gaze.
Verdict: Both models followed the complex prompt very well, but GPT Image 2 (Model A) is the winner due to its superior lighting and cleaner composition. While Qwen Image 2.0 (Model B) did a better job with the animals tumbling together, it suffered from minor anatomical clipping/artifacts in the center of the tumble where the paws meet.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to vector emblem style with professional framing.
- + Perfect spelling and elegant typography suitable for a vintage logo.
- + Detailed hatching and texture create a high-quality vintage feel.
- − The steam is a bit wispy and disconnected from the cloche.
Qwen Image 2.0
- + Clean, simple execution of the requested elements.
- + Good color palette adherence to warm brown and cream tones.
- − The steam appears to be coming through the cloche rather than from underneath it.
- − The composition feels slightly unbalanced due to the banner's asymmetrical scroll.
- − Lacks the 'vintage minimalist' sophistication, leaning more towards a generic illustration.
Verdict: GPT Image 2 is the clear winner as it perfectly captures the vintage emblem aesthetic with superior typography and a more cohesive layout. While Qwen Image 2.0 follows the basic prompt instructions, it lacks the professional graphic design quality and textural depth shown in GPT Image 2.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 2
- + Excellent text rendering with no spelling errors across all six steps and additional headers.
- + Comprehensive layout including crew silhouettes and a landing site map that follows the color palette instructions.
- + Professional illustration style with high-quality vector-like assets for the Saturn V and Lunar Module.
- − The step icons are more like detailed illustrations than the 'flat-vector' style requested in the prompt.
- − The 'Apollo 11' badge in the top right contains some slightly blurred text.
Qwen Image 2.0
- + Follows the 'flat-vector' style more closely with simplified shapes and icons.
- + Clean, vertical composition that logically flows through the mission steps.
- − Spelling error present in step 3: 'Translunjar' instead of 'Translunar'.
- − Text formatting is inconsistent, with some labels overlapping their icons (e.g., 'Lunar Orbit').
- − Missing some of the requested NASA-inspired palette's richness, feeling a bit sparse compared to Model A.
Verdict: GPT Image 2 provides a much more polished and professional-grade infographic with perfect spelling and a highly organized layout. While it leans slightly more toward detailed illustration than flat-vector, it successfully incorporates all prompt elements including the crew and specific palettes. Qwen Image 2.0 follows the aesthetic prompt well but is held back by a spelling error and poor text placement.
Explore each model
Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request