OpenAI's state-of-the-art image generation model with arbitrary resolution up to 4K and strong instruction following
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 2
#4 of 62 in Text-to-Image
Qwen Image 2512
#30 of 62 in Text-to-Image
Where the votes landed
GPT Image 2
100.0%
win rate
Ties
0.0%
Qwen Image 2512
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 2
- + Excellent photographic quality with realistic textures on the book and table.
- + Accurate interpretation of 'partially visible through the glass' for the green plant.
- + Strong spatial logic and consistent lighting from the left window.
- − The glass cube has double edges that look more like an open frameset than a single solid glass pane.
Qwen Image 2512
- + Successfully follows all prompt instructions including object placement and colors.
- + Good reflection of the blue sphere on the glass base.
- − The glass panes appear inconsistently thick and slightly distorted.
- − The book is floating slightly above the glass surface instead of resting on it.
- − The green plant in the background lacks the detail and clarity found in the other image.
Verdict: GPT Image 2 is the superior image due to its exceptional photographic realism and coherent spatial relationships. While Qwen Image 2512 follows the prompt accurately, it suffers from minor physics issues like a floating book and less refined textures compared to the realistic lighting and plant detail in GPT Image 2.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the 'imperfect framing' and 'candid' aspects of the prompt
- + Superior technical realism in reflections and the wet pavement texture
- + Highly realistic skin texture and age spots on the subject
- − The motion blur on the passing cars is slightly too frozen compared to the prompt's request for blur
- − A few minor anatomical glitches in the hands while holding tools
Qwen Image 2512
- + Good bokeh and depth of field effect
- + Effective inclusion of rain streaks and wet atmosphere
- + Strong emotional connection through the subject's expression
- − Failed the 'candid' prompt as the subject is posed and looking directly into the lens
- − Bicycle geometry is nonsensical, especially the handlebars and the dual front-wheel visual confusion
- − Fails the 'repairing' action, as the man is just squatting behind the bike
Verdict: GPT Image 2 is the clear winner for its adherence to the 'candid' and 'repairing' aspects of the prompt. While Qwen Image 2512 produces a more traditional portrait, it fails significantly on the technical details of the bicycle and creates a posed shot rather than the requested candid moment.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 2
- + Excellent photorealistic skin texture and lifelike eyes.
- + Natural incorporation of small beads in subtle braids.
- + Highly detailed and realistic engraving on the weathered armor.
- − The 'faint scars' are very difficult to see, leaning more toward just dirt.
- − The warm torchlight effect is quite subtle compared to the prompt's emphasis.
Qwen Image 2512
- + Perfect adherence to all prompt elements, including visible scars and bokeh sparks.
- + Strong composition with a clear light source and golden-hour lighting.
- + Great leather and bead details throughout the hair and chest piece.
- − The torch flame in the background looks slightly artificial.
- − Hair braids appear a bit stiff and repetitive compared to natural hair.
Verdict: Both models performed exceptionally well on this complex prompt. GPT Image 2 captures a more realistic, cinematic quality with superior skin rendering, whereas Qwen Image 2512 follows the specific prompt instructions like 'visible scars' and 'bokeh sparks' much more literally and effectively. Qwen Image 2512 wins slightly due to better adherence to the lighting and battle-worn features requested.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 2
- + Excellent text rendering with legible, coherent English describing realistic dishes.
- + Highly professional layout with expert use of white space and hierarchy.
- + Direct adherence to the section requirements (Appetizers, Pizza, Mains) with relevant high-quality imagery.
- − Small logo artifact in the 'Nova' text where it slightly overlaps the red lines.
Qwen Image 2512
- + Accurately reflects the 'grid' request for colorful food photos.
- + Good use of vibrant color accents through the icon system.
- + Appropriate minimalist font choices for a modern design.
- − Text is entirely nonsensical and full of AI artifacts/gibberish.
- − Layout logic is poor, with photo content not corresponding to the text sections (e.g., pizzas mixed everywhere).
- − Section headings are misspelled ('Appetiizers', '/Means').
Verdict: GPT Image 2 (Model A) is the clear winner as it produces a completely usable and professional-grade menu design with perfect English text and logical organization. Qwen Image 2512 (Model B) adheres to the visual grid request but fails significantly on text legibility and the logical placement of food items relative to their headers.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 2
- + Excellent text integration with a consistent fiery, glowing effect as requested.
- + Highly detailed food textures, especially on the meat patty and fresh lettuce.
- + Strong sense of motion with sauce droplets and flying embers that fit the 'exploded' theme.
- − The composition feels slightly crowded with the large text elements pressing against the burger.
Qwen Image 2512
- + Natural, clean lighting on the burger components makes them look very appetizing.
- + Good use of space in the composition, allowing the 'exploded' burger to take center stage.
- + Accurate text placement for the price starburst.
- − Missed the 'TIME' in 'LIMITED TIME ONLY', rendering it as 'LIMITED ONLY'.
- − The text effects are inconsistent; the main title is fiery, but secondary text is plain yellow.
- − The background is less dynamic and lacks the 'glowing embers' intensity requested.
Verdict: GPT Image 2 is the superior output because it followed all text instructions perfectly and maintained a consistent 'fiery' aesthetic across all elements. While Qwen Image 2512 has a clean layout, it failed on the specific wording of the secondary message and lacked the intense atmospheric energy described in the prompt.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 2
- + Perfect spelling on all menu items including complex terms.
- + Highly realistic chalk texture with slight smudging and authentic pressure variations.
- + Excellent composition that feels natural for a physical cafe space.
- − The slanted handwriting is slightly less decorative than the title request implied.
Qwen Image 2512
- + Clear, legible text with an attractive calligraphic style.
- + Good use of vertical space on a portrait-oriented board.
- − Spelling error present in 'Risitto' (should be Risotto).
- − Handwriting looks slightly more like a digital font than natural chalk strokes.
Verdict: GPT Image 2 is the superior output because it followed all text prompts perfectly, including difficult spelling, while maintaining a very realistic chalk aesthetic. Qwen Image 2512 had a spelling error in one of the primary menu items and the text style felt more calculated and less like authentic handwriting.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 2
- + Excellent adherence to the 'horse on top' spatial instruction
- + Highly surreal and creative interpretation of the prompt
- + Detailed textures on the spacesuit and lunar surface
- − The horse's front legs and hoof-gloves are anatomically confusing
- − The harness logic is physically impossible
Qwen Image 2512
- + High visual quality with realistic lighting
- + Dynamic and cinematic composition of the horse leaping
- − Failed the negative constraint; the astronaut is on top of the horse
- − Clipped tail at the bottom edge of the frame
Verdict: GPT Image 2 followed the specific and difficult instruction to place the horse on top of the astronaut, creating a truly surreal image as requested. Qwen Image 2512 ignored the spatial positioning constraint entirely, producing a standard 'astronaut on horse' image which makes it a failure for this specific prompt.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 2
- + Excellent photorealism with cinematic lighting
- + Natural textures on the capybara's fur and the jacket
- + Composition feels like an authentic candid photograph from the street
- − The capybara's paws lack clear claws/definition on the steering wheel
- − The 'T' on the cap is a bit generic
Qwen Image 2512
- + Perfectly captures the 'bored' and 'normal' expression of the passenger
- + Composition clearly shows the full interior context and the street ahead
- + Accurate double-paw placement on the steering wheel
- − The passenger's hands and phone interaction look slightly distorted
- − The capybara's paws look more like human hands wearing gloves than animal paws
- − Lighting is a bit flatter and less cinematic than the competitor
Verdict: GPT Image 2 (Model A) wins on sheer visual quality and realism, feeling like a genuine photographic still with beautiful lighting. However, Qwen Image 2512 (Model B) followed the character expression prompts more accurately, specifically regarding the woman's 'bored' look and the specific driving posture. Model A is preferred for its superior artistic execution and more believable integration of the capybara into the environment.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 2
- + Excellent typography with perfect spelling in all requested text fields.
- + Highly detailed and intricate gothic border featuring skulls, webs, and thorns.
- + Superior atmospheric depth with the inclusion of a bridge and city skyline matching the 'NYC' location.
- − The composition is a bit crowded with many competing dark elements.
Qwen Image 2512
- + Strong cinematic lighting with a clear focal point on the jack-o-lantern.
- + Clean and legible layout for the event details at the bottom.
- + Good representation of the twisted trees and misty background.
- − Misspelled the word 'Halloween' as 'Hallowern' in the main title.
- − The 'webs' in the border corners look a bit like geometric patterns rather than organic spider webs.
Verdict: GPT Image 2 is the clear winner as it followed all textural instructions perfectly, whereas Qwen Image 2512 had a significant misspelling in the main title. GPT Image 2 also went the extra mile by including architectural elements that hinted at the NYC location requested in the prompt.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 2
- + Excellent PBR material rendering with realistic sheen on fish and textures on stone
- + Perfect text layout and high-quality 3D typography
- + Complex and highly detailed composition that stays within the miniature theme
- − The garnish and base elements are slightly more complex than 'minimal'
Qwen Image 2512
- + Captured the 'soft refined textures' and 'cartoon scene' style perfectly
- + Good adherence to the 'minimal garnish' request
- + Clean isometric perspective
- − Text rendering is slightly inconsistent in alignment and font weight
- − Flag icon is placed awkwardly next to the text rather than having its own clear space
- − Lower overall detail in the sushi materials compared to the PBR request
Verdict: GPT Image 2 is the superior overall image, featuring much higher quality material rendering and more polished typography. While Qwen Image 2512 captures the 'cartoon' aesthetic well, it lacks the professional finish and crisp details found in GPT's version.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 2
- + Dynamic 'playfully chasing' action is well-captured
- + Beautiful rim lighting and backlighting effect on the fur
- + Includes all requested animals with distinct, active poses
- − The fox's front right paw has anatomical issues/blending
- − Floating butterflies lack depth and ground shadow
Qwen Image 2512
- + Excellent fur texture and facial detail on the animals
- + Perfectly captures the requested 'big expressive eyes'
- + More coherent composition with the animals grouped together
- − Static pose ignores the 'playfully chasing and tumbling' instruction
- − Anatomy issues with multiple paws blending into one another in the center lower area
Verdict: GPT Image 2 is the better interpretation of the prompt because it successfully captures the movement of chasing and tumbling in a meadow, whereas Qwen Image 2512 opted for a static, posed portrait. While Qwen offers slightly cleaner facial details, GPT Image 2 better utilizes the 'god rays' and lighting to create a cinematic, hyper-photorealistic environment.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 2
- + Excellent typography including the correct grave accent on 'Caffè'.
- + Perfectly balanced composition with a professional emblem frame.
- + Clean, vector-style execution that aligns with luxury branding.
- − The steam effect is very stylized and simple compared to Model B.
Qwen Image 2512
- + Dynamic and detailed rendering of the cloche and steam.
- + Strong retro vibe with bold script typography.
- + Excellent use of warm brown and cream tones with depth.
- − Incorrectly used an acute accent instead of a grave accent on 'Caffé'.
- − The composition feels slightly crowded without a containing border.
Verdict: GPT Image 2 is the superior logo as it correctly handles the typography of the name 'Caffè' and provides a more balanced, professional emblem layout. While Qwen Image 2512 has more impressive illustrative detail in the cloche and steam, its spelling error and lack of a framing element make it less effective as a functional logo.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 2
- + Excellent typography with perfect spelling and clear, readable labels.
- + Followed all six requested steps in a logical, chronological sequence.
- + Professional layout with balanced use of the NASA-inspired color palette.
- − Illustrations are slightly more detailed than a strictly 'flat-vector' style would dictate.
Qwen Image 2512
- + Successfully captured a clean, flat-vector aesthetic with muted colors.
- + Included the requested icons in a simplified, communicative style.
- − Frequent spelling errors in labels such as 'TranslauraJ' and 'Desceeint'.
- − Confusing and repetitive numbering for the mission steps.
- − Literal text from the prompt ('Steps stop at landing:') was accidentally included in the image design.
Verdict: GPT Image 2 is significantly superior, delivering a professional-grade infographic with perfect text rendering and a logical flow of the requested mission steps. In contrast, Qwen Image 2512 suffers from significant legibility issues, repetitive numbering, and severe spelling errors throughout the graphic.
Explore each model
Improved version of Alibaba's Qwen image model with better text rendering, finer natural textures, and more realistic human generation.