OpenAI's state-of-the-art image generation model with better instruction following and adherence to prompts
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 1.5
#6 of 62 in Text-to-Image
Qwen Image 2512
#30 of 62 in Text-to-Image
Where the votes landed
GPT Image 1.5
84.6%
win rate
Ties
0.0%
Qwen Image 2512
15.4%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1.5
- + Perfect adherence to the prompt's spatial instructions.
- + Excellent glass refraction showing the plant through the cube.
- + High-quality texture on the red book and wooden table.
- − The sphere is quite large relative to the cube, making the 'small' descriptor subjective.
- − The base of the cube looks slightly mirrors-like rather than simple glass.
Qwen Image 2512
- + Good lighting and realistic material textures.
- + Follows the prompt for all required elements.
- − Significant optical inconsistency where a second blue sphere appears to be inside the glass wall.
- − The sphere is off-center, making the composition feel less balanced.
- − The glass has a heavy teal tint compared to the clear glass in Model A.
Verdict: GPT Image 1.5 is the clear winner as it correctly handles the transparency and refractions of the glass cube, showing the plant behind it naturally. Qwen Image 2512 suffers from a major hallucination where a duplicate blue sphere is embedded within the glass pane on the right, which ruins the realism of the scene.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the 'repairing' action with tools visible on the ground.
- + Highly realistic lighting and reflections on the wet pavement.
- + Perfectly captured the candid feel with the subject engrossed in work.
- − The car in the background is sharp rather than showing the requested motion blur.
- − The rain is very subtle, almost difficult to see.
Qwen Image 2512
- + Strong bokeh and shallow depth of field as requested.
- + Effective motion blur on the background vehicles.
- + High facial detail and natural skin texture.
- − The subject is posing for a portrait rather than 'repairing' the bicycle as requested.
- − The bicycle anatomy is slightly warped, particularly the handlebars and frame connection.
- − Minimal visual evidence of rain beyond wet ground.
Verdict: GPT Image 1.5 is the winner because it correctly depicts the 'repairing' action with a candid, storytelling atmosphere, whereas Qwen Image 2512 produces a static portrait that ignores the core activity of the prompt. While Qwen better follows instructions for motion blur, the fundamental failure to show the man working makes it a less successful interpretation of the scene.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1.5
- + Excellent photorealistic skin texture with subtle grime and scars.
- + Rich warm lighting that naturally reflects off the intricate engravings of the armor.
- + Highly detailed cloth and leather textures on the underlayer.
- − The bokeh sparks are a bit large and slightly distracting.
- − The hair braids are somewhat messy compared to the specific 'beads' request.
Qwen Image 2512
- + Perfect adherence to the 'hair braided with small beads' prompt with clear colorful beads.
- + Strong composition with a visible torch source adding to the narrative.
- + Intricate and sharp engraving details on the plate armor.
- − The facial skin texture is slightly smoother and less 'battle-worn' than Model A.
- − The bokeh sparks appear a bit more synthetic in their placement.
Verdict: Both models performed exceptionally well, capturing the atmosphere and technical details requested. GPT Image 1.5 wins slightly on pure photorealism and texture depth, while Qwen Image 2512 followed the specific instruction for beads in the hair more accurately. Overall, GPT Image 1.5 feels more like a cinematic close-up due to the superior lighting and skin shaders.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering with no spelling errors
- + Includes specific logical categories and descriptions for items
- + High-quality, realistic food photography that matches the text labels
- − The 'grid' layout for photos is a bit irregular
- − Image is cropped tightly at the bottom
Qwen Image 2512
- + Strong minimalist aesthetic with a clean grid layout
- + Effective use of white space and bold typography
- − Contains significant gibberish text and spelling errors
- − Combines pizza and mains into one incoherent section
- − The food images are repetitive and low-detail
Verdict: GPT Image 1.5 is the clear winner as it produces a fully functional, professional-grade menu with perfect text and logically categorized items. Qwen Image 2512 follows the requested grid layout more strictly but fails significantly on text legibility and content accuracy.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1.5
- + Excellent photorealism with detailed textures on the meat patty and buns
- + Perfectly followed all text requirements including 'LIMITED TIME ONLY'
- + Dynamic and intense use of light and glowing embers that fills the frame
- − The composition is slightly crowded at the top edge
Qwen Image 2512
- + Clean layout with great use of negative space
- + Accurate 3D rendering of the 'MAGIC BURGER' title with fire effects
- + Clear starburst design for the price point
- − Missed the word 'TIME' in the secondary text, displaying only 'LIMITED ONLY'
- − The burger components are less 'exploded' and more stacked compared to the other model
- − The background is less immersive and dynamic than requested
Verdict: GPT Image 1.5 is the clear winner for its superior adherence to the text prompt and more convincing photorealistic detail. While Qwen Image 2512 has a clean design, it failed to include the word 'TIME' in the secondary message and produced a less dynamic composition compared to the fiery, high-motion output of GPT Image 1.5.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1.5
- + Perfect spelling of all menu items including 'Risotto'.
- + Excellent chalk texture with realistic smudges and dusty residues.
- + Very realistic handwriting that feels organic and non-systemic.
- − The layout is a bit sparse with significant empty space at the bottom.
- − The handwriting style is slightly less 'elegant cursive' for the title as requested.
Qwen Image 2512
- + Strong composition that fills the board effectively with a cozy cafe background.
- + Beautifully rendered cursive calligraphy for the title and items.
- + Text is very high contrast and easy to read.
- − Includes a spelling error in 'Risitto' (should be Risotto).
- − The chalk looks slightly like a digital brush rather than a physical chalk stick compared to Model A.
Verdict: GPT Image 1.5 wins on technical accuracy and realism, correctly spelling all menu items and providing a more authentic chalk-on-blackboard texture. While Qwen Image 2512 has a more appealing composition and background, the spelling error in 'Risitto' makes it less useful for a final output.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1.5
- + Excellent cinematic lighting and complex background details
- + Great texture on the horse's fur and the lunar dust
- + Correctly identifies the core concept of an astronaut riding a horse
- − Failed the negative constraint: the astronaut is on top, not the horse
- − The astronaut's left leg appears to be fused with the horse's neck
Qwen Image 2512
- + High clarity in the rendering of the astronaut and horse
- + Clean composition with a beautiful Earth curvature background
- − Failed the specific negative constraint: the astronaut is riding the horse, not vice versa
- − The lighting on the horse is somewhat flat compared to the astronaut
- − Astronaut's face and hands have slight anatomical inconsistencies
Verdict: Both GPT Image 1.5 and Qwen Image 2512 failed to follow the surreal constraint 'horse on top, not vice versa,' instead providing standard interpretations of an astronaut riding a horse. GPT Image 1.5 is preferred because it offers a much more detailed, cinematic environment with superior lighting and atmospheric effects compared to the simpler composition of Qwen Image 2512.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1.5
- + Excellent photorealism with cinematic lighting and depth of field.
- + The capybara's fur and expression are highly detailed and convincing.
- + Superior adherence to the specific hat design with legible 'TAXI' text.
- − The capybara's paws look slightly more like hands/claws than natural capybara anatomy.
Qwen Image 2512
- + Successfully captures all prompt elements including the businesswoman and the driver.
- + Good balance in the composition showing more of the car interior.
- − The capybara's front paws are incorrectly rendered, showing too many digits and a humanoid structure.
- − Lighting on the capybara's face feels a bit flat compared to the background.
Verdict: GPT Image 1.5 is the winner due to its superior photorealistic textures and lighting, which make the absurd scene feel more grounded. While both models followed the prompt instructions well, the fine details in the capybara's fur and the legibility of the taxi hat text give GPT Image 1.5 the edge over Qwen Image 2512.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1.5
- + Excellent typography with perfect spelling in all requested text fields.
- + Superior vintage aesthetic with a dark, textured parchment feel.
- + Cohesive lighting that blends the glowing elements naturally with the environment.
- − The 'frights' text on the scroll has a very slight trailing artifact/underline.
- − The background graveyard is a bit cluttered compared to the clean layout of the second image.
Qwen Image 2512
- + Clean, high-resolution rendering of the central jack-o-lantern.
- + Great use of depth of field with the twisted trees in the background.
- + Very clear and readable gothic-style font choice.
- − Contains a spelling error in the main title: 'Hallowern' instead of 'Halloween'.
- − The overall lighting feels a bit more like a modern digital painting than a vintage parchment poster.
- − The blue-tinted sky deviates slightly from the requested dark parchment color palette.
Verdict: GPT Image 1.5 is the clear winner because it correctly spelled every word in the complex text prompts, whereas Qwen Image 2512 failed with 'Hallowern'. Additionally, GPT Image 1.5 captured the 'vintage gothic parchment' texture much more effectively, providing a more authentic atmosphere for a physical invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering with clean, bold typography and a perfect flag icon.
- + Highly realistic textures on the sushi and wood materials that still maintain a miniature look.
- + Very high level of detail in the props, including the teapot, soy sauce bottle, and individual rice grains.
- − The composition is slightly crowded with too many extra items like the teapot and cup.
Qwen Image 2512
- + Follows the 'cartoon scene' and 'minimal garnish' instructions more closely regarding simplicity.
- + Excellent 3D miniature diorama feel with a clay-like soft texture.
- + Very clean layout that emphasizes the central plate.
- − The flag icon is integrated oddly next to the word 'Sushi' rather than as a distinct element.
- − The text styling is a bit more generic and less 'premium' than Model A.
- − Lower overall resolution and clarity compared to Model A.
Verdict: GPT Image 1.5 produced a superior image in terms of technical execution, featuring crisp text, a perfect flag icon, and beautiful PBR material rendering. While Qwen Image 2512 captured the 'cartoon' aesthetic well, its text integration was weaker and the image lacked the high-clarity polish found in GPT Image 1.5.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1.5
- + Captures the 'tumbling' and 'playful' aspect of the prompt much better than Model B.
- + Excellent rendering of golden hour light, god rays, and atmospheric dew sparkles.
- + Highly expressive and varied facial expressions on each animal.
- − Anatomical issues with the kitten, including an extra paw and a confusing orientation of limbs.
- − The butterfly's scale is a bit large compared to the animals.
Qwen Image 2512
- + Clean and anatomically correct subjects with very clear, high-resolution textures.
- + Well-organized composition that ensures all four animals are clearly visible and front-facing.
- + Beautifully detailed fur and sharp focus on the faces.
- − The animals are posing for a portrait rather than 'playfully chasing and tumbling' as requested.
- − The lighting and environment feel a bit more static and less magical than the god rays in Model A.
- − Butterflies feel somewhat pasted onto the scene rather than part of the action.
Verdict: GPT Image 1.5 does a much better job of capturing the joyful, dynamic energy of animals tumbling and chasing butterflies in a magical atmosphere, though it suffers from significant AI artifacts in the kitten's anatomy. Qwen Image 2512 produces a much cleaner, higher-quality technical image with perfect anatomy, but it ignores the action-oriented parts of the prompt in favor of a static group portrait. GPT Image 1.5 is the preferred choice for its superior interpretation of the 'wholesome vibe' and interaction requested.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1.5
- + Perfect adherence to all text requirements including 'Caffè Florian' and 'Est. 1720'.
- + Clean, minimalist vector aesthetic that works well as a logo.
- + Professional layout with a well-integrated banner.
- − Slightly generic steam effect compared to the illustrative quality of the other model.
Qwen Image 2512
- + Beautiful hand-drawn illustrative style with great woodcut texture.
- + High-quality rendering of the smoke/steam.
- + Accurate text rendering for both terms.
- − The composition is a bit more crowded than a typical 'minimalist' logo.
- − Small artifacts in the crossbar of the 'F' in Florian.
Verdict: Both models followed the prompt exceptionally well, producing accurate text and thematic elements. GPT Image 1.5 is preferred for its cleaner, more 'minimalist' vector logo aesthetic, whereas Qwen Image 2512 leaned more into a complex illustration style that is slightly less practical for a versatile logo.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1.5
- + Excellent text rendering with no spelling errors across all six steps.
- + Perfect adherence to the requested flat-vector style with crisp, consistent iconography.
- + Strong logical flow clearly dividing the mission stages into distinct panels.
- − Misses the 'translunar trajectory arc' as a standalone step, merging it into a panel with the moon.
Qwen Image 2512
- + Includes all specific icons requested in the prompt, such as the trajectory arc.
- + Captures the requested muted red and navy color palette effectively.
- − Significant text rendering issues with multiple misspellings like 'TranslauraJ' and 'Desceeint'.
- − Poor layout logic with repeated numbers and overlapping text.
- − Literal interpretation of prompt meta-text, including the phrase 'Steps stop at landing' as part of the graphic.
Verdict: GPT Image 1.5 is the clear winner as it produces a professional, usable infographic with perfect spelling and a clean layout. While Qwen Image 2512 includes more of the specific architectural requests, its execution is marred by significant typos, confusing numbering, and poor text placement.
Explore each model
Improved version of Alibaba's Qwen image model with better text rendering, finer natural textures, and more realistic human generation.