OpenAI's previous image generation model that accepts both text and image inputs and produces image outputs
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
GPT Image 1
#32 of 62 in Text-to-Image
Qwen Image 2512
#30 of 62 in Text-to-Image
Where the votes landed
GPT Image 1
0.0%
win rate
Ties
0.0%
Qwen Image 2512
100.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1
- + Excellent photographic quality with realistic textures on the book and wooden table.
- + Perfect logical coherence with the sphere clearly inside the cube and the book resting naturally on top.
- + Very clean glass rendering with believable reflections and refractions.
- − The sphere appears to be floating slightly above the bottom surface of the cube.
Qwen Image 2512
- + Natural handling of reflections on the interior faces of the glass cube.
- + Follows all spatial instructions in the prompt correctly.
- + Good bokeh effect on the background plant.
- − The cube's glass appears more like a mirror on some faces, obscuring the interior slightly.
- − The texture of the red book is a bit less realistic compared to Image A.
Verdict: Both models followed the prompt perfectly. GPT Image 1 is the winner due to its superior photographic clarity, specifically the realistic texture of the paper in the book and the more transparent quality of the glass, whereas Qwen Image 2512 has slightly distracting internal reflections that make the cube look partially silvered.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
GPT Image 1
- + Excellent shallow depth of field and bokeh quality
- + Highly realistic skin textures and lighting
- + Captures the 'repairing' action more convincingly
- − Slightly less motion blur on the cars than requested
- − Bicycle geometry is a bit simplified in the rear
Qwen Image 2512
- + Captures the 'imperfect framing' and wider street context well
- + Better sense of the wet pavement reflections
- + Includes more visible rain streaks
- − The subject is posing rather than repairing the bike
- − The bike proportions are warped, appearing too small for the man
- − Significant anatomical errors with the hands
Verdict: GPT Image 1 (Model A) provides a much more cinematic and realistic rendering with superior skin textures and a believable depth of field that aligns with a 50mm lens. Qwen Image 2512 (Model B) captures the requested street atmosphere well, but falls short due to significant anatomical issues with the hands and a lack of focus on the actual 'repairing' action.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1
- + Excellent depiction of ornate, swirled engraving on the plate armor
- + Subtle and realistic integration of grime and skin texture
- + Strong focus on a battle-worn facial expression
- − Lacks the requested leather straps and cloth underlayer detail
- − The beads in the hair are very sparse compared to the multiple braids shown in Model B
Qwen Image 2512
- + Perfectly captures all elements including leather straps and cloth textures
- + Accurate bead detailing in multiple complex braids
- + Dynamic lighting with visible torch and high-quality bokeh sparks
- − The scars look slightly 'painted on' compared to the skin texture
- − The armor engraving is slightly less intricate than Model A's
Verdict: Both models followed the prompt closely, but Qwen Image 2512 is the winner for including every specific detail from the prompt, including the leather straps, cloth layers, and extensive beaded braids. While GPT Image 1 has beautiful armor engraving, it missed several of the texture-specific keywords and a more varied use of lighting.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with clean, almost-perfect English text
- + High-quality, realistic food photography with vibrant colors
- + Strong minimalist aesthetic that looks like a real professional layout
- − Layout is slightly cut off at the bottom
- − Text descriptions beneath names are repetitive gibberish
Qwen Image 2512
- + Successfully creates a full document mockup view within a grid
- + Good use of color blocking for organizational sections
- − Text is largely illegible and contains numerous spelling errors
- − Food photos are small and lack the clarity of the competing model
- − Layout feels cramped and cluttered compared to the minimalist prompt
Verdict: GPT Image 1 is the clear winner as it produces a high-fidelity design with sharp, professional typography and appetizing photography that fits the 'modern minimalist' prompt. Qwen Image 2512 fails significantly on text legibility and overall image clarity, appearing more like a low-resolution thumbnail than a professional menu design.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with a consistent glowing ember effect on all text elements.
- + High-quality photorealistic textures on the burger ingredients.
- + Superb adherence to the 'exploded' concept with clear separation of all layers.
- − Failed to include the '6' in the price, displaying '.99' instead of '€6.99'.
- − The starburst shape is a bit simplified and digital looking.
Qwen Image 2512
- + Perfect accuracy on all required text strings, including the price.
- + Highly dynamic sense of motion with debris and sauce droplets flying out.
- + Rich, multi-layered composition with flames in the foreground and background.
- − The 'LIMITED ONLY' text is missing the word 'TIME'.
- − The typography style is less cohesive than Model A, with various fonts and effects mixed together.
Verdict: GPT Image 1 captures the 'exploded' burger aesthetic much better with its vertical separation and consistent glowing text, but it fails significantly on the specific price requested. Qwen Image 2512 provides a more visceral sense of motion and correctly renders the price, making it more functional as an actual advertisement despite missing one word in the secondary message.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
GPT Image 1
- + Excellent chalk texture throughout all characters
- + Perfect spelling for all requested menu items
- + Consistent letterforms that look authentically like the same person's handwriting
- − The 'cursive' requirement for the title is not fully met as it is more of a print style
- − Composition is a bit crowded vertically
Qwen Image 2512
- + Successfully followed the request for elegant cursive in the title and menu items
- + Great environmental storytelling with the blurred café background
- + Better use of space and layout for a menu board
- − One spelling error ('Risitto' instead of 'Risotto')
- − The chalk texture is a bit too smooth and looks slightly like a digital brush
Verdict: GPT Image 1 followed the instructions for text content perfectly with zero spelling errors and a very realistic chalk texture, even if it missed the specific 'cursive' stylistic cue. Qwen Image 2512 captured the 'elegant cursive' aesthetic and 'cozy café' atmosphere much better, but suffered from a spelling mistake in the main menu items. GPT Image 1 is the likely winner for its superior text accuracy and authentic-feeling chalk grit.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
GPT Image 1
- + Excellent cinematic lighting and texture on both the suit and horse fur.
- + Higher overall resolution and artistic detail in the background.
- + Stronger surreal atmosphere with the dark starfield and planetary curve.
- − Failed the negative constraint; the astronaut is riding the horse instead of vice versa.
Qwen Image 2512
- + Clear facial detail within the helmet visor.
- + Good anatomical representation of the horse and tack equipment.
- + Sharp contrast between the subject and the bright limb of the Earth.
- − Failed the crucial negative constraint; depicts a human riding a horse.
- − Lighting on the astronaut is inconsistent with the dark background space.
Verdict: Both GPT Image 1 and Qwen Image 2512 completely failed the specific negative constraint to place the horse on top of the astronaut. Because neither followed the core unique instruction, the comparison defaults to visual quality, where GPT Image 1 is superior due to its cinematic lighting, cohesive color palette, and intricate textures. Qwen Image 2512 looks more like a standard photo-composite with flatter lighting.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1
- + Excellent texture on the capybara's fur and the leather taxi hat.
- + Cinematic lighting that accurately reflects a New York night scene.
- + Subtle and professional expression on the capybara's face.
- − The passenger's hand holding the phone is slightly distorted and blurry.
- − The capybara's paws are somewhat indistinctly placed on the steering wheel.
Qwen Image 2512
- + Clearer depiction of both front paws on the steering wheel as requested.
- + The passenger's expression more accurately captures the 'bored/annoyed' look.
- + Very sharp focus and bright, vibrant colors throughout the scene.
- − The capybara's paws look more like primate hands with claws rather than capybara paws.
- − The taxi sign on the hat is a generic crest rather than the requested 'TAXI' text.
Verdict: GPT Image 1 offers a more photorealistic and cinematic atmosphere with high-quality textures, making the scene feel more authentic. Qwen Image 2512 follows the specific framing instructions well (paws on wheel, bored expression) but loses realism due to the odd, hand-like appearance of the capybara's paws and a less convincing taxi hat.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent typography style that matches the gothic aesthetic
- + Very clean layout with clear hierarchical design
- + High artistic quality of the jack-o-lantern illustration
- − Confused the 'Time' and 'Location' labels in the footer
- − Missing the requested '7pm' time entirely
Qwen Image 2512
- + Successfully included all requested text labels and event details without error
- + Detailed border with prominent thorns and webs
- + Strong cinematic lighting with vibrant contrast
- − Includes a spelling error in the main header ('Hallowern')
- − The banner scroll cuts into the bottom text area slightly awkwardly
Verdict: GPT Image 1 has a more sophisticated and polished design but fails on the factual data requested by mixing up the 'Time' and 'Location' labels. Qwen Image 2512 follows the prompt instructions much better regarding the event details, though it suffers from a major typo in the primary heading ('Hallowern').
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1
- + Excellent typography with clean, bold text and perfect spacing.
- + High-quality 3D clay-like textures that perfectly match the 'soft refined' request.
- + Superior lighting and shadow work on the diorama base.
- − The sushi variety is quite limited with only two types shown.
- − The flag icon is a bit large compared to the text.
Qwen Image 2512
- + Includes a wider variety of sushi (nigiri and maki).
- + Creative dioroma base with tiny grass and floral details.
- − Text rendering is lower quality with inconsistent outlines and spacing.
- − The 'JAPAN' text is slightly off-center to the left.
- − The overall image clarity is lower than Model A.
Verdict: GPT Image 1 is the superior choice because it strictly adheres to the 'clean' and 'high-clarity' requirements with professional-grade typography and lighting. While Qwen Image 2512 offers more variety in the sushi itself, its text rendering and overall composition feel less polished and slightly cluttered.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1
- + Excellent sense of motion and 'tumbling' as requested in the prompt
- + Superior lighting with realistic volumetric god rays and soft glows
- + Highly expressive, joyful facial expressions on all four animals
- − The fox kit has slightly unusual dark paws that look a bit like hooves or rounded stubs
- − The butterflies are less detailed compared to those in the second image
Qwen Image 2512
- + Exquisite detail in individual fur strands and butterfly wings
- + The animals are neatly grouped, making for a clear family-portrait style composition
- + Strong adherence to the '8K' and 'hyper-photorealistic' aesthetic
- − Static composition that fails to capture the 'playfully chasing' and 'tumbling' part of the prompt
- − Anatomical oddity where the kitten appears to lack a complete body behind the puppy's leg
- − The lighting feels a bit more artificial and post-processed compared to the natural look of image A
Verdict: GPT Image 1 captures the spirit of the prompt much more effectively by showing the animals in motion, playfully interacting as requested. While Qwen Image 2512 has higher sharpness in the fur textures and butterflies, its static composition and minor anatomical blending issues make it feel less like a cohesive scene and more like a staged collage.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1
- + Clean vector emblem style
- + Accurate spelling including the accent on the 'e'
- + Good textured brown and cream tones
- − Generated on a black background instead of the requested light background
- − Minimalist steam is a bit too simplified and looks thin
Qwen Image 2512
- + Perfect adherence to the light background and subtle texture request
- + Excellent illustrative vintage style
- + Clear, legible typography and banner
- − The 'i' in Florian has an unnecessary dot above it in a script that doesn't usually feature it that way
Verdict: Qwen Image 2512 followed the background instructions perfectly, providing a beautiful light-textured vintage logo with detailed steam and an elegant banner. While GPT Image 1 captured a more modernist minimalist style and got the accent mark correct, it failed to provide the light background requested in the prompt.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
GPT Image 1
- + Excellent adherence to the clean, flat-vector style requested.
- + Clear and legible typography with correct spelling for the most part.
- + Strict adherence to the specified NASA-inspired color palette.
- − Layout is a bit chaotic with text labels not aligning perfectly with their icons.
- − Has a spelling error in 'EARLLUNAR' and misses Lunar Orbit as a distinct step.
Qwen Image 2512
- + Strong composition that feels like a professional poster layout.
- + High level of detail on the spacecraft icons while maintaining a vector look.
- + Includes more of the requested sequential steps in a logical flow.
- − Numerous spelling errors in technical terms (e.g., 'TranslauraJ', 'Desceeint').
- − Text rendering is smaller and less crisp than its competitor.
- − Included the prompt text 'Steps stop at landing' literally into the image.
Verdict: GPT (Image A) better captures the requested 'flat-vector' aesthetic and has much clearer, more accurate text, though it misses some of the sequential steps. Qwen (Image B) has a superior infographic layout and more detailed icons, but suffers from significant spelling errors and the inclusion of literal prompt text in the design. GPT is the winner for its professional finish and stylistic accuracy.
Explore each model
Improved version of Alibaba's Qwen image model with better text rendering, finer natural textures, and more realistic human generation.