OpenAI's state-of-the-art image generation model with better instruction following and adherence to prompts
Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.
GPT Image 1.5
#7 of 62 in Text-to-Image
Qwen Image Max
#35 of 62 in Text-to-Image
Where the votes landed
GPT Image 1.5
100.0%
win rate
Ties
0.0%
Qwen Image Max
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to lighting instructions with a clear source from the left.
- + Realistic glass refraction and reflections on the sphere and table.
- + The placement of the plant behind the glass creates a natural sense of depth.
- − The sphere is quite large relative to the cube, making the 'small' descriptor in the prompt a bit loose.
Qwen Image Max
- + Sophisticated photographic quality with realistic wood grain and soft shadows.
- + Captures the book and sphere textures very well.
- − The sphere appears to be floating mid-air inside the cube with no support, which looks physically impossible.
- − The glass cube has internal mirror-like reflections that don't match the environment properly.
- − The plant is placed to the side rather than clearly 'behind' the cube as requested.
Verdict: GPT Image 1.5 is the clear winner as it follows the spatial instructions of the prompt more accurately, placing the plant behind the cube and resting the sphere on the base. Qwen Image Max produces a high-quality aesthetic but fails on physics and composition, making the sphere float and placing the plant toward the background side.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
GPT Image 1.5
- + Exceptional metallic textures with realistic light diffusion and dirt accumulation.
- + Incredible facial detail and lifelike eyes that appear moist and reflective.
- + Perfect adherence to the 'shallow depth of field' and 'bokeh sparks' request.
- − The braiding is somewhat loose compared to the more distinct braids in Model B.
Qwen Image Max
- + Distinct and intricate braiding with colorful beads as requested.
- + Strong character expression that conveys the 'battle-worn' theme through deep wrinkles and scars.
- + Clear engraving patterns on the plate armor.
- − The lighting on the armor looks a bit matte and lacks the realism of polished metal reflecting fire.
- − Some minor anatomical issues around the lips and asymmetrical eye rendering.
Verdict: GPT Image 1.5 delivers a superior technical performance with photorealistic textures, specifically on the engraved armor and skin, which makes for a more immersive portrait. While Qwen Image Max does an excellent job with the specific hair braiding and bead request, the overall image quality and lighting of GPT Image 1.5 are more cohesive and visually stunning.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
GPT Image 1.5
- + Excellent adherence to the 'exploded' request with clear vertical separation of all ingredients.
- + High-quality, photorealistic textures on the patty and vegetables.
- + Superior integration of the starburst element with the fiery background.
- − The 'MAGIC BURGER' text is slightly cut off at the top corners.
- − The lighting is very intense, which makes some areas look slightly over-sharpened.
Qwen Image Max
- + Clean, professional layout with well-balanced text elements.
- + Effective fiery effect on the typography that feels very integrated.
- + Good sense of motion and dynamic energy in the splash effect.
- − Failed the 'exploded' prompt as the bun is still resting on the patty.
- − The burger components are mostly bunched together rather than suspended in mid-air.
- − The starburst is more of a light flare than a traditional ad starburst shape.
Verdict: GPT Image 1.5 followed the prompt much more accurately, creating a true 'exploded' view where every ingredient is suspended individually. While Qwen Image Max produced a very clean and professional advertisement, it failed to separate the core components of the burger as requested. GPT Image 1.5 is the winner for its superior composition and adherence to the technical structure of the prompt.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
GPT Image 1.5
- + Excellent photorealistic texture on the capybara's fur
- + Accurate depiction of a classic taxi driver cap with text
- + Superior cinematic lighting and focus depth
- − The paws on the steering wheel look slightly anthropomorphized/distorted
Qwen Image Max
- + Clear composition showing both the driver and the passenger effectively
- + Great city background detail with recognizable street lights
- + Consistent clothing textures
- − The capybara's hands look very human-like and unsettling
- − Visible AI artifacts like the cracked ceiling texture in the car
- − The yellow cap is a generic baseball hat rather than a driver's cap
Verdict: GPT Image 1.5 wins due to its superior photorealistic quality and adherence to the specific 'taxi driver cap' detail. While Qwen Image Max offers a wider perspective, the distorted human-like hands on the capybara and the strange ceiling artifacts make it less visually convincing.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
GPT Image 1.5
- + Excellent color grading with a warm, vintage sepia tone
- + Great texture on the parchment and jack-o-lantern
- + Clean, professional typography that integrates well into the scene
- − The thorn border is a bit messy and indistinct in the corners
Qwen Image Max
- + Strong, defined gothic borders with clear thorn and web motifs
- + Perfect text accuracy for all requested fields
- + Good contrast between the central image and the parchment background
- − The lighting feels a bit disconnected between the cool background and warm pumpkin
- − The scroll banner is slightly warped on the right side
Verdict: GPT Image 1.5 produces a more atmospheric and cohesive piece of art with superior cinematic lighting and texture. While Qwen Image Max has a very clean layout and more defined borders, GPT Image 1.5 feels more like a polished, high-end vintage invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
GPT Image 1.5
- + Excellent 3D miniature diorama feel with detailed base layers.
- + High textural realism for materials like wood, ceramics, and rice.
- + Clean and professional typography layout with a centered flag icon.
- − Includes extra items like a teapot and soy sauce that contrast with the 'minimal garnish' request.
Qwen Image Max
- + Perfect adherence to the 45-degree isometric perspective.
- + Stronger '3D cartoon' aesthetic with cleaner, smoother surfaces.
- + More accurate execution of the 'minimal' request for the diorama base and scene contents.
- − The glass base has some minor perspective clipping on the bottom corner.
- − Typography and flag placement feel a bit crowded at the top.
Verdict: GPT Image 1.5 provides a richer, more detailed scene with impressive PBR material rendering, but goes beyond the 'minimal' requirement. Qwen Image Max better captures the requested '3D cartoon' style and isometric simplicity. Ultimately, GPT Image 1.5 is the winner due to the superior quality of its lighting and texture and the professional integration of the text and graphics.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
GPT Image 1.5
- + Successfully includes all four requested animals (dog, cat, bunny, fox).
- + Excellent rendering of soft fur textures and expressive eyes.
- + Dynamic 'tumbling' composition matches the playful vibe of the prompt.
- − Anatomy of the kitten's paws is slightly distorted with extra toes.
- − The fox's eyes appear a bit more doll-like than realistic.
Qwen Image Max
- + High clarity and excellent use of 'god rays' lighting.
- + Colorful and diverse floral environment.
- + Strong sense of interaction between the animals.
- − Failed to include the 'baby bunny' requested in the prompt.
- − Included two golden retrievers instead of the specified animals.
- − The fox has an unusually long, thin tail that looks slightly unnatural.
Verdict: GPT Image 1.5 adhered better to the specific list of animals provided, successfully including the rabbit which Qwen Image Max missed. While Qwen Image Max has vibrant lighting and colors, it hallucinated a second dog and failed a core part of the prompt requirements.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
GPT Image 1.5
- + Excellent typography with a mix of elegant script and serif styles.
- + Correct inclusion of all requested text elements.
- + Clean, professional-looking illustration of the cloche dome.
- − Ignored the request for a light background, providing a black one instead.
- − The steam element is very simple and lacks the detail seen in the other model.
Qwen Image Max
- + Perfect adherence to the background and texture prompt with a realistic parchment effect.
- + Highly detailed and artistic steam trailing from the cloche dome.
- + Very clean vector-style lines and excellent spacing.
- − Slightly generic typography compared to the more stylistically distinct text in the other model.
Verdict: Qwen Image Max followed every specific instruction, including the light background and subtle texture, whereas GPT Image 1.5 failed by providing a solid black background. While GPT Image 1.5 has slightly more interesting typography, Qwen Image Max is the superior logo overall due to its artistic steam rendering and full prompt adherence.
Explore each model
The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.