Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
Qwen Image 2.0
#34 of 62 in Text-to-Image
Seedream 4.0
#15 of 62 in Text-to-Image
Where the votes landed
Qwen Image 2.0
50.0%
win rate
Ties
0.0%
Seedream 4.0
50.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent texture on the red book cover.
- + Complex reflections within the glass panels.
- + Highly realistic wood grain and lighting.
- − The sphere appears to be floating rather than resting on the bottom surface.
- − The internal reflections are a bit busy and distracting.
Seedream 4.0
- + Natural and soft window lighting as requested.
- + Good spatial arrangement with the sphere resting on the base.
- + Clear visibility of the plant through the glass.
- − The plant appears to be growing inside the cube rather than behind it.
- − The glass cube edges are slightly distorted at the back right corner.
Verdict: Both models followed the prompt instructions well, but Seedream 4.0 provided a more balanced composition and better captured the 'soft window light' aspect. While Qwen Image 2.0 has superior texture detail on the book, its floating sphere feels less grounded than the one in Seedream 4.0.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
Qwen Image 2.0
- + Exceptional skin texture and facial detail
- + Very realistic 'imperfect' framing and camera feel
- + Convincing wet pavement reflections and rain atmosphere
- − The car in the background lacks the requested motion blur
- − An extra leg/person appears unintentionally on the right edge
Seedream 4.0
- + Successfully captured motion blur on the passing car
- + Comprehensive scene with repair tools and context
- + Accurate red bicycle and central framing
- − Skin and hair textures are slightly smoothed/AI-looking
- − The rain effect looks like a static overlay of vertical lines
- − The hand anatomy is a bit muddled during the repair
Verdict: Qwen Image 2.0 produces a much more authentic 'candid' feel with superior skin textures and high-fidelity lighting, though it fails the motion blur request for the car. Seedream 4.0 captures more of the specific prompt details like the motion blur and tools, but the overall image quality looks more synthetic and less like a real photo. Qwen Image 2.0 is preferred for its striking realism and successful execution of the 'no stylization' requirement.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent depiction of textured scars and rough, dirt-stained skin
- + Intricate engraving and realistic metal reflections on the pauldrons
- + Clear inclusion of small colored beads in the braided hair
- − The character has four fingers and no thumb visible on the resting hand
- − Multiple eyes/anatomical oddity in the way the pupils are rendered
Seedream 4.0
- + Strong cinematic lighting and more accurate shallow depth of field
- + Highly realistic eyes and facial features without anatomical errors
- + Very detailed texture on the leather straps and underlayer as requested
- − The hair braids are slightly less complex than the multi-braid style in the first image
- − The light source in the background is a bit blown out compared to the foreground
Verdict: While Qwen Image 2.0 provides excellent detail in the armor engravings and skin texture, it fails on basic anatomy with a deformed hand and strange eye details. Seedream 4.0 delivers a much more cohesive and lifelike portrait with natural-looking eyes and superior lighting that perfectly captures the 'warm torchlight' atmosphere.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent adherence to the 'grid' layout instruction with professional consistency.
- + Includes all three requested textual sections (Appetizers, Pizza, Mains) clearly.
- + Uniform photographic style across all food items creates a cohesive menu feel.
- − The placeholder text for food names is garbled and unreadable.
- − Layout is a bit repetitive, failing to logically separate the three header categories from the nine images.
Seedream 4.0
- + Creates a more dynamic, modern asymmetric grid layout.
- + Text rendering for headers is very clean and professional.
- − Failed to include food item names or prices, only displaying headers.
- − The 'Mains' category shows images of pizza and salad rather than distinct main courses.
- − Poor composition with large gaps of white space that look unfinished.
Verdict: Qwen Image 2.0 followed the prompt's structural requirements much more effectively, providing a consistent grid that includes prices and a logical layout of nine food items. Seedream 4.0 has better individual text rendering for headers, but it failed to organize the content properly, missing item descriptions and misplacing food types under the 'Mains' heading.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
Qwen Image 2.0
- + Perfect adherence to the requested price and text elements.
- + Excellent photorealistic detail on the textures of the meat and melting cheese.
- + Strong layout with embers and smoke enhancing the atmosphere.
Seedream 4.0
- + Highly dynamic sense of motion with the swirling effect.
- + Creative composition with the burger split into two flying segments.
- + The fiery background has a more grounded, physical feel with the coals at the bottom.
- − Failed the price requirement by displaying '€5.99' instead of '€6.99'.
- − Text rendering is slightly less crisp and professional than the competitor.
- − The burger looks more like two separate mini-burgers rather than one exploded stack.
Verdict: Qwen Image 2.0 is the superior choice because it followed all text prompts perfectly, including the specific price, and delivered a higher level of photorealistic detail on the food. While Seedream 4.0 offered a more creative sense of motion, it failed the numerical instruction and resulted in a slightly messy composition.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent text legibility and accuracy
- + Convincing chalk smudge and texture details
- + Highly realistic cafe background
- − The title font leans slightly more toward a clean print-style than 'elegant cursive'
Seedream 4.0
- + Successfully uses a cursive style for the title text
- + Excellent chalk texture and messy chalkboard aesthetic
- + Accurate completion of the 'Brown Butter Chocolate Chip Cookies' prompt constraint
- − The pricing text for the Risotto and Octopus is slightly cramped compared to model A
- − The '9' and '8' in the prices have slight structural inconsistencies
Verdict: Both models handled the complex text rendering task exceptionally well, adhering to the specific prices and menu items. Qwen Image 2.0 provides a cleaner, more legible layout with a high-quality background, while Seedream 4.0 better captured the requested 'elegant cursive' for the title and the gritty, tactile feel of a real chalkboard.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
Qwen Image 2.0
- + High visual quality and clarity in the render.
- + Excellent lighting and realistic textures on the horse's coat and spacesuit.
- + Imaginative environment with floating droplets and a beautiful view of Earth.
- − Completely failed the negative constraint; the astronaut is riding the horse, not the other way around.
Seedream 4.0
- + Dynamic composition with a cinematic feel.
- + Good use of reflections in the helmet visor showing celestial bodies.
- + High resolution with fine details in the horse's mane and suit fabrics.
- − Completely failed the negative constraint; it shows a standard 'astronaut riding a horse' scene.
- − The horse's neck anatomy and connection to the body are slightly awkward in this perspective.
Verdict: Both models failed to adhere to the specific request for 'horse on top, not vice versa,' instead generating the common 'astronaut riding a horse' trope. Qwen Image 2.0 is slightly preferred for its superior clarity, more logical lighting, and more aesthetically pleasing overall composition, whereas Seedream 4.0 has a more chaotic background and slightly less coherent anatomy.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent texture on the capybara's fur and the leather jacket.
- + High visual clarity and vibrant cinematic lighting.
- + Captures the 'professional expression' very well with a stoic look.
- − The passenger is sitting in the front seat instead of the back seat as requested.
- − The perspective makes the car interior feel a bit cramped and distorted.
Seedream 4.0
- + Perfectly follows the spatial instruction of having the businesswoman in the back seat.
- + Accurate rendering of a New York taxi exterior and interior layout.
- + Successful bokeh effect for the Manhattan street lights through the window.
- − The capybara's paws look slightly amorphous and less defined than in Model A.
- − The 'TAT' text on the hat is a bit nonsensical.
Verdict: While Qwen Image 2.0 has superior texture work and lighting, it completely failed the spatial requirement of placing the passenger in the back seat. Seedream 4.0 followed all prompt instructions accurately, including the positioning of the characters and the bored expression of the businesswoman, making it the more successful interpretation.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent typography with clean, legible text for all required sections.
- + Strong parchment aesthetic that maintains a vintage feel while being polished.
- + Clean composition with a well-defined thorny border and symmetrical elements.
- − The atmospheric lighting is a bit flat compared to the other model.
Seedream 4.0
- + Atmospheric cinematic lighting with deep shadows and a vibrant glow from the pumpkin.
- + Dynamic border design using three-dimensional horns and torn paper edges.
- + More 'spooky' feel with realistic textures on the trees and pumpkin.
- − Text rendering on the scroll banner is distorted and difficult to read.
- − Composition feels slightly cluttered with the large thorny border overlapping the background elements.
Verdict: Qwen Image 2.0 is the superior choice for a functional invitation because it rendered all specified text with perfect clarity and professional alignment. While Seedream 4.0 offered more impressive cinematic lighting and texture, its failure to legibly render the scroll banner text makes it less effective as a graphic design piece.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent text rendering with clear, bold typography.
- + Highly realistic textures on the fish and wood grain.
- + Clean, professional lighting that highlights the food naturally.
- − Failed the architectural aspect of the prompt by using a photographic style rather than a 3D cartoon style.
- − The flag icon is placed to the side rather than centered as requested.
Seedream 4.0
- + Perfectly captured the requested 3D cartoon/miniature aesthetic.
- + Accurate isometric perspective with a clear diorama-style raised base.
- + Followed center-alignment instructions for the text and flag more closely than the competitor.
- − The text 'SUSHI' is slightly less clean and suffers from minor aliasing/overlapping with the flag.
- − The textures are more simplified, leaning heavily into the cartoon style at the expense of 'PBR materials'.
Verdict: Qwen Image 2.0 produced a high-quality photographic image that looks professional but ignored the specific aesthetic request for a '3D cartoon miniature'. Seedream 4.0 followed the prompt's stylistic and structural instructions much more accurately, delivering the exact isometric diorama look and layout requested, despite slightly lower text refinement.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent anatomical details, especially in the paws and facial features of the animals.
- + Sophisticated lighting with soft god rays and very clear fur textures.
- + Well-balanced composition with an interaction that feels grounded and realistic.
- − The fox is slightly buried under the other animals, making its face harder to see.
- − The butterfly on the kitten's ear looks somewhat static/pasted on.
Seedream 4.0
- + Whimsical and energetic composition that perfectly matches the 'tumbling/chasing' part of the prompt.
- + Beautiful interpretation of 'dew sparkles' using bokeh-like light points.
- + Vibrant color palette that enhances the joyful vibe.
- − The anatomical structure of the kitten's body and the fox's legs is slightly distorted.
- − The butterfly wings have some merging artifacts with the background.
Verdict: Qwen Image 2.0 provides a more photorealistic image with superior anatomical accuracy and fine fur detail. Seedream 4.0 captures the playful energy and the requested 'sparkles' much better, but falls behind on the physical coherence of the animals. Qwen Image 2.0 is the winner for its masterclass handling of light and realism while still adhering to all prompt elements.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent typography rendering with the correct accent mark.
- + Clean vector-style execution with a unique steam interpretation.
- + Strong adherence to the requested warm brown and cream color palette.
- − The placement of the brand name inside the dome is a bit unconventional for a logo.
- − The banner beneath the dome feels slightly disconnected and bulky.
Seedream 4.0
- + Classic logo composition with well-balanced text arcs.
- + High-quality paper-like background texture adds to the vintage feel.
- + Accurate representation of the cloche dome and steam wisps.
- − The accent mark on 'Caffé' is slanting the wrong way (should be grave instead of acute).
- − The cloche handle is slightly off-center compared to the steam position.
Verdict: Both models followed the prompt closely, providing high-quality vintage emblems. Qwen Image 2.0 produced very crisp text with proper grammar, but Seedream 4.0 had a more traditional and aesthetically pleasing logo composition, despite a minor error in the accent mark direction.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
Qwen Image 2.0
- + Excellent adherence to the vertical poster layout and clean infographic style.
- + Correctly represents all 6 requested steps with appropriate icons.
- + Strong typography with almost perfect spelling, including the astronaut names and 'Tranquility'.
- − Minor typo in 'Translunjar'.
- − The icons for Earth and Lunar orbit are slightly overlapping/cluttered.
Seedream 4.0
- + Effective use of the requested NASA color palette (navy, red, white).
- + Clear iconography for the trajectory and orbit rings.
- − Failed to include 6 distinct steps, merging 'Descent Surface' as step 5.
- − Significant spelling errors such as 'Surfcce' and 'Translunar' is cut off.
- − Awkward composition with the oversized 'Landing 6' text and generic humanoid silhouettes.
Verdict: Qwen Image 2.0 followed the prompt's logical structure much more effectively, providing all six requested steps in a sophisticated vertical layout that feels like a real infographic. Seedream 4.0 struggled with the numbering, spelling, and general composition, resulting in a cluttered design that missed several key details.
Explore each model
ByteDance's image generation model with integrated text-to-image and image editing capabilities in a unified architecture, supporting up to 4K resolution