Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 [schnell] Black Forest Labs Qwen Image 2.0 Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

FLUX.1 [schnell]

19.2 arena score

#45 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image 2.0

21.7 arena score

#34 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 [schnell]

0%

win rate

Ties

0%

Qwen Image 2.0

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 [schnell]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell]

  • + Highly vibrant colors and crisp details
  • + Excellent lighting effects with soft window light
  • + Beautifully rendered monstera plant in the background
  • Added an extra blue sphere on top of the book which was not requested
  • The blue sphere inside the cube appears to be floating unnaturally

Qwen Image 2.0

  • + Accurately followed the spatial instructions for all objects
  • + Realistic glass textures and reflections
  • + Photorealistic rendering of the wooden table and red book
  • The glass cube construction looks more like a frame than a solid cube
  • Internal reflections of the blue sphere are slightly confusing/messy

Verdict: Qwen Image 2.0 is the winner because it adhered strictly to the prompt requirements, whereas FLUX.1 [schnell] hallucinated an additional blue sphere on top of the book. Qwen Image 2.0 also achieved a higher level of photographic realism, particularly in the texture of the book and the lighting consistency.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.1 [schnell]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent depiction of cinematic lighting and reflections
  • + High technical resolution and clarity
  • + Good sense of street atmosphere with bokeh and traffic
  • The 'repair' action is unrealistic as he is just holding the handlebars
  • The man's proportions feel slightly off compared to the bike size
  • Missing the requested motion blur on the background cars

Qwen Image 2.0

  • + Successfully captured the requested 'imperfect framing' and candid feel
  • + Accurately depicts the act of repairing (working on the pedal/chain)
  • + Includes motion blur on the passing car as requested
  • Texture on the man's hands is slightly inconsistent/muddied
  • Reflections on the pavement are less detailed than in Model A
  • The background figure is abruptly cut off in a distracting way

Verdict: Model B (Qwen Image 2.0) better followed the technical requirements of the prompt, specifically incorporating the 'imperfect framing', the 'motion blur', and a more realistic depiction of the man actually repairing the bike. While Model A (FLUX.1 [schnell]) is more aesthetically pleasing and has cleaner textures, it missed several specific descriptors and the action feels posed rather than candid.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 [schnell]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell]

  • + Extremely high-fidelity skin texture and pore detail
  • + Intense, lifelike gaze with sharp focus on the eyes
  • + Dramatic lighting that highlights facial structure and armor engraving
  • Missed the request for small beads in the braided hair
  • Did not include the bokeh sparks mentioned in the prompt
  • Composition is a bit too tight, losing the texture of the leather straps

Qwen Image 2.0

  • + Excellent adherence to specific details like hair beads and bokeh sparks
  • + Realistically depicts the 'battle-worn' aspect with prominent scars and dirt
  • + Superior interpretation of the ornate armor and cloth underlayer layering
  • Skin texture appears slightly over-processed and plastic in some areas
  • The hand anatomy on the sword hilt is slightly awkward
  • The lighting on the armor reflects a generic fire rather than specific torchlight

Verdict: Qwen Image 2.0 followed the prompt much more closely, including the beads in the hair and the bokeh sparks which FLUX.1 [schnell] ignored. While FLUX.1 [schnell] has more impressive photorealistic skin textures, Qwen Image 2.0 captured the specific 'battle-worn paladin' aesthetic and overall composition much more effectively.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 [schnell]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell]

  • + Successfully incorporates all requested menu sections: Appetizers, Pizza, and Mains.
  • + The layout closely resembles a traditional printed paper menu with hierarchy and bulleted descriptions.
  • + Text is highly legible, especially the main headings and titles.
  • The food photos are repetitive and lacks variety in the food subjects.
  • The body text is largely gibberish placeholder text.
  • The 'Orfefus' heading is a hallucination that doesn't fit the prompt.

Qwen Image 2.0

  • + Features very high-quality, vivid, and varied food photography that looks appetizing.
  • + Excellent use of the 'grid' layout requested in the prompt.
  • + The 'Mains' section correctly depicts main courses like steak, chicken, and pasta.
  • The text rendering is poor, with many characters appearing as distorted symbols or gibberish.
  • The layout lacks traditional menu descriptions, making it look more like a digital gallery than a functional menu.
  • Prices and item names are repetitive or illegible.

Verdict: FLUX.1 [schnell] creates a more realistic and usable menu layout that better follows the structural requirements of the prompt, including specific headings and a logical hierarchy. However, Qwen Image 2.0 provides significantly better food photography and grid composition, even though its text rendering is messy and it lacks the requested text-heavy sections. FLUX.1 [schnell] is the winner for better prompt adherence to the 'menu' concept.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 [schnell]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell]

  • + The lighting on the burger and sauce drippings is high quality and photorealistic.
  • + The background flame effects and embers are well-executed.
  • Major spelling error in the title showing 'AGIC BURGER'.
  • Fails to follow the 'exploded burger' prompt, showing a mostly assembled burger surrounded by crouton-like debris.
  • The price is repeated incorrectly and includes a confusing starburst design.

Qwen Image 2.0

  • + Perfect text rendering for all requested phrases with the specific fiery glowing effect.
  • + Includes accurate currency formatting and a well-integrated fiery starburst.
  • + Better captures the 'exploded' aspect with visible separation between the bun and patty.
  • The 'LIMITED TIME ONLY' text is slightly small and less prominent.
  • The background fire looks a bit more like a flat graphic overlay compared to the depth in the other image.

Verdict: Qwen Image 2.0 followed all instructions perfectly, including complex text requirements and the specific fiery aesthetic requested. FLUX.1 [schnell] failed significantly on text accuracy and the core concept of an exploded burger layout, instead producing a mostly intact burger with nonsensical text.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 [schnell]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell]

  • + The text has a clear chalk-like stroke quality.
  • + Good overall image clarity and lighting.
  • Significant spelling errors and gibberish throughout the menu items.
  • Failed to follow the specific date format and title requirements.
  • Text looks more like a digital marker than actual chalk texture at times.

Qwen Image 2.0

  • + Excellent prompt adherence with near-perfect spelling of all requested menu items.
  • + Realistic chalk texture including smudges and authentic handwriting variations.
  • + Successfully captured the requested date and cursive title style.
  • The 'Today's Specials' title is slightly cramped against the top edge.
  • Handwriting overlaps with some of the chalk smudges, making small parts slightly less legible.

Verdict: Qwen Image 2.0 significantly outperformed FLUX.1 [schnell] by following the text prompts accurately and rendering complex culinary terms without spelling errors. Qwen Image 2.0 also captured a much more authentic 'chalkboard' aesthetic with realistic smudges and varied handwriting, whereas FLUX.1 produced mostly incomprehensible text.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 [schnell]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell]

  • + Perfect adherence to the surreal instruction of the horse being on top.
  • + Cinematic lighting and high-quality rendering of the space background.
  • + Creative interpretation of an astronaut acting as a mount.
  • The horse has two heads/torsos merging, a significant anatomical artifact.
  • The astronaut's suit geometry is a bit nonsensical and distorted.

Qwen Image 2.0

  • + Crisp details on the horse's coat and the astronaut's suit.
  • + Good sense of motion and scale with the planet below.
  • Failed the negative constraint; the astronaut is on top of the horse.
  • The horse's front legs have anatomical issues and strange scaly textures.

Verdict: FLUX.1 [schnell] is the clear winner for its superior prompt adherence, successfully placing the horse on top of the astronaut as requested by the surreal prompt. While Qwen Image 2.0 has high clarity, it completely ignored the specific spatial instruction and produced a standard 'astronaut on horse' image.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 [schnell]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent fur texture and lighting on the capybara.
  • + Logical rendering of the human passenger looking at her phone.
  • + High-quality blurred background that feels cinematic.
  • The capybara only has one paw on the steering wheel, failing the prompt's specific instruction.
  • The cap is a simple beanie style rather than a traditional driver cap.

Qwen Image 2.0

  • + Followed the specific instruction for 'both front paws on the steering wheel'.
  • + The driver cap is more accurate to a traditional taxi uniform style.
  • + Realistic perspective through the car window with city light reflections.
  • The human passenger appears to be in the front passenger seat rather than the back seat.
  • The hands of the capybara have some slight anatomical irregularities.

Verdict: Both models captured the surreal nature of the prompt well, but FLUX.1 [schnell] produced a more polished and photorealistic image with better overall character expressions. Although Qwen Image 2.0 followed the 'both paws' instruction and provided a better cap, it failed to place the passenger in the back seat, making FLUX.1 [schnell] the more visually pleasing and coherent choice.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 [schnell]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell]

  • + Strong cinematic lighting with a vibrant glowing jack-o-lantern.
  • + Captures the moody night sky and thorny border well.
  • Several severe spelling errors and repeated text lines in the body.
  • Formatting of the event details is messy and nonsensical.
  • The parchment texture is less apparent compared to the other model.

Qwen Image 2.0

  • + Perfect text rendering for all requested fields including specific event details.
  • + Excellent vintage gothic parchment aesthetic with high-quality borders.
  • + Clear and balanced composition that feels like a real professional invitation.
  • The jack-o-lantern is a bit more generic in lighting style.
  • The scroll banner is slightly less integrated into the background.

Verdict: Qwen Image 2.0 followed the prompt instructions near-perfectly, successfully rendering complex text instructions without a single spelling error. While FLUX.1 [schnell] had more atmospheric lighting, it failed significantly on text coherence and layout, resulting in a cluttered invitation with several gibberish words.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 [schnell]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell]

  • + Perfectly captures the isometric miniature diorama aesthetic
  • + Extremely clean 3D render style with soft lighting
  • + Accurate placement of the Japanese flag icon
  • Missed the 'SUSHI' text requirement entirely
  • The sushi piece looks more like a 2D illustration on a 3D base rather than a fully 3D object

Qwen Image 2.0

  • + Followed all text instructions including 'JAPAN' and 'SUSHI'
  • + High-quality realistic PBR textures on the food items
  • + Good variety of sushi types on the plate
  • Failed the isometric diorama perspective, providing a standard high-angle photo style
  • The flag icon is placed to the right instead of centered as requested
  • The base is a simple wooden board rather than the requested 'small raised diorama base' style

Verdict: FLUX.1 [schnell] excelled at the stylistic isometric diorama request but failed to include the word 'SUSHI'. Qwen Image 2.0 followed all textual prompts correctly, but missed the specific 'isometric miniature' and 'diorama' art style, opting for a more realistic food photography look.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 [schnell]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent fur texture and lighting on the animals.
  • + Vibrant, warm colors that match the sunrise theme.
  • + Clean, professional composition with good depth of field.
  • Failed to include a distinct bunny; instead showed two kittens and a fox.
  • The positioning is static and missing the 'tumbling' action requested.

Qwen Image 2.0

  • + Stronger adherence to the 'tumbling' and 'chasing' actions described.
  • + Clearly represents all four requested species: puppy, kitten, bunny, and fox.
  • + Beautiful god rays and atmospheric lighting.
  • The fox kit's anatomy looks a bit awkward while on its back.
  • The kitten's size is slightly inconsistent relative to the puppy.

Verdict: While FLUX.1 [schnell] produced a more polished and visually striking image in terms of texture, it failed significantly on prompt adherence by missing the bunny and providing two kittens instead. Qwen Image 2.0 followed the complex prompt much more accurately, capturing the specific animals and the playful movement requested, despite slightly less refined animal anatomy.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 [schnell]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell]

  • + Clean vector emblem style
  • + Well-balanced circular composition
  • Serious spelling errors: 'Framilan' and '7720'
  • Missing the 'steam' element requested in the prompt

Qwen Image 2.0

  • + Perfect text rendering for name and date
  • + Includes the cloche, steam, and banner as requested
  • + Detailed shading and subtle paper texture
  • Logo is slightly less minimalist than Image A
  • The cloche handle is slightly asymmetric

Verdict: Qwen Image 2.0 is the clear winner as it followed every part of the prompt, including the correct spelling of 'Caffè Florian' and the 'Est. 1720' date. FLUX.1 [schnell] produced major hallucinations with the text and failed to include the steam element.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.1 [schnell]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent adherence to the 'flat-vector' and clean iconography style.
  • + Beautiful color palette that feels professional and NASA-inspired.
  • + Sophisticated composition with a centered, balanced layout.
  • Text is largely illegible gibberish.
  • The 'Saturn V' icon looks more like a generic toy rocket or shuttle.

Qwen Image 2.0

  • + Excellent adherence to all six requested steps with legible labels.
  • + Highly accurate icons for the Lunar Module and Saturn V.
  • + Strong supporting details like the crew names and 'Tranquility' marker.
  • The layout is a bit cramped with overlapping elements in 'Lunar Orbit'.
  • The white border around the central text is slightly less 'modern' than the requested flat style.

Verdict: Qwen Image 2.0 is the clear winner for its superior adherence to the informational requirements of the prompt, successfully illustrating all six steps with legible, accurate text and recognizable lunar hardware. While FLUX.1 [schnell] creates a more aesthetically pleasing 'modern' art piece, it fails to deliver a functional infographic and produces nonsensical text for its labels.

Next steps

Explore each model