Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 [schnell] Black Forest Labs Qwen Image Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

FLUX.1 [schnell]

18.7 arena score

#48 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image

20.3 arena score

#40 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 [schnell]

0%

win rate

Ties

0%

Qwen Image

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 [schnell]
Qwen Image

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent photorealism and texture detail.
  • + Vibrant colors and high resolution.
  • + Effective lighting following the prompt's direction.
  • Failed the spatial logic by placing an extra blue marble on top of the book.
  • The sphere inside is floating mid-air incongruously.

Qwen Image

  • + Perfect adherence to all spatial instructions and object counts.
  • + Realistic physics with the ball resting on the bottom of the cube.
  • + Clean composition with natural depth of field.
  • Slightly lower sharpness on the plant in the background.
  • Reflection on the bottom of the cube looks a bit like a double-layered floor.

Verdict: While FLUX.1 [schnell] has slightly higher textural fidelity, it failed the specific instructions by adding an extra blue sphere on top of the red book. Qwen Image followed every part of the prompt perfectly, including the precise placement and count of items, making it the superior choice for prompt adherence.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.1 [schnell]
Qwen Image

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent street atmosphere with realistic lighting and reflections
  • + Very high detail in skin texture and clothing materials
  • + Effective use of cinematic lighting on the wet pavement
  • Failed to include the requested motion blur on passing cars
  • Hands interacting with the handlebars are blurry and anatomically confused

Qwen Image

  • + Successfully captured the requested motion blur on the background car
  • + Better captures the feeling of 'light rain' through visual streaks
  • + Good candid posture for the subject
  • The red bicycle frame is physically broken/detached where the seat post meets the pedals
  • The background car has a warped, surreal appearance
  • Overall image clarity is lower than Model A

Verdict: Both models struggled with the complex physical interaction between the man and the bike. FLUX.1 [schnell] produced a much cleaner and more professional-looking image with superior textures, but it ignored the 'motion blur' requirement. Qwen-VL-Plus followed the prompt details more accurately, including the motion blur and rain streaks, but suffered from significant structural failures in the geometry of the bicycle and the car.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 [schnell]
Qwen Image

AI Judge Analysis

FLUX.1 [schnell]

  • + Intense and lifelike facial features and eyes
  • + Excellent skin texture and subtle realistic scarring
  • + Strong, cinematic use of shallow depth of field
  • Missed the request for small beads in the hair
  • Portrait is cropped too closely to see the 'ornate plate armor' and leather straps clearly

Qwen Image

  • + Excellent adherence to all prompt elements including beads, leather straps, and armor
  • + Clearly visible ornate engravings and battle damage on the plate
  • + Beautiful lighting and sparks from the torch
  • The sparks look a bit like digital stickers/stars rather than natural bokeh particles
  • Face lacks the high-fidelity skin realism found in the competitor

Verdict: While FLUX.1 [schnell] captures a more intense and realistic face, Qwen Image is the superior choice for prompt adherence, successfully incorporating the beads, leather straps, and full armor set requested. Qwen Image provides a more complete storytelling composition that aligns better with the 'battle-worn paladin' archetype.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 [schnell]
Qwen Image

AI Judge Analysis

FLUX.1 [schnell]

  • + Features a very clean and professional white space balance.
  • + Includes the requested sections for Appetizers, Pizza, and Mains.
  • + Text rendering of main headers is very legible.
  • The grid of food photos has inconsistent spacing and is not perfectly aligned.
  • Small body text is mostly gibberish.
  • The 'ORFEFUS' header shows some hallucination or spelling error.

Qwen Image

  • + Stronger adherence to the 'grid' request with colorful, well-defined tiles.
  • + Excellent use of bold sans-serif fonts and vibrant accents as requested.
  • + The layout feels more like a modern casual dining brand.
  • The pizza slice in the middle tile is cut off unnaturally at the top.
  • Headers 'Appetizers/' and 'Pizza/mains' contain slashes and slightly messy character rendering.
  • Combined 'Pizza/mains' into one section instead of separating them as requested.

Verdict: FLUX.1 [schnell] creates a more traditional, professional menu layout with excellent whitespace management, though it struggles with the grid alignment. Qwen Image provides a more modern and vibrant interpretation with a clear grid system and better color use, but has more artifacts in the font rendering and layout logic. Overall, FLUX.1 [schnell] feels more like a usable professional template.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 [schnell]
Qwen Image

AI Judge Analysis

FLUX.1 [schnell]

  • + The lighting on the burger and melting cheese is highly realistic.
  • + The background fire and embers look visceral and dynamic.
  • The main text is misspelled as 'AGIC BURGER'.
  • The pricing is cluttered, repetitive, and the starburst is aesthetically poor.
  • The burger is largely assembled rather than 'exploded' into mid-air components.

Qwen Image

  • + Perfectly adhered to the text requirements with correct spelling and fiery glowing effects.
  • + Successfully depicted the 'exploded' burger concept with individual ingredients suspended.
  • + The composition is professional, resembling a real promotional advertisement.
  • The burger ingredients look slightly more illustrative compared to the grit of Model A's textures.
  • Lower resolution in some of the background fire details.

Verdict: Qwen Image is the clear winner because it correctly followed all complex text instructions, including spelling and specific price placement in a starburst, whereas FLUX.1 [schnell] failed on both the spelling and the 'exploded' nature of the burger. Qwen Image also captured the 'fiery glowing effect' requested for the typography much better than its competitor.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 [schnell]
Qwen Image

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent chalk texture on the board and lettering
  • + Realistic handwriting style with natural variation
  • Significant spelling errors throughout almost every word
  • Failed to include 'April' correctly, writing 'Pril' instead

Qwen Image

  • + Much better spelling and legibility of the menu items
  • + Good composition with a clear cafe background
  • + Followed the pricing instructions accurately
  • Hallucinated the year as '20026' instead of '2026'
  • Text looks slightly like a digital font overlay rather than genuine chalk strokes
  • Missed the specific cursive requirement for the title

Verdict: While FLUX.1 [schnell] captures the physical texture of chalk and the requested messy handwriting style more authentically, it fails significantly at providing readable or correctly spelled text. Qwen Image provides a much more functional and legible menu, despite the error in the year and a slightly more 'digital' font appearance. Qwen Image is the winner for successfully rendering the complex menu items requested in the prompt.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 [schnell]
Qwen Image

AI Judge Analysis

FLUX.1 [schnell]

  • + Perfectly follows the specific instruction to have the horse on top of the astronaut
  • + Strong cinematic lighting and surreal atmosphere
  • + Higher overall image detail and resolution
  • The horse has two heads/necks, which is a significant anatomical error
  • The astronaut's anatomy is slightly fragmented

Qwen Image

  • + Anatomically correct horse and astronaut
  • + Clear, high-contrast composition
  • Completely failed the negative constraint to have the horse on top
  • Lacks the 'surreal' quality requested in the prompt beyond the setting

Verdict: FLUX.1 [schnell] is the clear winner because it successfully interpreted the difficult spatial instruction to have the horse on top of the astronaut, whereas Qwen Image defaulted to the standard trope. While FLUX.1 [schnell] suffered from a multi-headed horse artifact, its adherence to the core prompt concept makes it a much better result for the specific challenge compared to the generic output of Qwen Image.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 [schnell]
Qwen Image

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent fur texture rendering
  • + Clear legible text on the driver hat
  • + Captures the bored expression of the passenger perfectly
  • Paws are not positioned correctly on the steering wheel
  • The capybara's head is looking at the camera rather than the road
  • Light bokeh in the background is a bit generic

Qwen Image

  • + Natural profile view of the driver focusing on the road
  • + Hands/paws are correctly placed on the steering wheel
  • + Very convincing cinematic lighting and composition
  • The driver's hands look more like primate hands than capybara paws
  • Text on the car rooftop sign says 'YOXI' instead of 'TAXI'
  • The passenger's phone is held at a slightly awkward angle

Verdict: Qwen Image provides a more realistic and cinematically composed scene, placing the capybara in a natural driving position with both hands on the wheel as requested. While FLUX.1 [schnell] has superior texture for the capybara's fur and better text rendering on the hat, the frontal-facing pose feels less like a real driving scene compared to Qwen Image.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 [schnell]
Qwen Image

AI Judge Analysis

FLUX.1 [schnell]

  • + Successfully captures a moody night sky and atmospheric trees.
  • + Rich orange color palette feels appropriate for Halloween.
  • + High fidelity jack-o-lantern rendering.
  • Significant text hallucinations and repetition in the event details.
  • Failed to follow the 'dark parchment' texture requirement, opting for a digital gradient.
  • Grammatical errors in the scroll banner text.

Qwen Image

  • + Excellent adherence to the 'dark parchment', thorns, and webs border request.
  • + Near-perfect text accuracy on the event details and scroll banner.
  • + Better compositional balance with the text and central imagery.
  • Missing the word 'Halloween' in the main title text (reads 'Hallo Party').
  • The border thorns are a bit repetitive and mechanical.

Verdict: Qwen Image is the superior choice because it accurately captures the 'parchment' and 'border' stylistic elements while maintaining clear, legible event information. Although it missed a segment of the title text, FLUX.1 [schnell] suffered from significant text repetition and nonsensical formatting at the bottom of the poster which makes it unusable as an invitation.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 [schnell]
Qwen Image

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent high-clarity PBR materials and realistic textures
  • + Perfectly followed the 45-degree isometric diorama instruction
  • + Extremely clean and professional composition
  • Missed the 'SUSHI' text requirement at the top-center
  • The salmon texture has slightly unusual red markings/artifacts

Qwen Image

  • + Included all text elements: 'JAPAN', 'SUSHI', and flag icon
  • + Captures the 'miniature 3D cartoon scene' aesthetic very well
  • + Good use of diorama elements like chopsticks and a flag
  • The camera angle is more of a portrait view rather than the requested 45-degree top-down isometric view
  • Text rendering is slightly cramped and slightly off-center

Verdict: FLUX.1 [schnell] produced a much higher quality image in terms of technical rendering and isometric perspective, but it failed to include all the requested text. Qwen Image followed the text prompts more accurately and captured the 'cartoon' vibe well, but failed to deliver the specific isometric angle requested. FLUX.1 [schnell] is likely the preferred choice for its superior clarity and adherence to the layout constraints.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 [schnell]
Qwen Image

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent fur texture and lighting integration
  • + Vibrant colors and a very clean, high-resolution aesthetic
  • Failed to include a rabbit, instead generating two cat-like hybrid creatures
  • More of a static pose rather than the requested 'chasing' action

Qwen Image

  • + Successfully included all four specific animals: puppy, kitten, bunny, and fox
  • + Strong adherence to the 'god rays' and 'dew sparkles' requirements
  • + Captured the sense of movement and 'tumbling' much better
  • The fox kit has unusual dark, paw-like features on its forelimbs that look slightly anatomicaly incorrect
  • Slightly more 'AI-smooth' look compared to the sharper textures in Model A

Verdict: While FLUX.1 [schnell] produced a higher-fidelity image with more convincing fur textures, it failed the core logical task by omitting the bunny and doubling up on feline features. Qwen Image followed the complex multi-subject prompt perfectly, correctly identifying all four animal types and better capturing the dynamic 'chasing' interaction and specific atmospheric lighting requested.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 [schnell]
Qwen Image

AI Judge Analysis

FLUX.1 [schnell]

  • + Excellent typography style and layout for a vintage emblem.
  • + Professional vector aesthetic with a clean, centered composition.
  • + Accurate color palette and subtle parchment texture.
  • Failed to render the correct name, showing 'CAFEÉ FRAMILAN' instead of 'Caffè Florian'.
  • Misspelled the date as '7720' instead of '1720'.
  • Missing the steam element requested in the prompt.

Qwen Image

  • + Included the steam element above the cloche dome.
  • + Correctly rendered the date 'EST. 1720' as requested.
  • + Used the correct brand name, despite some messy overlapping characters.
  • The typography is poorly constructed with overlapping and jumbled letters.
  • The cloche is less refined and looks more like a basic icon than a professional logo.
  • The banner design is slightly chunky compared to the 'minimalist' request.

Verdict: This is a trade-off between technical execution and instruction following. FLUX.1 [schnell] produced a much more beautiful and professional-looking logo, but failed significantly on the specific text and date. Qwen Image followed the prompt instructions and text content much more accurately, but the typographical 'FLoraiAN' overlap is a major visual flaw. FLUX.1 [schnell] is likely the winner for a designer who can easily edit text, but Qwen Image is the winner for raw prompt adherence.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.1 [schnell]
Qwen Image

AI Judge Analysis

FLUX.1 [schnell]

  • + Strong aesthetic alignment with NASA-inspired color palette and flat-vector minimalism.
  • + Excellent graphic composition that feels like a professional poster design.
  • + Consistent iconography and clean, crisp lines throughout.
  • Text is largely illegible/nonsensical gibberish.
  • Fails to clearly delineate the 6 specific steps requested in the prompt.

Qwen Image

  • + Successfully incorporates readable text and specific names requested in the supporting details.
  • + Clearly labels specific steps (Launch, Orbit, etc.) making it functional as an infographic.
  • + Includes specific icons like the Lunar Module and astronaut silhouettes mentioned in the prompt.
  • Literal and awkward interpretation of the prompt, including 'Stop at landing' in the actual text.
  • The layout is cluttered and less balanced compared to a professional vector poster.
  • Several spelling errors ('Sarth', 'Aldlin', 'Tranquility') throughout the design.

Verdict: FLUX.1 [schnell] creates a much more visually appealing and professional-looking graphic that nails the flat-vector aesthetic, but it fails completely at conveying information through text. Qwen Image adheres much better to the specific step-by-step instructions and logical flow of an infographic, though it lacks the design polish of the other model and includes silly literal transcriptions of the prompt text. Qwen Image is the likely winner for actually attempting to fulfill the multi-step information requirements of the prompt.

Next steps

Explore each model