Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 [schnell] FP8 Black Forest Labs Qwen Image 2.0 Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

FLUX.1 [schnell] FP8

19.0 arena score

#48 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image 2.0

21.7 arena score

#34 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 [schnell] FP8

50.0%

win rate

Ties

0.0%

Qwen Image 2.0

50.0%

win rate

50.0% 0.0% ties 50.0%
Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 [schnell] FP8
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent photographic quality and realistic textures on the glass and wood.
  • + Accurate interpretation of soft window lighting from the left.
  • + Clean, modern composition with vibrant colors.
  • The glass object is a tall rectangular prism rather than a 'cube' as requested.
  • The sphere appears to be resting on an internal shelf rather than centered or floating inside the cube.

Qwen Image 2.0

  • + Perfectly adheres to the 'cube' shape for the glass container.
  • + Highly realistic texture on the red book cover and wooden table.
  • + Cleverly shows the green plant's distortion and visibility through the glass.
  • The blue sphere is floating without support, which may look slightly unnatural in a realistic scene.

Verdict: While FLUX.1 [schnell] FP8 produces a more aesthetically pleasing and luminous image, it fails to generate a cube, instead creating a tall prism. Qwen Image 2.0 followed the prompt geometry more precisely, correctly generating a cube and showing the plant behind the glass better, despite the slightly surreal floating effect of the sphere.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.1 [schnell] FP8
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent handling of wet pavement reflections and cinematic lighting.
  • + High overall image clarity and vibrant colors.
  • + Captures the mood of a busy city street effectively.
  • Misses the 'motion blur' request for passing cars, as they appear static.
  • The man is holding the handlebars rather than performing a repair action.
  • Skin texture and facial features look slightly smoothed and less 'natural' than requested.

Qwen Image 2.0

  • + Excellent adherence to the 'repairing' action, showing the man working on the pedal/chain.
  • + Superior skin texture and hyper-realistic facial details reflecting age.
  • + Better realization of the 'candid' look with imperfect framing and visible motion blur on the background car.
  • The composition is a bit cramped at the top of the bike.
  • The background lighting is less cinematic compared to Model A.

Verdict: Qwen Image 2.0 followed the specific nuances of the prompt much more accurately, particularly regarding the 'repairing' action and the 'motion blur' of passing vehicles. While FLUX.1 [schnell] FP8 produced a more polished, cinematic aesthetic, Qwen's superior skin texture and authentic candid feel better align with the 'no stylization' and 'natural' requirements.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 [schnell] FP8
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Extremely high skin texture and facial detail
  • + Dramatic warm lighting that accentuates the armor engraving
  • + Intense, lifelike eye rendering
  • Missed the request for hair braids with beads
  • Fails to show a clear 'cloth underlayer' or 'leather straps' as effectively as Model B
  • Less emphasis on 'battle-worn' scars compared to Model B

Qwen Image 2.0

  • + Excellent adherence to all prompt elements including braids with beads, scars, and leather straps
  • + Clearly visible ornate engraved plate armor with a cloth underlayer
  • + Great atmosphere with bokeh sparks and fire representing the torchlight source
  • The hand has anatomical issues with twisted fingers
  • Eyes appear slightly less lifelike and more glazed than Model A
  • Resolution in some fine skin areas feels a bit more smoothed out than Model A

Verdict: While FLUX.1 [schnell] FP8 offers superior facial realism and more dramatic lighting, it completely ignored the instruction for braided hair with beads. Qwen Image 2.0 followed every specific detail of the prompt, including the complex armor assembly and specific hairstyling, despite some minor anatomical issues with the hand.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 [schnell] FP8
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent structure that mimics a real physical menu with detailed price lists and descriptors.
  • + Clean white minimalism adhering well to the 'modern minimalist' prompt.
  • + High-quality sans-serif font choices for headers.
  • Text is largely gibberish or misspelled ('APPTIZERS', 'PIZZAL').
  • Photos are somewhat small and repetitive in composition.

Qwen Image 2.0

  • + Stunning, high-quality food photography with vibrant colors.
  • + Strict adherence to the 'grid' layout requested in the prompt.
  • + Clear and legible section headers for Appetizers, Pizza, and Mains.
  • The pricing and item names contain significant text artifacts and garbled characters.
  • Layout is a bit more 'Instagram feed' than a functional restaurant menu.

Verdict: FLUX.1 [schnell] FP8 does a better job of creating a layout that looks like an actual menu with lists of items and prices, though it suffers from spelling errors in the headers. Qwen Image 2.0 produces much more appetizing and realistic food photography in a clean grid, but the text beneath the photos is almost entirely unreadable artifacts. Qwen Image 2.0 is the visual winner for the quality of the 'vibrant food photos,' while FLUX.1 better understands the structural intent of a multi-section menu.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 [schnell] FP8
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Clean layout with high-contrast text.
  • + Appetizing food rendering and vibrant colors.
  • Numerous spelling errors including 'LIIMITED' and 'NEEY'.
  • Incorrect price rendering, missing the '6' for €69.
  • Fails the 'exploded' requirement as the burger is mostly assembled.

Qwen Image 2.0

  • + Perfect adherence to text requirements with 100% accurate spelling.
  • + Excellent 'fiery, glowing effect' on the typography as requested.
  • + Better 'exploded' composition with the top bun and ingredients clearly separated.
  • The starburst for the price is a bit cluttered.
  • The sauce texture on the meat is slightly messy compared to Model A.

Verdict: Qwen Image 2.0 outperformed FLUX.1 [schnell] FP8 by following every complex text and stylistic instruction perfectly, whereas FLUX.1 failed several spelling checks and price accuracy. Qwen Image 2.0 also achieved a better sense of motion and the requested 'exploded' effect while maintaining a high level of photorealism.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 [schnell] FP8
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Features a clean, centered composition with a sharp wooden frame.
  • + Includes the requested date accurately in the title.
  • Significant spelling errors and repetitive text throughout the menu items.
  • The title font is not cursive as requested.
  • The handwriting looks more like a digital marker than realistic chalk.

Qwen Image 2.0

  • + Excellent adherence to the menu text with perfect spelling.
  • + Highly realistic chalk texture including smudges and authentic handwriting variations.
  • + Captures the 'cozy café' atmosphere with superior lighting and background depth.
  • The board is slightly cut off at the edges compared to Model A.

Verdict: Qwen Image 2.0 is the clear winner as it followed every detail of the prompt, including the specific menu items, prices, and the request for a cursive title. In contrast, FLUX.1 [schnell] FP8 suffered from significant text hallucinations and failed to render the items accurately.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 [schnell] FP8
Qwen Image 2.0
100% wins 0% ties 0% wins

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent adherence to the 'horse on top' spatial instruction
  • + Cinematic composition with great lighting and scale
  • + Creative interpretation of equipment integrated with the horse body
  • Anatomical fusion is slightly messy where the second horse head emerges
  • Includes a strange architectural fragment at the bottom right

Qwen Image 2.0

  • + High visual clarity and sharp detailing
  • + Clean representation of an astronaut and a horse in space
  • + Pleasant color palette and atmosphere
  • Failed the primary prompt instruction of having the horse 'on top'
  • Standard cliché composition instead of the requested surreal reversal

Verdict: FLUX.1 [schnell] FP8 followed the specific and difficult instruction to place the horse on top of the astronaut (effectively making the horse the rider), whereas Qwen Image 2.0 ignored this and produced a standard astronaut riding a horse. While Qwen's image is cleaner, FLUX.1 succeeded in the core creative challenge of the prompt.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 [schnell] FP8
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent texture on the capybara's fur
  • + Clean and readable text on the taxi hat
  • + Good composition with the back-seat passenger clearly visible
  • The passenger is holding two phones simultaneously
  • The capybara's eyes and snout are somewhat distorted or looking at the camera rather than the road
  • The capybara is not sitting in a realistic position relative to the steering wheel

Qwen Image 2.0

  • + Highly realistic 'candid' photographic lighting and textures
  • + The capybara's anatomy and paw placement on the wheel are very believable
  • + Natural human expression and behavior in the background
  • The taxi driver hat's brim is slightly clipped and has no text
  • The passenger and driver are sitting very close together, making the interior feel cramped for a taxi

Verdict: Qwen Image 2.0 (Model B) wins due to its superior photorealistic aesthetic and better anatomical integration of the capybara with the car controls. FLUX.1 [schnell] FP8 (Model A) produces a more stylized, clean image but suffers from logic errors like the passenger holding two phones and a less natural posture for the driver.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 [schnell] FP8
Qwen Image 2.0
0% wins 0% ties 100% wins

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Features a visually striking central glowing pumpkin
  • + Uses a thematic dark aesthetic suitable for gothic themes
  • Numerous spelling errors in the banner and title text
  • Details like date and time are incorrect or hallucinated
  • Missing the webs and thorns requested for the border

Qwen Image 2.0

  • + Excellent text rendering with near-perfect spelling and typography
  • + Strong adherence to all prompt elements including thorns, webs, and twisted trees
  • + High visual quality with realistic textures on the parchment and pumpkin
  • The lighting on the pumpkin is slightly less cinematic compared to the overall background depth

Verdict: Qwen Image 2.0 outperformed FLUX.1 [schnell] FP8 by accurately rendering all requested text and specific decorative elements like the spiderweb and thorn border. While FLUX.1 [schnell] FP8 captured the moody atmosphere, it failed significantly on text legibility and prompt details.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 [schnell] FP8
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent 3D isometric diorama aesthetic
  • + Consistent stylized 3D textures and lighting
  • + Matches the requested soft cartoon miniature style
  • Text is garbled and repetitive
  • Missing requested 'SUSHI' text word
  • Flag icon is integrated poorly into the text

Qwen Image 2.0

  • + Perfect text rendering for both 'JAPAN' and 'SUSHI'
  • + Accurate and clear flag icon
  • + Diverse and realistic sushi types presented on the diorama base
  • Lean more toward photo-realism than the requested '3D cartoon' style
  • The blue background is slightly more saturated than the requested 'light blue'

Verdict: FLUX.1 [schnell] FP8 captured the requested 3D cartoon/miniature aesthetic and isometric perspective much more accurately, but failed significantly on the text. Qwen Image 2.0 followed the specific text instructions perfectly but produced a more conventional photograph-style image that missed the intended 'cartoon' art style.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 [schnell] FP8
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Vibrant, expressive lighting and atmosphere
  • + Sharp focus on facial features and eyes
  • + Whimsical, storybook aesthetic
  • Failed to include the baby bunny
  • Repeated kitten subjects instead of diverse animals
  • Anatomical issues with extra or merging limbs on the central animals

Qwen Image 2.0

  • + Perfect prompt adherence including puppy, kitten, bunny, and fox
  • + More realistic animal textures and anatomy
  • + Excellent lighting with visible god rays and dew
  • The butterfly on the right is floating without a clear body
  • Slightly less 'hyper-detailed' fur compared to Model A

Verdict: Qwen Image 2.0 is the clear winner as it successfully included all four requested animals (golden retriever, tabby kitten, bunny, and fox) with high anatomical accuracy, whereas FLUX.1 [schnell] FP8 missed the bunny entirely and had significant clipping/merging issues between the multiple kittens it generated. Qwen Image 2.0 also better captured the 'tumbling together' action and the specific lighting effects requested like god rays.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 [schnell] FP8
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent vintage tone and aesthetic color palette
  • + Good vector emblem composition
  • + Appropriate use of subtle background texture
  • Major spelling errors in the brand name ('AFe FLAMILAN' instead of 'Caffè Florian')
  • The primary illustration looks more like a dome or building than a cloche dome

Qwen Image 2.0

  • + Perfect text rendering for both the brand name and the establishment date
  • + Clean, professional vector illustration that clearly depicts a cloche dome
  • + Accurate interpretation of the steam and banner elements
  • The shading on the cloche is slightly more 3D than a 'minimalist logo' usually requires
  • The steam effect looks a bit like a flame

Verdict: Qwen Image 2.0 is the clear winner because it followed the text instructions perfectly, whereas FLUX.1 [schnell] FP8 failed significantly on the typography and spelling. Qwen Image 2.0 also correctly visualized a restaurant cloche dome, while FLUX.1 generated an architectural-looking shape.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.1 [schnell] FP8
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [schnell] FP8

  • + Excellent primary heading text rendering for 'APOLLO 11'.
  • + Very clean, professional layout that matches the 'modern vector infographic' aesthetics.
  • + Uses the requested NASA-inspired color palette effectively and subtly.
  • Nonsense filler text for all sub-labels.
  • Incorrect iconography; for example, 'Earth Orbit' is a circle with dots and 'Launch' is just a blank circle.
  • Fails to follow the chronological step sequence correctly in the layout.

Qwen Image 2.0

  • + Accurately depicts all six requested steps in the correct chronological order.
  • + Superior iconography that actually represents the objects described (Saturn V, Earth, Moon, Lunar Module).
  • + Perfect English spelling for all technical step labels.
  • Text is slightly heavy with black outlines, clashing with the 'clean/modern' request.
  • Layout is a bit crowded vertically.
  • Styling is a bit more illustrative than the requested 'flat-vector' infographic style.

Verdict: Qwen Image 2.0 is the clear winner because it successfully followed the complex instruction to include six specific steps with relevant icons and accurate text labels. FLUX.1 [schnell] FP8 produced a more aesthetically pleasing 'modern' design, but it failed the core task by using gibberish text and icons that did not correspond to the requested mission phases.

Next steps

Explore each model