Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 [dev] Black Forest Labs Qwen Image 2.0 Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

FLUX.1 [dev]

24.6 arena score

#16 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image 2.0

21.7 arena score

#34 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 [dev]

0%

win rate

Ties

0%

Qwen Image 2.0

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent photographic quality with realistic depth of field.
  • + Very smooth and clean rendering of glass reflections and light.
  • + Precisely follows the lighting direction (left side).
  • The sphere appears to be a light blue rather than a saturated blue.
  • The glass cube is rectangular rather than a perfect cube.

Qwen Image 2.0

  • + Perfect adherence to the object 'sphere' and 'cube' geometry.
  • + High contrast and saturated colors, especially for the book and sphere.
  • + Good reflection logic on the interior glass panels.
  • The glass panels look slightly thin and more like a frame than a solid glass cube.
  • Slightly more grain/noise in the background compared to model_a.

Verdict: Both models followed the prompt perfectly, including the specific spatial relationships of the objects. FLUX.1 [dev] produced a more aesthetically pleasing, high-end photographic look with superior lighting, while Qwen Image 2.0 captured the color saturation and geometric shapes of the internal objects more accurately.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.1 [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent depiction of rain and atmosphere through lighting and reflections
  • + High cinematic quality with smooth bokeh
  • + Accurate representation of an elderly man with naturalistic clothing texture
  • The man is standing and holding the bike rather than actively repairing it
  • The bike design feels slightly generic/modern for the requested 'candid street' aesthetic

Qwen Image 2.0

  • + Stronger adherence to the 'repairing' action with a squatting pose
  • + Excellent 'imperfect framing' that feels like a genuine candid snapshot
  • + Superior natural skin texture and weathered hand details
  • The car in the background lacks the requested motion blur
  • The depth of field is slightly deeper than requested for a 50mm look

Verdict: FLUX.1 [dev] produces a more aesthetically pleasing, cinematic image with beautiful lighting, but it feels staged. Qwen Image 2.0 better captures the 'candid street photo' prompt with a more convincing 'repairing' action and gritty, realistic skin textures, despite missing the motion blur on the background car.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent high-frequency facial textures and lifelike iris details.
  • + Subtle, realistic lighting integration on the metal armor.
  • + Soft, aesthetically pleasing bokeh bubbles.
  • Missed the 'small beads' in the hair braids.
  • The subject looks too clean and youthful for a 'battle-worn' description.
  • Lacks the requested 'ornate' engraving on the armor.

Qwen Image 2.0

  • + Perfect adherence to all prompt details, including scars, dirt, beads, and engravings.
  • + Strong emotional storytelling through a realistic battle-worn expression.
  • + Excellent texture work on leather straps and frayed cloth underlayers.
  • The eyes have an unnatural yellowish/reddish tint that looks slightly edited.
  • The hand on the sword has some anatomical blurring/merging issues.
  • The background fire is slightly overpowering compared to the 'warm torchlight' requested.

Verdict: Qwen Image 2.0 followed the prompt much more accurately, capturing specific details like the beads in the braids, the ornate engravings, and the gritty, battle-worn appearance. While FLUX.1 [dev] produced a cleaner, more high-fidelity portrait, it ignored several key descriptive elements, resulting in a character that looks more like a modern model in armor than a veteran paladin.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent typography with clean, legible header fonts
  • + Realistic menu layout with item descriptions and pricing
  • + Features a cohesive minimalist professional design
  • Failed the request for a 'grid' of food photos
  • Missing a dedicated 'Pizza' section header

Qwen Image 2.0

  • + Strictly followed the request for a grid layout with colorful food photos
  • + Included clear headings for Appetizers, Pizza, and Mains
  • + High-quality food photography within the grid cells
  • Text for individual menu items is garbled and illegible
  • Layout looks more like a photo gallery than a functional menu
  • Pricing and names are repetitious and lack realism

Verdict: FLUX.1 [dev] produced a more believable and professional menu design with high-quality typography, though it missed the specific grid layout instruction. Qwen Image 2.0 followed the structural prompt for a photo grid and headers perfectly, but the text rendering is poorly executed and less realistic for a professional setting.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [dev]

  • + Successfully rendered most of the requested text.
  • + Good vertical composition and separation of elements.
  • Completely missed the 'MAGIC BURGER' title text.
  • The burger looks like multiple burgers stacked rather than one 'exploded' burger.
  • Missed the required starburst element for the price.

Qwen Image 2.0

  • + Captured all three text elements accurately, including the starburst.
  • + Excellent fiery text effect and background atmosphere as requested.
  • + High photorealistic detail in the food textures and sauce.
  • The 'exploded' effect is minimal, with most components still touching.
  • The price text within the starburst lacks the requested fiery/glowing effect compared to the title.

Verdict: Qwen Image 2.0 is the clear winner as it followed all prompt instructions, including the complex text requirements and the starburst element which FLUX.1 [dev] omitted. While FLUX.1 [dev] provided a more literal 'exploded' layout, Qwen Image 2.0 delivered a much higher quality advertisement aesthetic with superior lighting and texture.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [dev]

  • + Displays excellent text rendering with zero spelling errors across the entire board.
  • + The handwriting appears consistent and has a convincing chalk-like texture.
  • + Centred composition creates a clean and organized aesthetic for the menu.
  • The text looks slightly too uniform, almost like a digital chalk font rather than manual handwriting.
  • The 'Ask us for our' text at the bottom has a slight merging of letters ('usfor').

Qwen Image 2.0

  • + Features highly realistic chalk smudges and dust that enhance the authentic atmosphere.
  • + The letter size and slant variations look naturally handwritten rather than font-based.
  • + Successful composition with the lighting casting a soft glow across the board.
  • Small spelling/spacing artifacts present, such as the double dash and spacing on the cookie line.
  • The title handwriting is less 'elegant cursive' as requested compared to Model A.

Verdict: Both models followed the complex text prompt remarkably well, but FLUX.1 [dev] is the winner due to its perfect technical execution of the specific menu items and date without any spelling errors. While Qwen Image 2.0 captured a more realistic 'messy' chalkboard atmosphere with convincing chalk dust, FLUX.1 [dev] provided a clearer, more professional-looking menu that perfectly adhered to the text requirements.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [dev]

  • + Features cinematic lighting with a soft atmospheric glow.
  • + Clean composition that focuses on the two main subjects.
  • + High resolution/clarity in the horse's coat and astronaut's suit.
  • Failed the negative constraint: the astronaut is riding the horse instead of the horse riding the astronaut.
  • Anatomical issues with the horse's legs, specifically the rear legs appearing detached or malformed.

Qwen Image 2.0

  • + Includes creative surreal elements like floating water droplets.
  • + Detailed texture work on the horse's shoulders and the astronaut's gear.
  • + Shows a dynamic sense of motion with the mane and tail.
  • Failed the negative constraint: the astronaut is riding the horse instead of the horse riding the astronaut.
  • The horse's front legs have awkward anatomy, particularly the jointing of the right leg.

Verdict: Both FLUX.1 [dev] and Qwen Image 2.0 failed the specific positional instruction 'horse on top, not vice versa,' resulting in standard images of an astronaut riding a horse. FLUX.1 [dev] produces a more cinematic, minimalist aesthetic, while Qwen Image 2.0 adds more surreal details like floating droplets, but both struggles with horse anatomy.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent fur texture and lighting on the capybara's face.
  • + Bored expression of the woman perfectly captures the requested 'normal' vibe.
  • + Atmospheric cinematic lighting with realistic depth of field.
  • Failed spatial placement: the passenger is sitting in the front seat next to the driver instead of the back seat.
  • The capybara's hands look more like paws from a different animal or distorted appendages.

Qwen Image 2.0

  • + Correctly placed the passenger in the back seat as requested.
  • + The capybara's hands/paws on the steering wheel are anatomically more convincing for a rodent.
  • + Captures the 'professional' taxi driver cap style more accurately.
  • The perspective from the window creates a slightly confusing composition with the car frame.
  • The businesswoman looks slightly less 'bored' and more just neutral, though still acceptable.

Verdict: While FLUX.1 [dev] produced a more aesthetically pleasing image with superior lighting, it failed a key spatial instruction by placing the passenger in the front seat. Qwen Image 2.0 followed the prompt more accurately by placing the woman in the back seat and provided a more realistic interpretation of the capybara's paws on the wheel.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [dev]

  • + Features a beautiful thorny border and central illustration.
  • + Good contrast and cinematic lighting on the jack-o-lantern.
  • Numerous text errors including 'Falloween Rantcl' and 'You Tre'.
  • Did not include the spider webs requested in the prompt.
  • Poor text layout with redundant information like '7pm, 7pm'.

Qwen Image 2.0

  • + Excellent text accuracy, rendering all requested words and details perfectly.
  • + Successfully captures the 'vintage gothic parchment' aesthetic with webs and thorns.
  • + Cinematic atmosphere with misty trees and a central glowing pumpkin.
  • The transition from the central scene to the parchment border is slightly harsh near the bottom.

Verdict: Qwen Image 2.0 is the clear winner as it followed every instruction, including the specific texture of the parchment, the inclusion of spider webs, and flawless text rendering. FLUX.1 [dev] struggled significantly with the text content and ignored the web requirement, resulting in a less functional invitation.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent handle on the 'miniature 3D cartoon' aesthetic with soft, clay-like textures.
  • + Well-executed isometric perspective and diorama base.
  • + Refined lighting and shadows create a cohesive 3D scene.
  • Text rendering is poor, with 'SUSHI' misspelled as 'SUSH CATON' and a random Kanji character.
  • Texture is a bit too soft, lacking the realistic PBR materials requested.

Qwen Image 2.0

  • + Perfect text rendering for 'JAPAN' and 'SUSHI'.
  • + High-quality material rendering on the wood and fish textures.
  • + Clear adherence to the layout instructions with the flag icon.
  • Failed the '3D cartoon' style, opting for a photorealistic look instead.
  • The wooden board is an interesting addition, but the overall diorama feel is less stylized than requested.

Verdict: FLUX.1 [dev] followed the stylistic cues for a 3D isometric cartoon diorama much better than Qwen Image 2.0, which produced a more photorealistic image. However, Qwen Image 2.0 was significantly better at rendering the requested text and flag accurately, whereas FLUX.1 [dev] struggled with spelling and letter forms.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [dev]

  • + Warm and appealing lighting effects
  • + Clean and cute character designs that are consistent in style
  • Failed to include a tabby kitten, instead showing two puppies
  • Characters look like stylized 3D renders rather than the requested hyper-photorealistic style
  • Static pose that does not represent 'chasing' or 'tumbling'

Qwen Image 2.0

  • + Followed the prompt exactly, including the puppy, tabby kitten, bunny, and fox kit
  • + Captured the action of tumbling and playing perfectly
  • + Achieved a high level of photorealistic detail in the fur and environment
  • The fox kit has a slightly distorted facial expression while tumbling
  • Minor butterfly scale issues relative to the animals

Verdict: Qwen Image 2.0 followed the prompt significantly better by including all four specific animal types and capturing the requested action of 'tumbling together' in a photorealistic style. FLUX.1 [dev] produced a much more stylized, cartoonish image that failed to include the kitten and lacked the dynamic movement requested.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent vector logo composition and balance
  • + Nice subtle paper texture on the background
  • + Includes all requested elements like the cloche and banner
  • Major spelling errors with 'Flarilaan' and 'Reseaurant'
  • Conflicting dates with '1941' and '11011' added unnecessarily

Qwen Image 2.0

  • + Perfect text rendering for both 'Caffè Florian' and 'Est. 1720'
  • + Very clean illustration style with great use of gradients
  • + Precisely follows all prompt requirements including the cloche and banner
  • Composition feels slightly cramped compared to a traditional logo
  • The steam effect is a bit bulky/solid

Verdict: Qwen Image 2.0 is the clear winner because it correctly spells all requested text, whereas FLUX.1 [dev] produced significant typos and added random nonsensical numbers. Qwen Image 2.0 also achieved a much cleaner 'vector emblem' look that aligns with modern logo design standards while maintaining the vintage aesthetic.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.1 [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 [dev]

  • + Excellent aesthetic alignment with the flat-vector infographic style request.
  • + Sophisticated composition with a central focus and balanced graphical elements.
  • + Adheres well to the specific NASA-inspired muted color palette.
  • The text is largely nonsensical gibberish or misspelled.
  • The icons do not clearly represent the specific requested steps (e.g., Saturn V, lunar module).

Qwen Image 2.0

  • + Perfect text rendering for all the requested steps of the mission.
  • + High logical adherence to the 'Steps' requested, specifically showing the orbit rings and landing module.
  • + Clean, readable layout that functions effectively as an educational poster.
  • The vector style is slightly inconsistent, with the landing module being more detailed than the other flat icons.
  • The composition is a simple vertical list which is less creative than the infographic layout in Image A.

Verdict: While FLUX.1 [dev] produced a more visually stunning and 'designer' quality vector layout, it failed significantly on text legibility and following the specific logical steps. Qwen Image 2.0 followed every instruction, including the specific requested steps and text labels with perfect accuracy, making it the superior infographic overall.

Next steps

Explore each model