Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 Kontext [dev] Black Forest Labs Qwen Image 2.0 Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

FLUX.1 Kontext [dev]

16.5 arena score

#58 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image 2.0

21.7 arena score

#34 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 Kontext [dev]

0%

win rate

Ties

0%

Qwen Image 2.0

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 Kontext [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent adherence to lighting instructions with a clear soft glow from the left.
  • + Superior texture rendering on the book and the wooden table surface.
  • + Highly realistic glass refraction and reflections on the sphere.
  • The glass cube is missing its front face, appearing more like a glass stand or table.

Qwen Image 2.0

  • + Accurately depicts the glass as a full enclosed cube.
  • + Follows all spatial requirements including the sphere being 'inside' and the plant 'behind'.
  • + Realistic reflections of the sphere on the inner glass surfaces.
  • The sphere appears to be floating unnaturally in the center of the air rather than resting on the bottom.
  • The lighting is somewhat flat compared to the requested window light.
  • The glass has slight structural inconsistencies where the vertical panes meet.

Verdict: FLUX.1 Kontext [dev] produced a much more aesthetically pleasing image with high-end textures and lighting, though it failed to render a fully enclosed cube. Qwen Image 2.0 followed the technical prompt more accurately by providing a closed cube and including the sphere fully inside, but the image quality and lighting feel less natural than FLUX.1 Kontext [dev].

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.1 Kontext [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent adherence to the 'light rain' and 'cinematic' prompts with visible raindrops and bokeh.
  • + High image resolution and clean rendering of clothes and the environment.
  • Fails the main action as the man is standing/mounting the bike rather than 'repairing' it.
  • The composition feels artificial and centered, lacking the 'imperfect framing' requested.

Qwen Image 2.0

  • + Perfect adherence to the 'repairing' action and 'imperfect framing' of a candid shot.
  • + Superior realistic skin textures and natural, non-stylized lighting.
  • + Captures the 'motion blur from passing cars' more accurately than Model A.
  • The 'light rain' is less visually obvious in the air compared to Model A.
  • Minor anatomical confusion in the overlapping of the hands near the bicycle crank.

Verdict: Qwen Image 2.0 is the clear winner for its superior grasp of the requested narrative and photographic style, successfully depicting a man actually repairing a bike with authentic candid framing. FLUX.1 Kontext [dev] produced a high-quality image, but it ignored the 'repairing' instruction, instead showing a man simply posing with a bicycle.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 Kontext [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent cinematic lighting with a focused warm torchlight glow.
  • + Ornate engraving on the armor is exceptionally high quality and coherent.
  • + Subtle and realistic implementation of scars and skin texture.
  • Failed to include the specific instruction for beads in the hair.
  • The background is very dark, losing some of the environment detail requested.

Qwen Image 2.0

  • + Strong adherence to specific prompt details like beads in the hair and dirt on the skin.
  • + Outstanding texture on leather straps and frayed cloth underlayers.
  • + Highly accurate rendering of the 'battle-worn' aesthetic with realistic grime.
  • The hand placement and anatomy on the sword hilt are slightly awkward.
  • The bright fire in the background creates harsher lighting than the 'warm torchlight' requested.

Verdict: Qwen Image 2.0 captures nearly every specific detail of the prompt, including the beads in the braids and the intricate textures of the under-armor layers, providing a truly 'battle-worn' feel. FLUX.1 Kontext [dev] creates a more polished and aesthetically pleasing cinematic portrait with superior armor engravings, but it misses several key descriptive elements like the hair beads.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 Kontext [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent grid composition with interlocking text and image blocks
  • + High-quality, vibrant food photography
  • + Cleaner typography execution with fewer rendering glitches in the large fonts
  • Text is mostly gibberish despite looking like words
  • Missing specific categories like 'Pizza' requested in the prompt

Qwen Image 2.0

  • + Directly follows the 'Appetizers/Pizza/Mains' section requirement
  • + Logical layout for a functional menu with prices
  • + Clearer, more recognizable food subjects in the photos
  • Smaller text is heavily distorted and difficult to read
  • The grid feels a bit generic compared to a modern magazine-style layout

Verdict: Qwen Image 2.0 followed the specific structural requirements of the prompt by including the requested sections for Appetizers, Pizza, and Mains. However, FLUX.1 Kontext [dev] produced a much more visually compelling and modern design that better captures the 'minimalist' and 'professional' aesthetic, despite missing some category keywords. Qwen Image 2.0 is the choice for layout accuracy, while FLUX.1 Kontext [dev] wins on artistic design and image quality.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 Kontext [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Strong, vibrant lighting on the coals and background.
  • + Clean, legible title typography.
  • Failed the 'exploded' instruction as the burger is fully assembled.
  • Spelling error in the secondary text ('LNHLY' instead of 'ONLY').
  • The starburst element looks like a flat graphic rather than part of the scene.

Qwen Image 2.0

  • + Successfully captured the 'exploded' mid-air effect with separated components.
  • + Perfect text rendering for all requested messages.
  • + Superior textures on the meat patty, cheese, and tomatoes.
  • The fiery effect on the text is slightly cliché but fits the prompt.

Verdict: Qwen Image 2.0 followed the prompt instructions much more accurately, successfully depicting an exploded burger while FLUX.1 Kontext [dev] rendered a static, fully assembled one. Qwen also provided perfect text spelling and a more professional food-photography aesthetic compared to the typos and flat graphics in FLUX.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 Kontext [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Successfully captures a realistic chalk texture on the letters.
  • + Includes the specific date and price points requested in the prompt.
  • Numerous spelling errors including 'Mashroom', 'Risoktso', and 'Octpus'.
  • Failed to render the title in 'elegant cursive' as requested, opting for a blocky style.

Qwen Image 2.0

  • + Excellent prompt adherence with near-perfect spelling for all items.
  • + Superior composition showing the board within a 'cozy café' environment rather than just a flat shot.
  • + Accurately rendered the title in a cursive-leaning script as requested.
  • The 'chalk smudge' effect is slightly overdone, making some text a bit harder to read.
  • Missing the specific year '2026' in the header (though the date is otherwise correct).

Verdict: Qwen Image 2.0 is the clear winner as it correctly spelled all complex menu items and followed the stylistic instruction for cursive handwriting. While FLUX.1 Kontext [dev] had a nice chalk texture, it suffered from severe typographical errors like 'Mashroom' and 'Risoktso' and failed to place the board in a café setting.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 Kontext [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Successfully followed the specific instruction for the horse to be on top of the astronaut.
  • + High realism in the textures of the space suit and the horse's fur.
  • + Clean, cinematic lighting that feels professional.
  • The composition is a bit rigid with both subjects floating in a neutral pose.
  • The astronaut's hands are slightly distorted and lack fine detail.

Qwen Image 2.0

  • + Excellent sense of motion and dynamic energy in the horse's pose.
  • + Beautiful rendering of stars, earth, and floating water droplets.
  • + Highly detailed textures on the horse's neck and the astronaut's suit.
  • Completely failed the negative constraint/specific instruction by placing the astronaut on top of the horse.
  • The horse's anatomy is slightly warped with the hind legs appearing elongated and thin.

Verdict: FLUX.1 Kontext [dev] followed the core challenge of the prompt, successfully depicting a horse riding an astronaut, which is a difficult spatial reasoning task. Qwen Image 2.0 produced a visually stunning and dynamic image but failed to follow the instruction for the unconventional positioning, resulting in a standard 'astronaut riding a horse' scene.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 Kontext [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent photographic lighting and skin texture for both the human and capybara.
  • + Captures the bored, indifferent expression of the passenger perfectly.
  • + High-quality fur rendering and clothing detail.
  • The capybara only has one hand on the steering wheel, failing that specific prompt instruction.
  • The capybara's head looks more like a groundhog or beaver than a distinct capybara.

Qwen Image 2.0

  • + Perfect adherence to the pose, showing both paws on the steering wheel.
  • + The capybara's anatomy and face shape are highly accurate to the species.
  • + The yellow driver cap feels more authentic to a classic taxi driver style.
  • The passenger appears to be sitting in the front passenger seat or too close to the driver rather than the back seat.
  • The image has slightly more digital artifacts and less realistic lighting compared to Image A.

Verdict: FLUX.1 Kontext [dev] produced a more cinematically polished and realistic image with a better 'bored' expression for the passenger, but failed the specific request to have both paws on the wheel. Qwen Image 2.0 followed the technical pose instructions more accurately and captured the specific features of a capybara much better, though it struggled with the spatial depth of the car interior. Qwen Image 2.0 is the winner for better anatomical accuracy and prompt adherence regarding the driving pose.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 Kontext [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Strong, high-contrast lighting on the Jack-o-lantern
  • + Clean main title text rendering
  • The banner text is completely illegible and garbled
  • Fails to provide the dark parchment background requested, opting for flat black
  • Location name is misspelled as 'The Argiiah's'

Qwen Image 2.0

  • + Excellent adherence to all text requirements with perfect spelling
  • + Superior composition with a moody parchment texture and atmosphere
  • + Includes all requested elements like webs, thorns, and a night sky
  • Jack-o-lantern rendering is slightly less vibrant than Model A's

Verdict: Qwen Image 2.0 is the clear winner as it followed every instruction, including the complex text in the banner and the specific parchment background. FLUX.1 Kontext [dev] failed significantly on text legibility for the smaller strings and ignored the parchment texture requirement.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 Kontext [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent adherence to the '3D cartoon' and 'miniature' aesthetic.
  • + Clean, minimalist composition that feels perfectly balanced.
  • + Accurate text rendering for both 'JAPAN' and 'SUSHI'.
  • The flag icon is stylized and abstract rather than being the actual Japanese flag.
  • Interpretation of sushi is overly simplified, leaning more towards toy-like than realistic PBR materials.

Qwen Image 2.0

  • + Accurate Japanese flag icon included at the top.
  • + High-quality realistic PBR textures on the sushi and wooden base.
  • + Good layout and text clarity.
  • Missed the 'cartoon' style requirement, opting for a realistic photography look instead.
  • The text is slightly off-center compared to the rest of the composition.

Verdict: FLUX.1 Kontext [dev] followed the stylistic instructions for a '3D cartoon miniature' much better than Qwen Image 2.0, which produced a realistic photograph. While Qwen Image 2.0 had better material realism and an accurate flag, FLUX.1 Kontext [dev] captured the intended isometric diorama aesthetic more effectively.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 Kontext [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent depiction of dynamic movement and joy
  • + Exceptional fur detail and lighting coherence
  • + Clean composition with a very cinematic bokeh effect
  • Failed to include the fox and the bunny
  • The kitten species is not clearly a 'tabby' in the foreground

Qwen Image 2.0

  • + Complete prompt adherence including all four specific animals
  • + Captures the 'god rays' and 'tumbling together' aspects perfectly
  • + High variety of wildflowers and textures
  • The fox kit has unnatural blue eyes and a slightly distorted head shape
  • The butterfly placement is a bit flat against the kitten's head

Verdict: Qwen Image 2.0 is the clear winner for prompt adherence as it successfully included all four requested animals (dog, cat, bunny, and fox) in a tumbling pile, whereas FLUX.1 Kontext [dev] missed half of the subjects. While FLUX.1 Kontext [dev] offered slightly cleaner lighting and resolution, Qwen Image 2.0 captured the specific 'god rays' and composition described in the prompt much more accurately.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 Kontext [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Perfectly follows the minimalist vector emblem style.
  • + Accurate and clean typography with correct accents.
  • + Excellent balance and alignment of elements.
  • The cloche dome icon is slightly abstract and could be mistaken for a building dome.
  • Missed the 'banner' element for the 'Est. 1720' text.

Qwen Image 2.0

  • + Incorporates all requested elements including the steam and banner.
  • + Strong vintage aesthetic with nice shading on the cloche.
  • + Includes subtle paper texture on the background as requested.
  • The typography is less professional with inconsistent spacing.
  • The banner is asymmetrical and terminates abruptly on the right side.
  • Less 'minimalist' than requested due to complex shading.

Verdict: FLUX.1 Kontext [dev] produced a much cleaner and more professional logo that fits the 'minimalist' and 'vector' keywords perfectly, though it missed the banner element. Qwen Image 2.0 followed the prompt more literally by including the banner and steam, but the execution of the typography and banner illustration is unpolished/unbalanced.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.1 Kontext [dev]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Follows the color palette accurately with a strong navy and muted red.
  • + Good flat vector aesthetic for the main icons.
  • Severe spelling errors in the headers and supporting text (e.g., 'APOLO', 'DESCNT').
  • Iconography is cluttered and confusing, making the steps difficult to follow.
  • Text overlaps and garbled characters make it non-functional as an infographic.

Qwen Image 2.0

  • + Excellent adherence to logical flow, showing a clear top-to-bottom progression of steps.
  • + Very high-quality typography with legible, correctly spelled text.
  • + Accurate iconography for the Saturn V, Lunar Module, and various orbits.
  • Minor typo in 'Translunjar' (contains an extra 'j').
  • The composition is a bit tight on the vertical axis, causing 'Lunar Orbit' to overlap its ring.

Verdict: Qwen Image 2.0 is the clear winner as it provides a functional, legible, and highly accurate infographic that follows the requested narrative steps perfectly. In contrast, FLUX.1 Kontext [dev] produced a disorganized layout with significant spelling errors and poorly rendered text that fails the basic requirements of an infographic.

Next steps

Explore each model