Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 Kontext [pro] Black Forest Labs Qwen Image 2.0 Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

FLUX.1 Kontext [pro]

19.5 arena score

#44 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image 2.0

21.7 arena score

#34 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 Kontext [pro]

0%

win rate

Ties

0%

Qwen Image 2.0

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 Kontext [pro]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent photographic realism with natural depth of field.
  • + Accurate spatial relationships between the book, cube, and sphere.
  • + Subtle and realistic light interaction on the glass surfaces.
  • The blue sphere has a felt-like texture rather than appearing as a smooth geometric object.
  • The blue sphere appears to be floating without a clear shadow on the bottom glass.

Qwen Image 2.0

  • + Strong adherence to all prompt elements including placement and color.
  • + Complex and realistic reflections of the sphere in the glass panels.
  • + Highly detailed textures on the book cover and wooden table.
  • The glass cube has some structural inconsistencies in its edges.
  • The blue sphere is perfectly centered in a way that feels slightly less natural than Model A.

Verdict: Both models followed the prompt perfectly, capturing the specific arrangements of objects and lighting. FLUX.1 Kontext [pro] produces a more convincing, high-quality photograph with natural lighting, while Qwen Image 2.0 excels at rendering complex optical reflections within the glass cube.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.1 Kontext [pro]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent atmosphere with visible rain and cinematic mood
  • + Highly realistic skin textures and natural lighting
  • + Captures a strong sense of emotion and focus in the subject
  • The subject is holding the handlebars rather than 'repairing' the bike as requested
  • Fails to include the requested motion blur from passing cars

Qwen Image 2.0

  • + Accurately depicts the act of repairing the bicycle pedal/chain
  • + Excellent wet pavement reflections that enhance the realism
  • + Successful implementation of 'imperfect framing' for a candid street photography feel
  • The rain is almost invisible compared to Image A
  • The car in the background lacks the requested motion blur
  • Minor anatomy issues with how the right hand interacts with the pedal

Verdict: While FLUX.1 Kontext [pro] creates a more evocative and atmospheric image with superior skin textures and visible rain, it fails to show the man actually repairing the bike. Qwen Image 2.0 adheres better to the specific action of the prompt and excels with the candid framing and wet pavement reflections, despite the rain itself being less prominent.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 Kontext [pro]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent photorealistic skin and hair textures
  • + Highly detailed and authentic-looking engraved plate armor
  • + Beautifully subtle lighting and shallow depth of field
  • Missed the request for small beads in the braids
  • The scars and dirt are very faint compared to the 'battle-worn' prompt

Qwen Image 2.0

  • + Successfully included beads in the hair as requested
  • + Stronger adherence to the 'battle-worn' aesthetic with visible scars and grime
  • + Includes armor, leather, and cloth layers clearly
  • Eyes look somewhat artificial and unnatural
  • The skin texture lacks the realistic pore detail of Model A
  • The hand placement on the sword hilt appears slightly awkward and poorly rendered

Verdict: FLUX.1 Kontext [pro] creates a much more cinematic and realistic portrait with superior texture and lighting, though it missed the specific detail of the beads. Qwen Image 2.0 followed the specific item descriptions (beads, scars, cloth) more closely but sufferedจาก a lower level of overall visual polish and slightly distorted eyes.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 Kontext [pro]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent typography with clean, bold sans-serif fonts
  • + Clean white space and professional alignment of text and images
  • + High clarity in the food photography
  • Asymmetrical layout feels slightly unbalanced compared to a grid
  • Prices are nonsensical (e.g., 243 for a salad)
  • Text contains minor gibberish, though letters are well-formed

Qwen Image 2.0

  • + Perfect grid layout for food photos as requested
  • + High-quality, appetizing food photography with vibrant colors
  • + Strong visual balance across the entire canvas
  • Text rendering is poor with garbled characters and illegibility
  • Lacks detailed text-based menu descriptions
  • Header alignment is slightly disconnected from columns

Verdict: Qwen Image 2.0 followed the grid layout instructions much more effectively, creating a vibrant and visually balanced design, though its text rendering is poor. FLUX.1 Kontext [pro] produced superior typography and a more realistic professional feel, but opted for an alternating layout rather than the requested grid. Qwen Image 2.0 is the slightly better choice for a design mock-up due to its superior composition and photography, provided the text is eventually replaced.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 Kontext [pro]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent typography and layout for 'MAGIC BURGER' and 'LIMITED TIME ONLY'.
  • + Exceptional photorealism on the burger materials, especially the dripping cheese and bun texture.
  • + Strong atmospheric lighting that blends the product with the background naturally.
  • Repeated the price text '€6.99' twice in different positions.
  • Failed to include the requested 'starburst' shape for the price.

Qwen Image 2.0

  • + Accurately followed the 'starburst' instruction for the price graphic.
  • + Impressive fiery texture applied directly to the font of the main title.
  • + Good sense of motion with flying morsels and smoke trails.
  • The 'LIMITED TIME ONLY' text is small and slightly difficult to read against the busy background.
  • The composition feels a bit more cluttered compared to the clean layout of the opponent.

Verdict: FLUX.1 Kontext [pro] creates a more professional and delicious-looking advertisement with superior font rendering and realistic product textures. However, Qwen Image 2.0 followed the specific structural prompt instructions more closely by including the requested starburst and applying aLiteral fire effect to the text.

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 Kontext [pro]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent text legibility and spelling accuracy throughout the board.
  • + Strong adherence to the requested chalk texture and handwriting style.
  • + Perfectly centered and balanced composition within the frame.
  • The title is not in 'elegant cursive' as requested, but rather a blocky print style.
  • The background is very blurry and provides less 'cozy café' atmosphere than the competitor.

Qwen Image 2.0

  • + Successfully used a cursive/script style for the 'Today's Specials' title reflecting the prompt.
  • + Highly realistic chalkboard texture with smudges and natural chalk dust effects.
  • + Captures the 'cozy café' atmosphere better with a more detailed background scene.
  • The text layout is a bit cluttered with prices detached from the items.
  • The handwritten style of the menu items is slightly less consistent than Model A.

Verdict: Both models followed the prompt well, but FLUX.1 Kontext [pro] produced much cleaner and more professional-looking text with perfect spelling. Qwen Image 2.0 followed the stylistic instruction for cursive title handwriting and better captured the ambient lighting of a café, but the text placement feels slightly disjointed compared to the clean layout of FLUX.1 Kontext [pro].

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 Kontext [pro]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent adherence to the specific 'horse on top' request
  • + High cinematic quality with realistic space lighting
  • + Creative and surreal execution of the prompt's unusual requirement
  • Anatomical fusion issues where the horse's back leg meets the astronaut's shoulder
  • The astronaut has hooves instead of boots

Qwen Image 2.0

  • + High visual clarity and sharp details on the astronaut suit
  • + Good composition with interesting floating water droplets
  • Failed the core prompt instruction of 'horse on top, not vice versa'
  • Generic interpretation of a common AI trope

Verdict: FLUX.1 Kontext [pro] successfully interpreted the difficult and specific instruction to have the horse riding the astronaut, creating a truly surreal image as requested. Qwen Image 2.0 ignored the specific positioning instruction entirely, providing a standard astronaut-on-horse image which makes it a failure for this specific prompt.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 Kontext [pro]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent photorealistic texture on the capybara's fur
  • + Accurate depiction of a capybara's facial expression
  • + The passenger is correctly positioned in the back seat
  • The passenger is holding a phone but looking away from it, missing the 'looking at her phone' prompt
  • Only one paw is clearly on the steering wheel

Qwen Image 2.0

  • + Perfect adherence to the passenger behavior of looking at her phone with a bored expression
  • + Successfully placed both front paws on the steering wheel
  • + Very clear and colorful city lights in the background
  • The passenger appears to be in the front passenger seat rather than the back seat
  • The capybara's head is disproportionately large for its body
  • The lighting on the capybara is slightly flatter and less integrated than model A

Verdict: Both models captured the essence of the prompt, but Model A (FLUX.1 Kontext [pro]) is the superior image due to its more realistic fur textures, accurate taxi depth with the passenger clearly in the back seat, and better lighting. While Model B (Qwen Image 2.0) followed the character actions more closely (looking at the phone and two paws on the wheel), the composition feels cramped, and the passenger appears in the front seat, which deviates from the prompt's layout.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 Kontext [pro]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent font choice for the gothic title and banner.
  • + Great utilization of the entire image space for a dense atmospheric effect.
  • + Accurate rendering of the thorn and web border theme.
  • Includes a line of hallucinated gibberish text in the event details.
  • Uses a comma instead of a period for the date, deviating slightly from the prompt.

Qwen Image 2.0

  • + Captures the 'dark parchment' aesthetic perfectly with a paper-like texture.
  • + Follows all text instructions exactly with zero hallucinations or typos.
  • + Clearer central jack-o-lantern rendering with strong cinematic lighting.
  • The 'twisted trees' look a bit like generic 3D assets rather than integrated vintage art.
  • The cloud background is slightly less 'moody night sky' and more of a gray fog.

Verdict: Qwen Image 2.0 is the winner because it followed the text instructions perfectly without adding hallucinated lines of text, which FLUX.1 Kontext [pro] failed at. While FLUX.1 had a more integrated vintage poster feel, the accuracy of the date and details in Qwen Image 2.0 makes it more functional as an invitation.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 Kontext [pro]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent 3D cartoon aesthetic with soft, pillowy textures.
  • + Perfect adherence to the 45-degree isometric perspective.
  • + Typography is beautifully integrated into the overall art style.
  • The rice grains look more like rounded bumps than individual rice grains.
  • The flag icon is stylized but slightly lacks the precise circular proportion of a Japan flag.

Qwen Image 2.0

  • + Realistic food textures especially on the eel and tuna.
  • + Accurate and clear rendition of the Japanese flag.
  • + Clean, bold typography that is easy to read.
  • The perspective is more of a standard high-angle photo than a 45-degree isometric diorama.
  • The composition feels a bit crowded with many sushi types despite 'minimal' and 'miniature' prompts.
  • Lacks the specific '3D cartoon' miniature style requested.

Verdict: FLUX.1 Kontext [pro] captures the requested 3D cartoon diorama aesthetic much more effectively, presenting a clean, cohesive isometric scene. While Qwen Image 2.0 has higher realism in the food textures, it fails to deliver the specific 'miniature 3D cartoon' style and feels more like a stock photo on a plain background.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 Kontext [pro]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent fur texture rendering and soft lighting.
  • + Cohesive facial expressions with a unified joyful vibe.
  • + Great bokeh effect and integrated butterfly placement.
  • The 'bunny' looks more like a white kitten with elongated ears than a rabbit.
  • Less dynamic interaction; animals are mostly sitting still.

Qwen Image 2.0

  • + Strong prompt adherence with animals actually 'tumbling together' and playing.
  • + Distinct and accurate biological features for the bunny and fox.
  • + Beautiful god rays and a diverse color palette of wildflowers.
  • Composition feels slightly crowded with some odd anatomy in the fox's pose.
  • One butterfly is hovering very close to the cat's ear in an unnatural way.

Verdict: Qwen Image 2.0 is the winner as it much more accurately depicts the specific species requested, particularly the rabbit and the fox, while capturing the 'tumbling' action from the prompt. FLUX.1 Kontext [pro] creates a beautiful, soft image, but its animals look somewhat generic and lack the playful interaction described.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 Kontext [pro]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Perfect adherence to the minimalist vector emblem style.
  • + Excellent typography with correctly rendered Italian accent mark.
  • + Superior subtle texture that gives it an authentic vintage paper feel.
  • Simple typo in the banner background text ('EEST' instead of 'EST').

Qwen Image 2.0

  • + Strong visual hierarchy and attractive color palette.
  • + Perfectly rendered text for both the name and the date.
  • + More creative integration of the steam element inside the cloche.
  • Less 'minimalist' than requested, bordering on a complex illustration.
  • The layout of the banner at the bottom is slightly unbalanced on the right side.

Verdict: FLUX.1 Kontext [pro] creates a much more authentic vintage minimalist logo with a perfect vector-style aesthetic, despite a small typo in the establishing date. Qwen Image 2.0 is more polished in its text rendering but fails the 'minimalist' part of the prompt by including complex gradients and highlights. FLUX.1 is the preferred choice for a brand identity task.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.1 Kontext [pro]
Qwen Image 2.0

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Clean vector illustration style that feels artistic and modern.
  • + Beautifully rendered lunar surface and atmospheric navy background.
  • + Successfully incorporates NASA-inspired colors into the graphics.
  • Confused information layout; labels like 'Earth' are pointed at the Moon.
  • Bizarrely includes Saturn (a planet) instead of the Saturn V rocket.
  • Text and labels are disorganized and include nonsensical terms like '2,f ar Collins'.

Qwen Image 2.0

  • + Perfect adherence to the requested logical steps (Launch through Landing).
  • + Highly legible typography and correct spelling of mission milestones.
  • + Excellent preservation of the flat-vector iconography requested in the prompt.
  • Minor typo in 'Translunjar' (added 'j').
  • Slightly less 'artistic' or dynamic composition compared to Model A.

Verdict: Qwen Image 2.0 is the clear winner for its superior logical coherence and adherence to the step-by-step instructions. While FLUX.1 Kontext [pro] produced a more visually striking illustration, its content was factually hallucinatory, placing Saturn in a lunar mission infographic and mislabeling planets. Qwen Image 2.0 correctly followed the sequential iconography and maintained a clean, professional layout that functions as an actual infographic.

Next steps

Explore each model