Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 Kontext [dev] Black Forest Labs Qwen Image Alibaba

Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.

FLUX.1 Kontext [dev]

17.1 arena score

#54 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image

20.8 arena score

#38 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 Kontext [dev]

0%

win rate

Ties

0%

Qwen Image

0%

win rate

Shared challenges 13

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 Kontext [dev]
Qwen Image

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent texture on the red book cover and wood grain
  • + Accurately represents the transparency and refraction of the plant through the glass
  • + Natural soft lighting with realistic window reflections on the sphere
  • The glass cube lacks a distinct bottom edge where it meets the table, appearing slightly fused

Qwen Image

  • + Very clean geometric construction of the glass cube
  • + Good composition with a wider field of view of the table
  • + Accurate placement of all requested elements
  • The green plant in the background is overly blurred and lacks detail
  • The sphere appears slightly flatter with less realistic reflections compared to the other model

Verdict: Both models followed the prompt perfectly, including the complex spatial relationships of the objects. FLUX.1 Kontext [dev] is the winner due to its superior photographic realism, particularly in the textures of the book and the realistic way the light interacts with the materials, whereas Qwen Image feels slightly more like a CGI render.

Candid Street Photography

Text-to-Image

“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”

FLUX.1 Kontext [dev]
Qwen Image

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent sharpness and color vibrance
  • + Very clean face and hand rendering
  • + Captures the reflection on the wet pavement effectively
  • The subject is posing with the bike rather than 'repairing' it
  • The composition feels centered and studio-lit rather than 'candid' or 'imperfectly framed'
  • The man's suit looks too formal for a candid street photography repair scene

Qwen Image

  • + Successfully captures the 'repairing' action from the prompt
  • + Composition feels more authentic to candid street photography with its offset framing
  • + Excellent application of motion blur on passing vehicles and natural-looking light rain
  • Resolution is lower with softer details compared to the competitor
  • Noticeable anatomical artifacts in the hands while working on the bicycle seat
  • Colors are a bit muted and muddy in certain areas

Verdict: FLUX.1 Kontext [dev] produces a much higher quality image in terms of resolution and detail, but it fails the specific action of the prompt, showing a man simply sitting on a bike. Qwen Image adheres much better to the narrative of 'repairing' the bike and captures the requested 'candid' and 'imperfect' aesthetic far more accurately, despite the lower technical clarity.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 Kontext [dev]
Qwen Image

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent realism in skin texture and facial lighting.
  • + High-quality engraving on the armor that feels integrated into the design.
  • + Subtle and professional bokeh effects that enhance the portrait feel.
  • Missed the request for braided hair with beads.
  • The armor appears slightly more like leather or brass than classical plate steel.

Qwen Image

  • + Accurately included the braided hair and multicolored beads as requested.
  • + Fantastic texture on the leather straps, buckles, and chainmail underlayer.
  • + Visible torch and sparks add dynamic energy to the composition.
  • The sparks have a somewhat digital, artificial look compared to the rest of the scene.
  • The scars look like painted-on marks rather than integrated skin tissue.

Verdict: While FLUX.1 Kontext [dev] creates a more lifelike and cinematic face, Qwen Image followed the specific prompt details much better, particularly regarding the braided hair and beads. Qwen Image also delivered higher detail on the various material layers like the chainmail and leather straps, making it the more accurate interpretation of the prompt despite some minor issues with spark realism.

Modern Clean Menu

Text-to-Image

“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”

FLUX.1 Kontext [dev]
Qwen Image

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + High resolution food photography with realistic textures
  • + Effective use of negative space for a modern aesthetic
  • + Bold, clean typography that aligns with the professional layout request
  • Nonsensical text that doesn't clearly delineate the requested categories
  • Odd food combinations like pasta on pizza and fried textures that look slightly messy

Qwen Image

  • + Excellent adherence to category sections for Appetizers, Pizza, and Mains
  • + Extremely clean and vibrant grid layout that feels like a real template
  • + Legible prices and header text that enhance the menu functionality
  • Small body text and sub-headers contain garbled AI lettering
  • Slightly more 'clipart' feel to some food photos compared to the realism in the other model

Verdict: Qwen Image is the superior choice because it perfectly captures the structure of a menu, including the requested sections for Appetizers and Pizza/Mains with a clear price list. While FLUX.1 Kontext [dev] has higher visual fidelity in its food photography, its layout is more of a collage and fails to function as a legible menu.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 Kontext [dev]
Qwen Image

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent primary title rendering with clean 3D effects.
  • + Vibrant and detailed fire/burning coal effect at the base.
  • + High image resolution and crisp details on the bun and patty.
  • Failed to produce an 'exploded' burger, showing a fully assembled one instead.
  • Spelling error in the secondary text ('LNHLY' instead of 'ONLY').
  • The starburst for the price is very flat and lacks the requested glowing effect.

Qwen Image

  • + Successfully captured the 'exploded' effect with suspended ingredients.
  • + Perfectly rendered all requested text with no spelling errors.
  • + Better integration of the glowing, fiery effect on the text and starburst.
  • The 'LIMITED TIME ONLY' text is quite small compared to the other elements.
  • Some ingredients look slightly cartoonish or plastic-like compared to Model A.

Verdict: Qwen Image followed the complex prompt instructions much more accurately, successfully executing the 'exploded' burger concept and providing flawless text rendering. FLUX.1 Kontext [dev] produced a higher quality static burger, but failed significantly on the composition (no explosion) and text accuracy (spelling error).

Chalkboard Menu

Text-to-Image

“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”

FLUX.1 Kontext [dev]
Qwen Image

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Features a more realistic chalk-on-blackboard texture with dust and smudges.
  • + Captures a wider range of handwriting styles from bold block to casual script.
  • Has significant spelling errors like 'Risoktso', 'Mashroom', and 'Octpus'.
  • Contains repetitive text and garbled characters in the date and price sections.

Qwen Image

  • + Perfect legibility for the main menu items with no spelling errors.
  • + Excellent composition that sets the chalkboard within a cozy café environment.
  • + Follows the prompt's formatting for the menu items more accurately.
  • Incorrectly renders the year as '20026' instead of '2026'.
  • The text looks more like a digital font or a clean vector rather than realistic porous chalk texture.

Verdict: While FLUX.1 Kontext [dev] has a more authentic chalk texture and physical board appearance, it fails significantly on spelling and text coherence. Qwen Image provides a much more professional and legible menu with better environmental context, though its year was slightly off and the text feels a bit too clean compared to real chalk. Qwen Image is the better choice for a functional and aesthetically pleasing result.

The Reversed Rodeo

Text-to-Image

“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”

FLUX.1 Kontext [dev]
Qwen Image

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Successfully followed the difficult logic of the horse riding the astronaut.
  • + High detail in the astronaut face and textures.
  • + Strong surrealist composition with floating spheres.
  • The astronaut's hands and the horse's back leg integration are slightly awkward.
  • The horse's chest merges directly into the astronaut’s back without a clear saddle or mounting mechanism.

Qwen Image

  • + Very cinematic lighting and background composition.
  • + High quality rendering of the space suit and horse's muscles.
  • Completely failed the semantic instruction to have the horse on top.
  • Standard, cliché interpretation of the prompt topic.

Verdict: The main differentiator was the complex logical instruction 'horse on top, not vice versa'. FLUX.1 Kontext [dev] followed this surreal instruction perfectly, whereas Qwen Image defaulted to the common trope of an astronaut riding a horse, ignoring the specific reversal requested.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 Kontext [dev]
Qwen Image

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent photorealistic texture on the capybara's fur and the woman's face.
  • + Captures the bored, indifferent expression of the human passenger perfectly.
  • + High-quality lighting that feels natural for a night-time city environment.
  • The capybara only has one hand on the steering wheel, failing the prompt's request for both paws.
  • The proportions of the capybara's head to its body feel slightly uncanny.

Qwen Image

  • + Successfully placed both paws on the steering wheel as requested.
  • + The yellow taxi hat is more stylish and fits the 'taxi driver' uniform aesthetic better.
  • + Better overall composition and sense of motion through the window.
  • The hands on the steering wheel look more like primate hands than capybara paws.
  • The human passenger's hand holding the phone is slightly mangled around the fingers.
  • The text on the taxi sign says 'YOXI' rather than an accurate taxi light.

Verdict: Qwen Image followed the specific instruction to have both paws on the steering wheel, whereas FLUX.1 Kontext [dev] only depicted one. However, FLUX.1 Kontext [dev] achieved a much higher level of photorealism and captured the requested 'bored' expression of the passenger more effectively.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 Kontext [dev]
Qwen Image

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent typography rendering for the primary title and date.
  • + Clear thorny border that frames the image well.
  • + Strong, vibrant jack-o-lantern glow.
  • Internal banner text is illegible gibberish.
  • The background is very flat and lacks the 'moody night sky' and 'twisted trees' requested in detail.
  • Misspelled 'The Arches' as 'The Argiiah's'.

Qwen Image

  • + Beautiful gothic aesthetic with atmospheric trees and a moody sky.
  • + Excellent parchment texture and intricate web/thorn border.
  • + Correct spelling of all event details and the scroll banner text.
  • Typos in the main title ('Hallo Party' instead of 'Halloween Party').
  • The scroll banner is small and placed awkwardly in the center.
  • Resolution of the smaller text is slightly soft compared to Model A.

Verdict: Qwen Image is the superior choice because it captures the requested 'vintage gothic' atmosphere perfectly with twisted trees and a parchment background, whereas FLUX.1 Kontext [dev] is very flat. Although Qwen Image has a typo in the main header, its ability to correctly render the secondary text and the complex visual scene makes it more effective as a design piece.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 Kontext [dev]
Qwen Image

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent typography with clean, bold characters.
  • + Extremely clean, minimalist aesthetic that follows the 'soft texture' requirement.
  • + Perfect centering of the main subject.
  • The flag icon is stylized to the point of being unrecognizable as the Japanese flag.
  • The sushi design is overly simplistic and doesn't fully capture 'Japan's signature dish' variety.

Qwen Image

  • + Successfully includes the flag icon both in text and as a 3D miniature object.
  • + Beautifully detailed isometric diorama with varied sushi types and garnishes.
  • + Strong adherence to the 'miniature 3D cartoon scene' and '45° top-down' prompt.
  • The text 'JAPAN' is slightly less crisp than Model A.
  • Slightly more visual noise compared to the 'ultra-clean' request, though still very high quality.

Verdict: Qwen Image is the winner as it much better captures the 'miniature diorama' and 'Japan' theme by including a recognizable flag and a variety of sushi pieces. While FLUX.1 Kontext [dev] has superior typography, it failed to produce a correct flag icon and its interpretation of the scene was perhaps too minimal.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 Kontext [dev]
Qwen Image

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Dynamic sense of movement and playful energy
  • + Consistent lighting and color palette
  • Failed to include the fox and the bunny
  • Anatomy issues with the center kitten appearing somewhat distorted
  • Includes two kittens instead of the diverse set of animals requested

Qwen Image

  • + Successfully included all four requested animals (dog, kitten, bunny, fox)
  • + Excellent lighting with visible god rays and dew sparkles as requested
  • + High detail in fur textures and expressive eyes
  • The fox looks slightly more stylized/illustrative than hyper-photorealistic
  • Composition is a bit crowded with all animals in one center cluster

Verdict: Qwen Image is the clear winner because it followed the prompt instructions to include four specific types of baby animals, whereas FLUX.1 Kontext [dev] only generated a dog and two kittens. Additionally, Qwen Image captured the specific atmosphere details like god rays and dew sparkles much more effectively.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 Kontext [dev]
Qwen Image

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent typography rendering with perfect spelling
  • + Clean vector aesthetic that feels professional
  • + Accurate use of color tones
  • The dome looks more like a building or cupcake wrapper than a cloche
  • Missing the requested 'banner' for the establishment date

Qwen Image

  • + Better representation of a cloche dome with steam
  • + Successfully included the banner element
  • + Good use of subtle paper texture
  • Severe spelling and layout issues in the main brand name
  • Typography is messy and overlapping
  • The capital letters and lower case letters are jumbled

Verdict: FLUX.1 Kontext [dev] produced a much more usable and polished logo with perfect text rendering, despite missing the banner element and having a slightly abstract cloche. Qwen Image captured more of the specific prompt elements (cloche shape, banner, steam) but failed significantly on the typography, making the brand name unreadable.

Apollo 11: Journey to Tranquility

Text-to-Image

“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”

FLUX.1 Kontext [dev]
Qwen Image

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Strong minimalist aesthetic with clean lines.
  • + Uses the requested color palette effectively.
  • + Icons are stylized and uniform in appearance.
  • Extreme spelling errors including the main heading 'APOLO'.
  • The icons are abstract to the point of being unreadable for the specific mission steps.
  • The layout is cluttered and difficult to follow as a sequence.

Qwen Image

  • + Logical infographic flow that tells a story from launch to landing.
  • + Clear and recognizable iconography for planets, ships, and modules.
  • + Much better text legibility even with minor spelling errors.
  • Included instructions like 'Stop at landing' as actual text on the image.
  • Missing step numbers 1 and 5 in the sequence.
  • Typography is a bit basic compared to the modern vector prompt.

Verdict: Qwen Image is the clear winner for its functional layout and recognizable icons that actually map to the requested mission steps, despite mistakenly including part of the prompt text. FLUX.1 Kontext [dev] produced a visually striking style but failed significantly on legibility, spelling the primary title incorrectly and creating icons that bear no resemblance to the actual Saturn V or Lunar Module.

Next steps

Explore each model