Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 Kontext [pro] Black Forest Labs Qwen Image Max Alibaba

Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.

FLUX.1 Kontext [pro]

20.3 arena score

#41 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image Max

21.5 arena score

#35 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 Kontext [pro]

0%

win rate

Ties

0%

Qwen Image Max

0%

win rate

Shared challenges 8

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 Kontext [pro]
Qwen Image Max

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent photographic realism with natural lens blur and lighting.
  • + Clean and accurate glass transparency that matches the prompt's layout perfectly.
  • + Highly detailed textures on the book and wooden table.
  • The sphere is slightly off-center and appears to be made of felt rather than glass or plastic.
  • The glass cube has very thin, almost frame-like edges instead of solid thick glass walls.

Qwen Image Max

  • + Realistic glass thickness and reflections, including internal reflections of the sphere.
  • + Strong adherence to all spatial requirements of the prompt.
  • + Dynamic lighting with visible sunbeams across the table.
  • The sphere and the book both appear to be floating or clipping without clear points of contact.
  • The plant in the background has a slightly artificial, plastic-like texture compared to Model A.

Verdict: Both models followed the prompt instructions perfectly, including the specific spatial arrangement of the four main objects. FLUX.1 Kontext [pro] is preferred for its superior photographic quality and natural depth of field, whereas Qwen Image Max, while catching better glass physics and light rays, suffers from objects looking like they are hovering unnaturally.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 Kontext [pro]
Qwen Image Max

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Natural, lifelike skin texture and lighting
  • + Elegant engraving on the plate armor
  • + Coherent composition with a smooth bokeh background
  • Largely missed the 'braided with small beads' requirement, showing only one braid and one bead
  • The 'battle-worn' aspect is very subtle compared to the prompt

Qwen Image Max

  • + Excellent adherence to all prompt details, including scars, dirt, and numerous small beads in hair
  • + Great textural contrast between the worn leather and intricate armor
  • + Strong atmospheric lighting that clearly shows the torch source
  • Slightly more artificial appearance in the sparks/skin compared to Model A
  • Composition is a bit cluttered with the torch very close to the head

Verdict: Qwen Image Max followed the prompt much more accurately, capturing the specific details like scars, dirt, and beads that FLUX.1 Kontext largely overlooked. While FLUX.1 Kontext produced a softer, more conventionally 'beautiful' portrait, Qwen Image Max better realized the gritty, battle-worn paladin archetype requested.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 Kontext [pro]
Qwen Image Max

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent typography rendering with clean, glowing effects
  • + High appetizing quality with realistic melted cheese textures
  • + Strong vertical composition that feels professional
  • Missed the 'starburst' container for the price
  • Redundant price tags (includes price twice)
  • The 'exploded' effect is more of a minor separation than a dynamic explosion

Qwen Image Max

  • + Perfect adherence to the 'starburst' requirement for the price
  • + Excellent sense of motion with dynamic angles and debris
  • + Vibrant, fiery text effects that match the background atmosphere
  • The 'MAGIC BURGER' text has slight irregularities in the flame mask
  • Logo placement is a bit crowded near the top edge

Verdict: Qwen Image Max followed the complex formatting instructions more accurately, specifically by including the starburst for the price and creating a much more dynamic 'exploded' feel. While FLUX.1 Kontext [pro] produced a very appetizing burger with clean text, it failed on the starburst instruction and included the price twice, making the ad layout feel less intentional.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 Kontext [pro]
Qwen Image Max

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent photographic lighting and skin/fur textures
  • + Accurately represents the passenger and capybara in a cinematic style
  • + The capybara's expression is very calm and fits the prompt well
  • The passenger is holding a phone to her ear like a call rather than looking at it
  • The capybara's paw on the steering wheel looks slightly distorted

Qwen Image Max

  • + Perfect adherence to the requested interior perspective showing the dashboard and city street
  • + High level of detail on the capybara's clothing and the taxi interior
  • + The passenger is correctly looking down at her phone as requested
  • The hands driving the car appear more primate-like/humanoid than capybara paws
  • The capybara's head is disproportionately large compared to its body and the car seat

Verdict: Both models followed the prompt well, but Qwen Image Max captured the entire scene better, including the requested city street background and the bored businesswoman looking at her phone. FLUX.1 Kontext [pro] has slightly more realistic photographic quality and better texture on the capybara, but the composition is tighter and misses the 'looking at her phone' detail for the passenger.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 Kontext [pro]
Qwen Image Max

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent typography for the main title and banner text.
  • + Atmospheric lighting with the glowing jack-o-lantern reflected on the ground.
  • Added nonsensical hallucinatory text ('Your: Vorkleat: Iight & Spans') not in the prompt.
  • The 'parchment' look is very dark, making it look more like a digital painting than a poster.

Qwen Image Max

  • + Perfectly followed all text instructions with zero typos or additions.
  • + Clearly defined parchment texture and a beautiful thorn/web border.
  • + Excellent composition that balances the central pumpkin with the twisted trees and bats.
  • The border is a bit repetitive in its thorn pattern.
  • Lighting is slightly less dramatic than Model A.

Verdict: Qwen Image Max is the winner as it followed all specific text instructions perfectly without adding the nonsensical hallucinations found in FLUX.1 Kontext. While FLUX.1 Kontext has arguably more cinematic lighting, Qwen Image Max captured the 'parchment' look and the specific gothic border elements much more effectively.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 Kontext [pro]
Qwen Image Max

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent typography with a charming 3D bubble effect
  • + Clean and simple composition that feels like a true miniature
  • + Smooth, refined 'cartoon' textures that match the prompt perfectly
  • The sushi rice looks somewhat like pebbles or spheres rather than individual grains

Qwen Image Max

  • + High variety of sushi types and realistic food textures
  • + Sophisticated glass-like base material rendering
  • + Strong adherence to the requested isometric 45-degree angle
  • The 'JAPAN' text is slightly clipped at the top
  • Composition feels a bit more crowded compared to the 'minimal' request

Verdict: Both models followed the prompt very well, but FLUX.1 Kontext [pro] captured the 'cartoon scene' and 'miniature' aesthetic more effectively with its soft, playful textures and perfectly centered layout. Qwen Image Max produced more realistic and detailed sushi models, but the minor text cropping and slightly busier composition make it feel less like a polished diorama compared to the FLUX.1 Kontext [pro] output.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 Kontext [pro]
Qwen Image Max

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent soft lighting and atmospheric god rays.
  • + Matches the 'big expressive eyes' and 'ultra-detailed fur' descriptions perfectly.
  • + Consistent, high-quality artistic style across all characters.
  • Failed to include the rabbit, instead generating two cat-like creatures in the middle.
  • The animals are sitting rather than 'tumbling' or 'chasing'.

Qwen Image Max

  • + Successfully captured the 'tumbling' and 'playful' interaction requested.
  • + Includes various butterflies which matches the 'chasing' prompt.
  • + Good anatomical variety between the puppy, kitten, and fox.
  • Failed to include the rabbit, instead generating two retriever puppies.
  • The fur textures are slightly more coarse and less 'fluffy' than Model A.
  • Some artifacts present, such as the fox's floating paw near the puppy's head.

Verdict: FLUX.1 Kontext [pro] offers superior aesthetic quality with beautiful lighting and textures, but it failed to capture the requested action and substituted the rabbit for a second kitten. Qwen Image Max follows the complex interaction of 'tumbling' much better, but it lacks the fine fur detail of the first model and also missed the rabbit in favor of a duplicate puppy. FLUX.1 Kontext is slightly preferred for its stunning lighting and higher overall visual polish.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 Kontext [pro]
Qwen Image Max

AI Judge Analysis

FLUX.1 Kontext [pro]

  • + Excellent typography rendering with the correct 'Caffè' spelling and accent
  • + Clean vector aesthetic that perfectly matches the 'minimalist' prompt requirement
  • + Sophisticated use of subtle texture on the paper background
  • Included a typo in the banner, writing 'EEST.' instead of 'EST.'
  • The steam icon is very simplified, appearing almost like an afterthought

Qwen Image Max

  • + Beautifully detailed cloche and steam illustration
  • + Correct spelling for all text elements including 'Est. 1720'
  • + Rich vintage color palette and convincing parchment texture
  • Less 'minimalist' than Model A, leaning more towards a detailed illustration
  • Text alignment is slightly cramped between the cloche base and the banner

Verdict: FLUX.1 Kontext [pro] captures the minimalist vector emblem style more accurately but suffers from a spelling error in the banner. Qwen Image Max follows the text instructions perfectly and provides a more visually impressive illustration, although it is less 'minimalist' than requested. Qwen Image Max is the preferred choice for its professional finish and technical accuracy in text rendering.

Next steps

Explore each model

The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.