Head to head
Esc

Models · slot A

to navigate to pick

FLUX.1 Kontext [dev] Black Forest Labs Qwen Image Max Alibaba

Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.

FLUX.1 Kontext [dev]

17.2 arena score

#54 of 62 in Text-to-Image

Skill signature · Text-to-Image

Qwen Image Max

21.3 arena score

#35 of 62 in Text-to-Image

Vote tally

Where the votes landed

FLUX.1 Kontext [dev]

0%

win rate

Ties

0%

Qwen Image Max

0%

win rate

Shared challenges 8

Challenge by challenge

The strongest take from each model on every shared challenge, with the AI judge's read.

Geometric Composition

Text-to-Image

“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”

FLUX.1 Kontext [dev]
Qwen Image Max

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent adherence to lighting instructions (soft light from left)
  • + Beautifully rendered reflections and transparency in the glass
  • + Realistic book texture and binding
  • The glass cube has no bottom face, appearing as a five-sided shell

Qwen Image Max

  • + Perfect visual logic with the blue sphere appearing to float inside a solid cube
  • + High-quality glass edge detailing
  • + Strong composition and clear table edge
  • The plant is partially behind but mostly to the side of the cube
  • The light source feels more overhead than from a window on the left

Verdict: Both models followed the complex spatial instructions perfectly. FLUX.1 Kontext [dev] captured the requested soft window lighting much better, while Qwen Image Max created a more structurally convincing glass cube. FLUX.1 Kontext [dev] is the winner for its superior atmospheric quality and more accurate interpretation of the 'partially visible through glass' instruction for the plant.

Fantasy Warrior

Text-to-Image

“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”

FLUX.1 Kontext [dev]
Qwen Image Max

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Strong cinematic lighting with realistic reflections on the armor.
  • + Highly detailed engraving on the plate armor.
  • + Lifelike eyes and skin texture.
  • Missed the request for braided hair with beads.
  • The skin looks a bit too clean for a 'battle-worn' character.

Qwen Image Max

  • + Excellent adherence to the 'braided hair with beads' prompt detail.
  • + Superior texture on leather straps and distressed cloth underlayers.
  • + Visually distinct scars and dirt that convey 'battle-worn' more effectively.
  • The torch in the background is a bit distracting and less out-of-focus than requested.
  • Some of the bokeh sparks look like static noise rather than light orbs.

Verdict: Qwen Image Max followed the complex prompt more accurately, specifically including the braided hair with beads and the weathered textures on the leather and cloth. While FLUX.1 Kontext [dev] produced a very high-quality cinematic portrait, it ignored several key descriptive elements of the character's appearance.

Magic Burger Explosion: Fiery Photorealism Challenge

Text-to-Image

“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”

FLUX.1 Kontext [dev]
Qwen Image Max

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent legibility of the main title text
  • + High-quality rendering of the charcoal and flame background
  • + Perfectly followed the 'starburst' request for the price
  • Failed the 'exploded' instruction as the burger is fully assembled
  • Spelling error in the secondary text ('LNHLY' instead of 'ONLY')
  • The burger lighting is slightly flat compared to the background

Qwen Image Max

  • + Successfully captured the 'exploded'/dynamic suspended motion
  • + Perfect text rendering for all requested phrases with no spelling errors
  • + Superior integration of the fiery glow effect on the typography
  • The 'starburst' for the price is more of a light flare than a graphic starburst
  • Some sauce and lettuce artifacts appear slightly messy during the 'explosion'

Verdict: Qwen Image Max is the clear winner as it successfully interpreted the 'exploded' requirement and maintained perfect text accuracy. FLUX.1 Kontext [dev] produced a static burger and failed to spell 'ONLY' correctly, whereas Qwen Image Max delivered a dynamic, high-fidelity ad with impressive fiery effects.

The Capybara Taxi Driver

Text-to-Image

“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”

FLUX.1 Kontext [dev]
Qwen Image Max

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent fur texture rendering on the capybara.
  • + The passenger's bored expression perfectly matches the prompt.
  • + Atmospheric lighting that creates a realistic nocturnal urban mood.
  • Failed the instruction to have both paws on the steering wheel.
  • The capybara's head shape is slightly stylized, looking more like a squirrel or groundhog.

Qwen Image Max

  • + Followed the instruction to have both paws on the steering wheel.
  • + The capybara's anatomical features are more accurate to the species.
  • + Higher levels of detail in the taxi interior and the New York street background.
  • The capybara's hands look slightly more primate-like than capybara-like paws.
  • The taxi interior shows some minor deterioration artifacts on the ceiling.

Verdict: Both models captured the surreal yet mundane nature of the prompt very well. Qwen Image Max is the winner for better prompt adherence regarding the placement of the paws on the steering wheel and a more accurate representation of a capybara's facial structure, whereas FLUX.1 Kontext [dev] struggled with the specific paw placement and species likeness.

The Halloween Invitation

Text-to-Image

“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”

FLUX.1 Kontext [dev]
Qwen Image Max

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Strong, high-contrast title text
  • + Vibrant, glowing jack-o-lantern central focus
  • Banner text is distorted and mostly illegible
  • The location name 'The Arches' is misspelled
  • Lacks the requested parchment texture and night sky

Qwen Image Max

  • + Excellent text rendering with no spelling errors
  • + Accurately represents all prompt elements including parchment and night sky
  • + Sophisticated gothic aesthetic with detailed thorns and webs
  • Central pumpkin lighting is slightly less vibrant than Model A

Verdict: Qwen Image Max followed every instruction perfectly, producing a cohesive design with a beautiful parchment texture and flawless text. FLUX.1 Kontext [dev] struggled significantly with the banner text and location name, and failed to include the requested moody night sky and parchment background.

Isometric Miniature Diorama Scenes

Text-to-Image

“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”

FLUX.1 Kontext [dev]
Qwen Image Max

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent clay-like soft textures that match the cartoon scene request.
  • + Clean and bold text rendering for both JAPAN and SUSHI.
  • + Simple and effective isometric composition.
  • The flag icon is a strange abstract shape rather than a recognizable Japanese flag.
  • The sushi design is overly simplified to the point of looking like a toy block.

Qwen Image Max

  • + High visual quality with realistic PBR material textures for the fish and glass base.
  • + Accurate representation of the Japanese flag icon.
  • + Professional typography and layout that feels like a finished graphic design.
  • Slightly more complex than the 'minimal' garnish requested.
  • The shadows under the text are a bit heavy compared to the soft lighting of the scene.

Verdict: Qwen Image Max is the winner as it perfectly follows all prompt instructions, including the specific flag icon, while providing superior PBR materials and a more appealing variety of sushi. FLUX.1 Kontext [dev] captured the 'cartoon' feel well but failed on the flag icon and produced a very basic interpretion of the subject matter.

Adorable Baby Animals in Sunny Meadow

Text-to-Image

“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”

FLUX.1 Kontext [dev]
Qwen Image Max

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Excellent sense of motion and action with the animals mid-pounce
  • + Dynamic backlighting with soft bokeh effect in the background
  • + Very clean, high-resolution textures on the puppy's fur
  • Failed to include the baby bunny and red fox kit, showing multiple kittens instead
  • Lacks the variety of wildflowers requested, focusing mostly on daisies

Qwen Image Max

  • + Successfully included all requested animal types: puppy, kitten, fox, and bunny (partially visible/small)
  • + Beautiful lighting with visible god rays and a more diverse range of wildflowers
  • + Higher level of detail in the fur textures and butterfly variety
  • Anatomical issues with the standing dog's front paw being oddly placed on the other dog
  • The bunny is very small/hidden compared to the other animals

Verdict: While FLUX.1 Kontext [dev] creates a more dynamic action shot with better photographic clarity, it significantly fails the prompt by omitting several requested animals. Qwen Image Max follows the prompt much more accurately, including the fox, kitten, and puppy in a more lush and colorful meadow, despite some minor anatomical clipping in the legs.

Vintage Cafe Logo

Text-to-Image

“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”

FLUX.1 Kontext [dev]
Qwen Image Max

AI Judge Analysis

FLUX.1 Kontext [dev]

  • + Very clean minimalist aesthetic
  • + Perfect text rendering with the correct accent mark
  • + Clear vector emblem style suitable for modern branding
  • Missed the request for a banner element
  • Steam is overly simplified and less recognizable
  • Background texture is almost non-existent

Qwen Image Max

  • + Accurately included the banner for 'Est. 1720'
  • + Beautiful vintage paper texture and hatching details
  • + Cloche shape and steam are well-defined and artistic
  • The text 'Caffè' has a slightly inconsistent accent mark shape
  • Less 'minimalist' than Model A, opting for a more illustrative look

Verdict: Qwen Image Max is the winner as it followed every detail of the prompt, specifically the 'banner' and 'subtle texture' which FLUX.1 Kontext [dev] omitted. While FLUX.1 Kontext [dev] produced a very clean vector logo, Qwen Image Max better captured the requested vintage atmosphere and complexity while maintaining high text quality.

Next steps

Explore each model

The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.