Black Forest Labs' 12 billion parameter distilled image generation model optimized for speed, capable of generating high-quality images in just 4 inference steps
Settled by community votes across 8 shared challenges, with an AI judge weighing in on each.
FLUX.1 [schnell]
#48 of 62 in Text-to-Image
Qwen Image Max
#35 of 62 in Text-to-Image
Where the votes landed
FLUX.1 [schnell]
0%
win rate
Ties
0%
Qwen Image Max
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent handling of glass thickness and internal reflections.
- + High aesthetic quality with vibrant colors and sharp details.
- − Failed the spatial prompt by adding an extra blue sphere on top of the book.
Qwen Image Max
- + Perfect adherence to the spatial requirements of the prompt.
- + Realistic soft window light and shadows on the wooden table.
- + Accurate visibility of the plant through the glass panels.
- − The sphere appears to be floating rather than resting at the bottom, which may be a slight physical inconsistency depending on interpretation.
Verdict: Qwen Image Max is the winner because it followed every spatial instruction in the prompt perfectly. FLUX.1 [schnell] produced a visually striking image but failed the logical constraints of the prompt by placing an unnecessary second sphere on top of the book.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.1 [schnell]
- + Extremely high-quality skin texture and lifelike eyes.
- + Compelling lighting with strong highlights from the torchlight.
- − Missed the 'bokeh sparks' and 'small beads' in the hair.
- − The crop is perhaps too close, obscuring the detail of the leather straps and cloth.
Qwen Image Max
- + Excellent adherence to all prompt details including beads, scars, dirt, and sparks.
- + Beautifully rendered ornate engraving on the plate armor and textured leather straps.
- + Captures the 'battle-worn' aesthetic perfectly through his expression and skin texture.
- − The sparks are slightly uniform in shape, appearing a bit like a filter.
Verdict: While FLUX.1 [schnell] creates a stunningly realistic face, Qwen Image Max is the superior image for this prompt as it includes every specific detail requested, such as the beads in the braids, the bokeh sparks, and the intricate texture of the leather and ornate armor. Qwen Image Max better interprets the 'battle-worn' theme and provides a more balanced composition for a portrait of this type.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent photorealistic texture on the burger bun and patty.
- + Deep, rich color palette with nice lighting on the cheese.
- − Failed the primary text prompt, spelling it 'AGIC BURGER'.
- − The prices are repetitive and messy, including a '€699' error.
- − The burger is mostly assembled rather than 'exploded' into individual components.
Qwen Image Max
- + Perfect text adherence with correct spelling and glowing fiery effects.
- + Captures the 'exploded' motion effect much better with separate components flying outward.
- + Highly professional ad composition with a clear focal point and starburst.
- − The 'LIMITED TIME ONLY' text is slightly less fiery than the main title.
- − One small piece of lettuce looks slightly flat compared to the rest of the image.
Verdict: Qwen Image Max is the clear winner as it followed every instruction, including complex text rendering and the 'exploded' layout. FLUX.1 [schnell] failed on basic spelling and price accuracy, and produced a mostly whole burger despite the prompt requesting a deconstructed view.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent character expressions, perfectly capturing the requested 'bored' look on the passenger.
- + Strong lighting coherence and high-quality rendering of fur and phone.
- + Clear and legible 'TAXI' text on the hat.
- − The capybara's paws are not both on the steering wheel, as one is resting on its lap.
- − The perspective makes the capybara look overly large in comparison to the car frame.
Qwen Image Max
- + Successfully placed both paws on the steering wheel as requested.
- + Highly detailed and realistic car interior, including a textured dashboard and worn ceiling.
- + Excellent composition that shows more of the Manhattan street environment.
- − The hands on the steering wheel look more like primate hands than capybara paws.
- − The human passenger's face is slightly distorted and less clear than the main subject.
Verdict: Qwen Image Max is the winner for its superior composition and taxi interior realism, accurately placing both 'paws' on the wheel as the prompt required. While FLUX.1 [schnell] captures the specific bored expression of the passenger better, it fails the technical instruction regarding the placement of the paws and has a more cramped composition.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Strong cinematic lighting with a vibrant glowing jack-o-lantern.
- + Captures the moody atmosphere with stylized trees and bats well.
- − Significant text errors including misspelled words and repeated lines.
- − Instruction for a 'small scroll banner' was misinterpreted as multiple generic banners.
Qwen Image Max
- + Excellent adherence to all text requirements with perfect spelling.
- + High visual quality on the border details, including the requested thorns and webs.
- + Superior gothic aesthetic that feels like a polished vintage poster.
- − The jack-o-lantern is a bit more realistic and less 'central' in terms of light spill compared to Model A.
Verdict: Qwen Image Max is the clear winner as it followed every specific text instruction perfectly, whereas FLUX.1 [schnell] struggled with severe typos and hallucinated extra text fields. Qwen Image Max also delivered a more intricate border and a more authentic vintage gothic aesthetic.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.1 [schnell]
- + Clean aesthetic with a very simple, minimalist diorama base
- + Accurate 45-degree isometric projection
- + Included the flag icon as requested
- − Failed to include the word 'SUSHI' in the text
- − The sushi piece is a strange hybrid of nigiri and a roll
- − Text rendering is slightly faint and lacks the requested boldness
Qwen Image Max
- + Perfect text rendering for both 'JAPAN' and 'SUSHI'
- + High-quality PBR material rendering, especially on the glass base and fish textures
- + More visually appealing variety of sushi pieces
- − The diorama base is a bit small for the plate size
- − Composition feels slightly more crowded than the requested 'minimal' scene
Verdict: Qwen Image Max is the clear winner as it followed all text instructions perfectly, including both requested words and the flag icon. While FLUX.1 [schnell] captured a cleaner minimalist vibe, it failed to render the word 'SUSHI' and produced a confusingly anatomical piece of food. Qwen Image Max also demonstrated superior material work with realistic lighting and textures.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.1 [schnell]
- + Excellent fur texture and soft lighting
- + Vibrant, high-contrast colors following the sunset theme
- + Very clean composition with no major anatomical distortions
- − Missed the baby bunny entirely, replacing it with a second cat-hybrid creature
- − Lacks the 'tumbling together' action described in the prompt
- − Animals are mostly static rather than chasing butterflies
Qwen Image Max
- + Included all requested animals: puppy, kitten, fox, and the puppy acts as a substitute for the bunny or there is a tumbling group effect
- + Captured the 'god rays' and 'tumbling together' action much more effectively
- + Better representation of butterflies interacting with the animals
- − The fox's anatomy is slightly awkward with thin limbs
- − The golden retriever puppy is duplicated instead of including a distinct bunny
- − The kitten has an extra-long, slightly unnatural tail
Verdict: Qwen Image Max is the winner because it successfully captured the dynamic action of the prompt, showing the animals actually tumbling and interacting with the butterflies under clear god rays. While FLUX.1 [schnell] has slightly superior fur rendering, it failed the specific prompt requirements by omitting the baby bunny and posing the animals statically.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.1 [schnell]
- + Clean vector aesthetic
- + Symmetrical composition
- − Significant spelling error ('FRAMILAN' instead of 'Florian')
- − Incorrect year ('7720' instead of '1720')
- − Missing the requested steam element
Qwen Image Max
- + Perfect text rendering for both the name and the date
- + Includes all requested elements including the steam and banner
- + Excellent vintage texture and shading
- − The parchment background is slightly more distressed than a 'subtle texture'
Verdict: Qwen Image Max followed every specific detail of the prompt, including the complex text and the steam effect. FLUX.1 [schnell] failed significantly on the text, outputting 'FRAMILAN' and '7720', making it unusable for the requested logo.
Explore each model
The Max series of Tongyi Qwen’s image generation model excels across a wide range of generation tasks. Compared with the Plus series, it significantly reduces the “AI-like” feel in generated images, enhancing their realism. It delivers more lifelike material textures for human subjects, finer and more detailed natural textures, and more visually appealing text rendering.