Black Forest Labs' enhanced 12-billion parameter flow transformer with 6x faster generation than FLUX.1 [pro], delivering superior composition, detail, and artistic fidelity
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX1.1 [pro]
#50 of 62 in Text-to-Image
Qwen Image
#40 of 62 in Text-to-Image
Where the votes landed
FLUX1.1 [pro]
0%
win rate
Ties
0%
Qwen Image
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent realism with highly detailed glass reflections and textures.
- + The plant is clearly visible through the glass as requested.
- + High clarity and sharpness on the red book's fabric texture.
- − The glass object is a tall rectangular prism rather than a cube.
- − The blue sphere is quite large relative to the container, deviating from 'small sphere'.
Qwen Image
- + Correctly interprets the 'cube' geometry and the 'small' scale of the sphere.
- + Accurate composition with the plant partially visible behind the glass.
- + Natural lighting and balanced layout follow the prompt's instructions well.
- − The bottom of the glass cube has an odd mirrored/opaque base rather than being fully transparent glass.
- − Lower overall resolution and detail compared to Image A.
Verdict: Both models followed the prompt successfully, including the complex request of showing a plant through glass. FLUX1.1 [pro] produced a much more photorealistic image with superior textures, but it failed to generate a cube (producing a tall rectangle instead). Qwen Image followed the spatial and geometric instructions more accurately, providing a true cube and a smaller sphere, though its lighting and textures are slightly less sophisticated.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent atmospheric lighting and realistic pavement reflections
- + Follows the motion blur request for passing cars well
- + High level of detail on the man's skin and clothing texture
- − The bicycle structure is anatomically broken, missing the rear half of the frame
- − The man's posture is awkward and doesn't clearly show 'repairing'
Qwen Image
- + The bicycle structure is far more coherent and complete
- + Captures the action of 'repairing' the seat/post effectively
- + Clean composition with good street atmosphere
- − Lacks the requested motion blur for the cars
- − Rain effect looks slightly more artificial/added-on compared to the lighting in image A
Verdict: FLUX1.1 [pro] creates a much more cinematic and atmospheric image with superior textures and lighting, but it fails significantly on the physical logic of the bicycle. Qwen Image provides a much more coherent subject (the bike) and better illustrates the specific action, making it a more successful 'photo' despite being slightly less atmospheric. Qwen Image is the winner because the broken bicycle in FLUX1.1 [pro] is a major visual distractor.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX1.1 [pro]
- + Extremely realistic skin texture and lifelike eyes.
- + Natural integration of bokeh sparks into the lighting environment.
- + Consistent and subtle battle-worn aesthetic.
- − The braids and beads are very difficult to see and lack detail.
- − Less of the ornate armor is visible due to the extreme close-up crop.
Qwen Image
- + Excellent adherence to the 'braided hair with small beads' prompt with colorful details.
- + Highly detailed engraving on the plate armor and clear leather straps.
- + Strong composition that includes more of the character and the light source.
- − The sparks have a very digital, artificial look like starburst filters.
- − The facial scars look a bit painted on rather than integrated into the skin.
- − Some anatomical issues with hair emerging through the metal pauldron.
Verdict: FLUX1.1 [pro] excels in raw realism and cinematic lighting, creating a much more believable person, though it missed the prominence of the requested beads. Qwen Image captured every specific detail of the prompt, including the beads and ornate engravings, but the final output looks more like a high-end video game render than a real photograph.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent presentation for a menu with realistic flat-lay cutlery.
- + Complex layout that shows multiple pages of a menu for better context.
- − Text rendering is mostly gibberish or illegible at smaller scales.
- − The 'grid' for food photos is a bit cluttered and lacks the minimalist feel requested.
Qwen Image
- + Perfect adherence to the 'grid' structure for food photos.
- + Bold, legible sans-serif headers that closely match the requested design aesthetic.
- + Stronger minimalism with clean use of white space and vibrant color blocks.
- − The placeholder text 'Pizzaurant' is a bit literal and nonsensical.
- − The food photos look slightly more AI-generated/cartoonish compared to the realistic textures in the other model.
Verdict: Qwen Image is the winner because it better captures the 'minimalist' and 'grid' aspects of the prompt with a layout that looks like a real modern graphic design piece. While FLUX1.1 [pro] provides a more complex scene with cutlery, it fails to deliver the specific grid-based modern design requested, and its text is much less coherent.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent photographic quality and texture on the burger patties and bun.
- + Clean, modern typography that is easy to read.
- + Accurate rendering of the price at the bottom.
- − Failed the 'exploded' instruction as the burger remains largely intact.
- − Failed to place the price in a starburst as requested.
- − The glowing effect is blue rather than fiery as requested in the prompt.
Qwen Image
- + Followed the 'exploded' instruction better with visible gaps between components.
- + Successfully included the price in a glowing starburst element.
- + Adhered well to the fiery glowing effect for both the text and background.
- − The burger ingredients look slightly more illustrative compared to the photorealism of Model A.
- − Minor artifacts in the sauce splatters.
- − The price text is slightly less crisp than the main title.
Verdict: Qwen Image is the winner because it followed all the complex prompt instructions, including the exploded burger layout and the specific starburst price tag, which FLUX1.1 [pro] missed. While FLUX1.1 [pro] has slightly higher realism in the textures of the meat, it failed to execute the core creative concepts of the fiery glow and the exploded view.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI judge analysis unavailable for this challenge.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent cinematic lighting and dynamic composition.
- + High level of detail in the horse's fur and the astronaut's suit texture.
- − Completely failed the logical constraint of the prompt (horse on top).
- − The horse's body has an anatomical merging error where the harness disappears into the chest.
Qwen Image
- + Clean, high-resolution rendering with a clear space background.
- + Includes a realistic depiction of Earth and moon as part of the composition.
- − Completely failed the spatial instruction for the horse to be on top.
- − The horse's legs have illogical anatomical orientation and hooves appear deformed.
Verdict: Both models failed the specific prompt instruction to place the horse on top of the astronaut, instead providing the cliché astronaut-riding-a-horse image. FLUX1.1 [pro] is the better image overall due to its cinematic lighting and artistic flair, whereas Qwen Image appears more generic and has significant anatomical issues with the horse's legs.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent atmospheric lighting and bokeh
- + Beautifully rendered textures on the capybara fur
- − Failed to place the capybara in the driver's seat, placing both subjects in the back seat instead
- − Missed the requirement of both paws on the steering wheel
Qwen Image
- + Perfect adherence to all prompt elements including placement, attire, and action
- + High visual clarity and convincing photorealistic style
- + Correct anatomical placement of paws on the steering wheel
- − The taxi sign on top of the car has a typo ('YOXI')
- − The capybara's hands look slightly more primate-like than typical capybara paws
Verdict: While FLUX1.1 [pro] produced a more cinematic and artistic image, it failed significantly on the spatial requirements of the prompt by placing the driver and passenger in the same row. Qwen Image followed every instruction perfectly, correctly positioning the capybara behind the wheel and the woman in the back, making it the clear winner for prompt adherence.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent internal cinematic lighting and atmospheric depth in the illustration
- + Clean thorn border that frames the central artwork well
- + Captures the moody night sky with subtle moon and bat placement
- − Significant text repetition and typos (e.g., 'You inivted to a n night')
- − Layout is vertically skewed despite the square format request
- − Font choice for event details is a generic sans-serif which clashes with the gothic theme
Qwen Image
- + Stronger adherence to the 'vintage parchment' and 'gothic border' elements
- + Text is much more legible and formatted correctly for an invitation
- + Uses appropriate gothic typography that matches the visual style
- − Minor spelling error in the primary header ('HalleParty')
- − The bats look slightly more cartoonish compared to the cinematic lighting of the background
- − The scroll banner is slightly off-center
Verdict: Qwen Image is the better invitation because it successfully integrates the requested gothic typography and layout structure, whereas FLUX1.1 [pro] suffers from heavy text repetition and poor font selection. While FLUX1.1 [pro] has a more sophisticated illustrative style for the jack-o-lantern, Qwen Image captures the 'vintage parchment' texture and spiderweb border far more effectively.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent typography rendering with consistent spacing
- + High variety of sushi types adding visual interest
- + Refined PBR textures with a soft subsurface scattering look on the fish
- − Failed to include the requested flag icon next to the text
- − The scale of the garnish (tomatoes) is slightly out of place for a sushi dish
Qwen Image
- + Successfully included the flag icon as requested in the prompt
- + Accurate 45-degree isometric projection and composition
- + Clean, toy-like aesthetic that fits the 'miniature 3D cartoon' description well
- − Lower complexity in sushi design compared to Model A
- − Text is slightly less refined in terms of kerning and character weight
Verdict: Both models followed the prompt very closely, producing high-quality isometric dioramas. Model A (FLUX1.1 [pro]) has superior rendering on the food materials and more variety in the sushi, but Qwen Image followed the instructions more accurately by including the requested flag icon and adhering more strictly to the 'minimal' garnish constraint.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX1.1 [pro]
- + Beautiful backlighting and bokeh effect.
- + Excellent texture on the puppy's fur and expressive eyes.
- + Cohesive warm color palette.
- − Failed to include all requested animals (missing the fox).
- − The butterflies are glowing symbols rather than realistic insects.
- − The animals are sitting still rather than 'tumbling and chasing'.
Qwen Image
- + Followed the prompt instructions more accurately by including all four animals: puppy, kitten, bunny, and fox.
- + Better adherence to the 'chasing and tumbling' action requested.
- + Successfully captured various lighting elements like god rays and dew sparkles.
- − The fox's anatomy looks slightly more like a plush toy than a real kit.
- − The background blurring is a bit less sophisticated than Model A.
Verdict: Model B (Qwen Image) is the winner because it successfully included all four requested animals (puppy, kitten, bunny, and fox), whereas FLUX1.1 [pro] missed the fox entirely. Qwen Image also better captured the active 'chasing' vibe of the prompt and included more accurate butterfly renderings and dew drops, despite FLUX1.1 [pro] having slightly more realistic fur textures on the characters it did generate.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent ornamental detail and vector style clarity
- + Elegant balanced composition with a classic badge feel
- + Accurate rendering of the secondary text 'Est. 1720'
- − Failed to spell the main brand name correctly, rendering 'Franilian' instead of 'Florian'
Qwen Image
- + Successfully captured the minimalist cloche dome icon
- + Strong texture application on the background
- + Good color palette adherence following the warm brown and cream tones
- − Severely jumbled and cluttered typography for 'Caffè Florian'
- − The layout is less cohesive compared to a traditional vector emblem
Verdict: While FLUX1.1 [pro] failed the primary spelling of 'Florian', its overall aesthetic, composition, and professional vector lines are significantly higher quality. Qwen Image captured the minimalist cloche better but produced illegible, overlapping text that makes the logo unusable for branding purposes.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent visual layout and hierarchy that feels like a professional infographic.
- + Adheres closely to the colors and flat-vector style requested in the prompt.
- + Produces a large amount of visual content including multiple callouts and a central graphic.
- − Text consists largely of gibberish and misinterpreted labels (e.g., 'Earna Orbit').
- − Includes Saturn (planet with rings) instead of icons for orbital paths, which is scientifically inaccurate.
Qwen Image
- + Text is much more legible and contains actual names and mission terminology.
- + The iconography for the Earth, Moon, and Lunar Module is clear and distinct.
- + Correctly followed the color palette instructions.
- − Failed to include all 6 requested steps, stopping effectively at step 4.
- − Included meta-text from the prompt '(Stop at landing)' inside the final image design.
- − Layout is a bit cramped and lacks the sophisticated design balance seen in the other model.
Verdict: FLUX1.1 [pro] created a much more visually appealing and professional-looking infographic layout that captures the requested aesthetic, though the text content is mostly nonsense. Qwen Image produced much clearer, more legible text and name labels, but failed to complete the sequence of requested steps and included the prompt's instructions within the graphic. FLUX1.1 [pro] is the winner for its superior composition and adherence to the layout style requested.
Explore each model
Alibaba's Qwen image model