Black Forest Labs' enhanced 12-billion parameter flow transformer with 6x faster generation than FLUX.1 [pro], delivering superior composition, detail, and artistic fidelity
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX1.1 [pro]
#49 of 62 in Text-to-Image
Qwen Image 2.0
#34 of 62 in Text-to-Image
Where the votes landed
FLUX1.1 [pro]
0%
win rate
Ties
0%
Qwen Image 2.0
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent photographic clarity and lighting
- + Realistic glass refraction and material textures
- + Accurate 1:1 aspect ratio composition
- − The glass container is a tall rectangular prism rather than a 'cube'
- − The plant in the background is quite blurry
Qwen Image 2.0
- + Strong adherence to all spatial prompt elements
- + Realistic lighting interaction on the table surface
- + Good glass transparency showing the plant behind
- − The glass geometry is slightly warped/inconsistent
- − Reflections on the side panels of the glass are physically confusing
Verdict: Both models followed the prompt instructions well, including the spatial arrangement of the sphere, book, and plant. FLUX1.1 [pro] produced a much cleaner, more aesthetically pleasing image with superior textures, though it failed to make the glass a true cube. Qwen Image 2.0 captured the 'cube' shape better but suffered from minor architectural glitches in the glass panels.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent cinematic atmosphere with realistic lighting and reflections.
- + Strong adherence to the 'light rain' and 'wet pavement' aspects of the prompt.
- + Captures the bokeh and 50mm lens look effectively.
- − Serious anatomical and structural errors: the man has three legs/feet and the bicycle is missing its back half.
- − The man is leaning on the bike rather than actively 'repairing' it.
Qwen Image 2.0
- + Excellent 'repairing' action that accurately reflects the prompt's subject matter.
- + Incredible skin texture and facial detail that feels very candid and non-stylized.
- + Good motion blur on the background car as requested.
- − The framing is very tight, clipping parts of the bike and the man's feet.
- − The rain effect is more subtle compared to Image A.
Verdict: While FLUX1.1 [pro] creates a more visually striking and cinematic atmosphere, it suffers from major AI artifacts including a third leg and a broken bicycle geometry. Qwen Image 2.0 provides a much more coherent and realistic interpretation of 'repairing' with superior skin textures and fewer structural errors, making it the more successful image overall.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX1.1 [pro]
- + Extremely high skin texture detail and realistic eyes
- + Excellent cinematic lighting with warm highlights
- + Natural-looking wet and messy hair
- − The 'braided' hair with beads is less distinct compared to the other model
- − Composition is very tight, cutting off much of the armor detail
Qwen Image 2.0
- + Perfect adherence to the beaded braid request
- + Shows more of the ornate plate armor and leather straps
- + Clearer representation of scars and battle-worn features
- − Skin texture appears somewhat plastic or over-sharpened compared to human skin
- − Hand anatomy on the sword hilt is slightly awkward
- − Lighting feels a bit flatter despite the fire in the background
Verdict: FLUX1.1 [pro] excels in photorealistic skin textures and lifelike eyes, creating a more emotionally resonant portrait. However, Qwen Image 2.0 followed the technical aspects of the prompt more closely, demonstrating the beaded braids and ornate armor details far more clearly. FLUX1.1 [pro] is the winner for its superior visual quality and realism, though it missed some specific decorative elements.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent professional mock-up presentation including cutlery and multiple pages.
- + Sophisticated layout that utilizes whitespace effectively.
- + Follows the hierarchical prompt instructions for specific menu sections.
- − Images within the grid are smaller and less appetizing than the competition.
- − The text, while formatted well, is also gibberish.
- − The 'appetizers/pizza/mains' sections are not clearly separated within the vertical layout.
Qwen Image 2.0
- + High-quality, realistic food photography in a clear grid.
- + Strong adherence to the 'bold sans-serif' and 'vibrant' prompt requirements.
- + Very clean and modern UI approach suitable for a digital menu or flyer.
- − Text rendering is garbled and incoherent upon closer inspection.
- − The grid layout is slightly repetitive in its pizza choices for different sections.
Verdict: Both models struggled with legible text, but Model B (Qwen Image 2.0) produced far superior food photography that felt more 'vibrant' and aligned with the professional casual dining aesthetic. FLUX1.1 [pro] created a more realistic scene of physical menu cards, but the internal layout of the menu itself was less cohesive and the food images were cluttered.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent photographic rendering of the burger texture.
- + Very clean typography and professional layout.
- + Atmospheric lighting and depth of field.
- − Failed the 'exploded burger' requirement, as the burger is mostly assembled.
- − Missing the required starburst for the price.
- − Text is plain white/blue rather than the requested fiery effect.
Qwen Image 2.0
- + Successfully applied the fiery, glowing effect to all text elements.
- + Strictly followed all prompt components including the starburst and exploded layout.
- + Good sense of motion with drips and rising smoke.
- − The price text inside the starburst lacks the fiery glow seen in the main title.
- − Lighting on the burger feels slightly flatter compared to Model A.
Verdict: Qwen Image 2.0 is the clear winner as it followed every specific instruction in the prompt, including the 'exploded' assembly, the starburst, and the fiery text effects. While FLUX1.1 [pro] produced a more realistic-looking burger, it failed to execute several key creative constraints of the advertisement layout.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX1.1 [pro]
- + Clean and legible text layout
- + Warm, inviting background atmosphere with bokeh lighting
- − Failed to render the full text for the third menu item correctly
- − Split the second menu item into two separate lines with incorrect prices ($228 and $98)
- − The chalk texture looks a bit like a digital font overlay
Qwen Image 2.0
- + Excellent adherence to the full text prompt, including the third menu item
- + Very realistic chalk texture with smudges and natural variations
- + Correct spelling and prices for all items
- − The composition is slightly tight on the left edge
- − The handwriting is a bit more 'neat' than 'elegant cursive' in the title
Verdict: Qwen Image 2.0 followed the complex text prompt much more accurately, correctly rendering every menu item and price while maintaining a very realistic chalk-on-blackboard texture. FLUX1.1 [pro] struggled with the logic of the menu, hallucinating extremely high prices and failing to complete the third item properly.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent cinematic lighting and environmental depth
- + Highly detailed textures on the astronaut suit and horse mane
- + Dynamic, powerful composition with the horse rearing up
- − Completely failed the negative constraint to have the horse on top of the astronaut
- − The horse's anatomy becomes messy where it connects to the saddle area
Qwen Image 2.0
- + Crystalline, sharp visual quality
- + Includes creative elements like floating water droplets
- + Clearer space background with Earth visible
- − Failed the specific spatial instruction for the horse to be on top
- − The horse's leg anatomy is slightly awkward and elongated
Verdict: Both models failed the specific prompt instruction to place the 'horse on top' of the astronaut, instead providing the standard astronaut-riding-horse interpretation. FLUX1.1 [pro] produced a much more cinematic and atmospheric image with superior lighting, whereas Qwen Image 2.0 felt more like a standard digital composite but offered cleaner overall lines.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent atmospheric lighting and bokeh effect for a cinematic New York feel.
- + High-quality fur texture and realistic lighting on the capybara.
- + The woman's expression perfectly captures the 'bored and normal' requirement.
- − The perspective makes the capybara look more like a passenger than a driver from the interior angle.
- − Missing the steering wheel and the 'paws on wheel' detail.
Qwen Image 2.0
- + Successfully includes both front paws on the steering wheel as requested.
- + Realistic taxi exterior and window reflections.
- + Captures the interaction of the capybara wearing the driver's cap clearly from a side profile.
- − The woman appears to be in the front passenger seat rather than the back seat.
- − The blue car interior looks less like a standard New York yellow taxi.
- − The capybara's hands look slightly more primate-like than typical capybara paws.
Verdict: Both models captured the surreal nature of the prompt well, but they succeeded in different areas. FLUX1.1 [pro] produced a more aesthetically pleasing, cinematic image with better lighting and professional expressions, but it failed to show the capybara actually driving. Qwen Image 2.0 followed the specific structural instructions much better (paws on wheel, side profile), but it failed to place the passenger in the back seat as requested.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent cinematic lighting and atmosphere.
- + Beautifully detailed thorn border.
- + Vibrant glowing effect on the central jack-o-lantern.
- − Repetitive and redundant text with several typos and layout issues.
- − Banner text is partially illegible/misspelled.
- − Composition feels cramped with too much black space at the bottom.
Qwen Image 2.0
- + Perfect text rendering with zero spelling errors.
- + Accurately follows the 'Vintage gothic' and 'parchment' aesthetic requested.
- + Clean, balanced composition that fits all elements into the square format naturally.
- − Lighting is slightly flatter compared to Model A.
- − The jack-o-lantern feels a bit more like a stock asset than an integrated part of the scene.
Verdict: While FLUX1.1 [pro] excels in mood and lighting, it fails significantly on text accuracy and layout, producing redundant lines and typos. Qwen Image 2.0 provides a much better graphic design result, following all text instructions perfectly while maintaining a high-quality vintage aesthetic.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent adherence to the 'cartoon' and '3D isometric' miniature style requested.
- + Beautiful soft lighting and refined textures that match the PBR material prompt.
- + Text is well-integrated into the design aesthetic.
- − Failed to include the requested flag icon.
- − A few garnish elements like the tomatoes feel out of place for a traditional sushi scene.
Qwen Image 2.0
- + Successfully included all requested text and the Japan flag icon.
- + Very realistic food rendering with high clarity.
- + Correctly followed the 'small raised diorama base' instruction with a wooden block.
- − Completely ignored the '3D cartoon' and 'miniature' stylistic prompt, opting for realism instead.
- − Composition is slightly off-center compared to the prompt's request.
Verdict: FLUX1.1 [pro] followed the stylistic instructions for a 3D cartoon isometric scene much better than Qwen Image 2.0, which ignored the cartoon aesthetic for a realistic photo style. Although Qwen Image 2.0 was the only one to include the flag icon, FLUX1.1 [pro] is the preferred choice for accurately capturing the requested miniature diorama vibe.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX1.1 [pro]
- + Exceptional backlighting and lens flare effects that create a magical atmosphere
- + Extremely high level of detail in the fur texture and whiskers
- + Cohesive orange/golden color palette that enhances the 'warm golden sunrise' prompt
- − Failed to include the fox kit requested in the prompt
- − Missing the 'tumbling together' action, featuring static sitting poses instead
- − Butterflies are simplified glowing shapes rather than realistic insects
Qwen Image 2.0
- + Accurately included all four requested animals: puppy, kitten, bunny, and fox kit
- + Perfectly captures the 'tumbling together' and 'playfully chasing' action requested
- + Clearly defined god rays and realistic monarch-style butterflies
- − The fox kit has slightly strange anatomy where it connects with the kitten
- − The kitten's tail is missing or obscured in a way that looks a bit unnatural
- − The background rendering is slightly more cluttered than the soft bokeh of the competitor
Verdict: Qwen Image 2.0 is the clear winner for prompt adherence, successfully including all four specific animals and the 'tumbling' interaction that FLUX1.1 [pro] completely missed. While FLUX1.1 [pro] produced a more aesthetically polished and artistically lit single-frame portrait, it failed to fulfill the core requirements of the complex prompt.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent vintage woodcut aesthetic with beautiful frame details
- + Highly sophisticated vector emblem composition
- + Captures the 'vintage minimalist' texture and feel perfectly
- − Major spelling errors in the main name ('Cafe é Fratilian')
- − The steam is very subtle and barely visible
Qwen Image 2.0
- + Perfect adherence to text prompt spelling for 'Caffè Florian'
- + Clearer representation of the cloche and steam elements
- + Simple and clean layout
- − The 'Est. 1720' banner has a strange, disconnected curl on the right side
- − Styling is slightly more modern/clip-art than 'vintage minimalist'
- − Steam looks a bit like a flame
Verdict: While FLUX1.1 [pro] created a much more visually stunning and authentic vintage emblem, it failed significantly on the text spelling. Qwen Image 2.0 followed the text instructions perfectly and provided a clearer cloche dome, but the illustration quality is less sophisticated and has a minor rendering artifact on the banner.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX1.1 [pro]
- + Excellent navy, white, and muted red color palette.
- + Creative and balanced layout with high artistic flair.
- + Clean vector-style illustrative elements.
- − Poor text rendering with several spelling errors and nonsensical words.
- − Failed to follow the requested chronological numbering of steps.
- − Confusing iconography that doesn't clearly represent a Saturn V or specific mission stages.
Qwen Image 2.0
- + Highly accurate adherence to the specific 6-step sequence requested.
- + Clear and mostly correct typography for technical labels.
- + Excellent iconography including the Saturn V and Lunar Module.
- − Simple, vertical composition feels less like a professional 'infographic poster' and more like a basic list.
- − Minor spelling error in 'Translunjar'.
- − The white border around the image feels like a mock-up rather than a clean vector file.
Verdict: Qwen Image 2.0 is the clear winner for its superior prompt adherence, following the specific 6-step sequence and providing recognizable icons for the Saturn V and Lunar Module. While FLUX1.1 [pro] has a more sophisticated aesthetic and better color use, its inability to render legible text or follow the logical order of the mission stages makes it a failure as an infographic.
Explore each model
Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request