Black Forest Labs' open-weights image generation model with frontier performance, available for non-commercial local deployment
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.2 [dev]
#20 of 62 in Text-to-Image
Qwen Image 2.0
#34 of 62 in Text-to-Image
Where the votes landed
FLUX.2 [dev]
0%
win rate
Ties
0%
Qwen Image 2.0
0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.2 [dev]
- + Excellent photographic quality with realistic textures on the wood and glass
- + Accurate spatial arrangement with the plant clearly visible behind the cube through the glass
- + Perfect adherence to lighting instructions with logical reflections
- − The sphere is slightly smaller than what might be expected from the term 'small blue sphere' relative to the cube size
Qwen Image 2.0
- + Successfully includes all required elements: cube, sphere, book, and plant
- + Good texture on the book cover
- − The sphere appears to be floating mid-air inside the cube without physical support
- − The perspective and reflections on the glass panels are physically inconsistent
- − The plant looks more like a reflection or foreground element rather than being behind the cube
Verdict: FLUX.2 [dev] follows the prompt with much higher realism, providing a coherent physical scene where the sphere sits on the bottom of the cube and the plant is correctly refracted through the glass. Qwen Image 2.0 struggles with 3D space, resulting in a floating sphere and confusing reflections that make the cube look more like a series of mirrors.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.2 [dev]
- + Excellent adherence to technical prompts like motion blur and shallow depth of field
- + Highly realistic skin textures and wet surfaces
- + Captures the requested 'imperfect framing' effectively
- − The hands and bike handlebars have significant structural merging and AI artifacts
- − The red bicycle appears to have two sets of handlebars overlapping
Qwen Image 2.0
- + Natural and realistic pose for someone repairing a bicycle chain
- + Subtle and organic integration of rain droplets on the man's jacket
- + Clean facial details and skin texture
- − Failed to include the requested motion blur for passing cars
- − The reflection of the trees in the puddle does not match the sparse background trees
- − The background car has a strangely distorted front-end
Verdict: FLUX.2 [dev] followed the technical aspects of the prompt more closely, successfully implementing motion blur and a cinematic 50mm feel, though it struggled with the structural logic of the bicycle. Qwen Image 2.0 produced a very grounded and realistic character moment, but missed the specific 'motion blur' requirement. FLUX.2 [dev] is the slightly better choice for its successful atmospheric execution despite the minor anatomical/mechanical errors.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.2 [dev]
- + Excellent realization of 'warm torchlight' with realistic light falloff on the face and armor.
- + Highly detailed engraving on the plate armor and realistic leather strap texture.
- + Natural, high-quality skin textures with subtle 'faint scars' as requested.
- − The braids are somewhat simple compared to the 'small beads' request which are present but sparse.
- − The hair texture is slightly soft compared to the hyper-realism of the armor.
Qwen Image 2.0
- + Strong adherence to the 'hair braided with small beads' part of the prompt with intricate styling.
- + Effective 'battle-worn' look with a grit and intensity that fits the paladin archetype.
- + Good representation of the cloth underlayer and leather pauldrons.
- − Anatomical issues with the hand resting on the sword (distorted fingers and anatomy).
- − The 'warm torchlight' feels more like a large fire/explosion in the background, lacking the focus of a single light source.
- − Over-sharpened facial textures create a slightly artificial look.
Verdict: FLUX.2 [dev] produces a much more coherent and aesthetically pleasing portrait with superior lighting and material rendering. While Qwen Image 2.0 captures the specific detail of the beaded braids more effectively, it suffers from significant anatomical errors in the hand and a less nuanced approach to the requested lighting.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.2 [dev]
- + Excellent professional layout resembling a real multi-page menu.
- + Strong typography hierarchy with specific price alignments.
- + Effective use of vibrant color accents to separate sections.
- − Text is somewhat gibberish despite clean rendering.
- − The grid of images is a bit cluttered compared to Model B.
Qwen Image 2.0
- + Clean and modern grid-based layout.
- + Very high-quality and consistent food photography.
- + Follows the minimalist aesthetic closely with clear category headers.
- − Text rendering is messy with many character artifacts.
- − Logical layout issue: pizza images are placed under all three category headers.
- − Lacks the specific sub-details like descriptions typical of a 'professional layout'.
Verdict: FLUX.2 [dev] produces a much more realistic menu layout that includes item descriptions, varied pricing, and a structure that feels like a functional document. While Qwen Image 2.0 has superior food photography, it fails the logic of the prompt by placing pizzas under the 'Appetizers' and 'Mains' headers and suffers from significant text scrambling.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.2 [dev]
- + Excellent adherence to the 'exploded' layout with clear separation between all layers.
- + Text is perfectly legible and integrates cleanly into the fiery style.
- + The sauce droplets create a strong sense of dynamic motion.
Qwen Image 2.0
- + Texturing on the bun and meat patty is highly photorealistic.
- + The fiery effect on the main title is very detailed and visually impressive.
- + Good use of sparks and smoke to enhance the dark atmosphere.
- − Failed to properly 'explode' the burger; layers are mostly touching or overlapping.
- − The price starburst is visually cluttered compared to Model A.
Verdict: FLUX.2 [dev] followed the 'exploded burger' layout much more effectively, creating a clean marketing composition where every ingredient is visible. While Qwen Image 2.0 has superior texture detail on the food itself, it failed to separate the components in mid-air as requested, resulting in a more traditional stacked burger appearance.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.2 [dev]
- + Excellent chalk texture throughout the board.
- + Strong adherence to the specific menu text and prices.
- + Realistic cursive title with the requested date.
- − The 'with' on the second item has some overlapping smudge artifacts.
- − Visual background is slightly more cluttered compared to the other model.
Qwen Image 2.0
- + Natural lighting with realistic chalk smudges on a dark green board.
- + Consistent handwriting style that looks authentic.
- + Very clean and readable layout.
- − Slightly missed the 'elegant cursive' requirement for the title, using more of a print-hybrid.
- − The text is a bit thinner, making it slightly less impactful as a display board.
Verdict: Both models followed the complex text instructions perfectly, including specific dates and prices. FLUX.2 [dev] stands out for its superior chalk texture and the elegant cursive specified in the prompt, while Qwen Image 2.0 provides a more aesthetically pleasing background composition but slightly misses the cursive title requirement.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.2 [dev]
- + Excellent cinematic lighting and atmosphere
- + Realistic textures on the astronaut suit and horse fur
- + Natural, balanced composition with the curved Earth background
- − Failed the specific spatial instruction for the horse to be on top of the astronaut
Qwen Image 2.0
- + High level of intricate detail on the horse's skin and mane
- + Creative surreal elements like floating water droplets
- + Sharp, clear rendering of the space background
- − Failed the specific spatial instruction for the horse to be on top of the astronaut
- − Artificial, slightly 'plastic' look to the horse's scales
Verdict: Both models failed the negative/inverse logic of the prompt, which requested the horse to be on top of the astronaut. Instead, both FLUX.2 [dev] and Qwen Image 2.0 produced high-quality traditional interpretations of an astronaut riding a horse. FLUX.2 [dev] is the winner because it achieves a more cohesive cinematic look with superior lighting and realistic textures, whereas Qwen Image 2.0 feels slightly more over-processed and stereotypical.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.2 [dev]
- + Excellent adherence to the 'front paws on the steering wheel' instruction.
- + The passenger is correctly positioned in the back seat as requested.
- + The capybara's facial expression is remarkably professional and calm.
- − The paws appear slightly mutated with too many long digits.
- − The 'T' logo on the hat is slightly generic rather than a traditional NYC taxi badge.
Qwen Image 2.0
- + High photographic realism in the lighting and skin textures.
- + Captures the bored expression of the passenger very effectively.
- + The capybara's fur texture is extremely detailed.
- − The passenger is sitting in the front passenger seat instead of the back seat.
- − The capybara's paws are poorly rendered and do not appear to be firmly gripping the wheel.
- − The composition feels a bit cramped compared to the other model.
Verdict: FLUX.2 [dev] followed the spatial instructions much better by placing the passenger in the back seat and correctly positioning the capybara's paws on the wheel. While Qwen Image 2.0 has slightly more natural nighttime lighting, it failed the core compositional requirement of having the businesswoman in the back. FLUX.2 also better captured the specific 'professional' look of the driver.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.2 [dev]
- + Excellent typography with perfect spelling in all requested areas.
- + High-contrast cinematic lighting on the jack-o-lantern.
- + Sophisticated border design with thorns and webs cleanly integrated.
- − The transition from the central dark area to the parchment edge is a bit abrupt.
- − The banner is quite small and less decorative than Model B.
Qwen Image 2.0
- + Good vintage parchment texture throughout the composition.
- + Elegant scroll banner with a more natural placement.
- + Accurate rendering of the requested text elements.
- − The jack-o-lantern looks a bit like a stock photo slapped onto the background.
- − Border thorn details are slightly less crisp than Model A.
Verdict: Both models followed the complex text requirements perfectly. FLUX.2 [dev] is the winner because it creates a more cohesive artistic piece with superior lighting and a polished, professional layout, whereas Qwen Image 2.0 feels a bit more like a collage of individual elements.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.2 [dev]
- + Perfectly follows the 45-degree isometric perspective and diorama style.
- + Excellent miniature 3D cartoon aesthetic with soft, high-quality PBR-like textures.
- + Accurate text layout and centering as requested in the prompt.
- − The sushi types are slightly more generic compared to Image B.
Qwen Image 2.0
- + Higher realism in the food textures, particularly the fish grain and glaze.
- + Clean text rendering and well-designed flag icon.
- − Fails to capture the 'miniature 3D cartoon' style, leaning too far into photo-realism.
- − The camera angle is a standard high-angle shot rather than a true 45-degree isometric projection.
- − The diorama base is just a wooden board rather than a stylized raised platform.
Verdict: FLUX.2 [dev] is the clear winner as it perfectly adheres to the stylistic requirements of 'isometric miniature 3D cartoon' and 'diorama base', whereas Qwen Image 2.0 produced a more standard realistic food photo. FLUX.2 [dev] also followed the specific text placement and layout instructions more accurately.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.2 [dev]
- + Excellent fur texture and lighting consistency.
- + Beautiful, clean composition with soft bokeh and god rays.
- + Includes all requested animals plus an extra bunny for a dense, cute aesthetic.
- − The animals are largely static and posing rather than 'playfully chasing and tumbling'.
- − The kitten has slightly oversized, slightly uncanny eyes.
Qwen Image 2.0
- + Perfectly captures the 'tumbling' and 'playful' action described in the prompt.
- + Superior interaction between the different animals.
- + Stronger depiction of god rays and dynamic sunrise atmosphere.
- − The puppy has a fifth leg appearing near its chest (anatomical artifact).
- − The kitten's paws look slightly blurred or poorly defined during the action.
Verdict: FLUX.2 [dev] produces a much cleaner, more polished image with beautiful lighting, but the animals are just sitting still. Qwen Image 2.0 much better captures the 'tumbling' action and dynamic energy of the prompt, though it suffers from a significant anatomical error with the puppy's legs. FLUX.2 is the superior overall image despite the less active interpretation.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.2 [dev]
- + Excellent typography including the requested accent mark on 'Caffè'.
- + Perfect adherence to the 'vector emblem style' and 'minimalist' keywords.
- + Great use of subtle texture on the background and logo elements.
- − The steam icon is very small compared to the dome.
Qwen Image 2.0
- + Classic serif typography looks elegant and fitting for the era.
- + Creative integration of the flame/steam motif directly into the dome's reflection.
- − The illustration style is more of a 3D-shaded mascot than a minimalist vector emblem.
- − The banner lacks the complexity and 'Est. 1720' integration seen in the other model.
- − The steam looks more like a stylised flame than vapours.
Verdict: FLUX.2 [dev] followed the prompt much more accurately, producing a true minimalist vector emblem that looks like a professional logo design. While Qwen Image 2.0 produced a beautiful illustration, it used 3D shading and gradients that deviated from the 'minimalist vector' requirement. FLUX.2 [dev] also handled the banner and integrated text more effectively.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.2 [dev]
- + Excellent vector art style with professional-looking icons.
- + Captures detailed mission elements like the Saturn V and Lunar Module with high fidelity.
- + Includes a clean header featuring the crew names.
- − The layout is scattered and does not represent a logical chronological flow.
- − Text rendering contains significant gibberish and spelling errors like 'Trarquility' and 'Translunar' is misplaced.
- − Included extra, redundant icons that were not part of the requested steps.
Qwen Image 2.0
- + Perfectly follows the requested chronological sequence from top to bottom.
- + Accurate text rendering for labels like 'Launch', 'Earth Orbit', and 'Landing'.
- + Excellent adherence to the clean, modern flat-vector infographic style.
- − One minor typo in 'Translunjar'.
- − The icons are slightly more simplistic compared to the other model.
Verdict: Qwen Image 2.0 is the clear winner because it actually functions as an infographic, following the requested six-step sequence in a logical vertical layout with mostly accurate text. FLUX.2 [dev] produces higher quality individual illustrations, but the layout is disorganized, and the text labels are largely unintelligible or misplaced.
Explore each model
Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request