Black Forest Labs' compact, open-source image generation model with sub-second inference, optimized for production and near real-time applications with multi-reference support
Settled by community votes across 13 shared challenges, with an AI judge weighing in on each.
FLUX.2 [klein] 4B
#32 of 62 in Text-to-Image
Qwen Image 2.0
#34 of 62 in Text-to-Image
Where the votes landed
FLUX.2 [klein] 4B
100.0%
win rate
Ties
0.0%
Qwen Image 2.0
0.0%
win rate
Challenge by challenge
The strongest take from each model on every shared challenge, with the AI judge's read.
Geometric Composition
Text-to-Image“A glass cube on a wooden table. Inside the cube is a small blue sphere. On top of the cube sits a red book. A green plant is behind the cube, partially visible through the glass. Soft window light from the left.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Excellent adherence to lighting instructions with realistic soft shadows and highlights from the left window.
- + Highly realistic glass textures including dust-like specs and plausible reflections.
- + Perfect spatial composition where the book fits naturally on the cube.
- − The sphere is slightly off-center, though this is a minor aesthetic choice.
Qwen Image 2.0
- + Accurately includes all requested elements including the green plant and wooden table.
- + Clearer view of the sphere which appears more substantial in the scene.
- − Physical logic issues where the sphere appears to be floating mid-air without support.
- − The glass cube has strange internal glass dividers or thick edges that weren't requested and look unnatural.
- − The lighting is flat compared to the specific 'light from the left' request.
Verdict: FLUX.2 [klein] 4B followed the prompt with significantly higher realism and better adherence to the lighting instructions. While Qwen Image 2.0 included all elements, it struggled with the physics of the scene, creating a floating sphere and awkward internal geometry within the glass cube.
Candid Street Photography
Text-to-Image“A candid street photo of an elderly Japanese man repairing a red bicycle in light rain, reflections on wet pavement, shallow depth of field, 50mm lens, natural skin texture, imperfect framing, motion blur from passing cars, cinematic but realistic, no stylization.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Excellent atmosphere with visible rain and glowing reflections
- + Cinematic composition with beautiful urban depth of field
- + Accurate representation of the requested red bicycle in a street setting
- − Physical interaction with the bike is slightly awkward
- − Motion blur on the background cars is subtle rather than pronounced
Qwen Image 2.0
- + Strong 'imperfect framing' that heightens the candid feel
- + Exceptional skin texture and facial detail
- + Clearer action of 'repairing' the bicycle chain
- − The visible person on the right is cut off awkwardly
- − Wetness and rain effects are less visible compared to Model A
- − Background car is fairly sharp, missing the requested motion blur
Verdict: FLUX.2 [klein] 4B better captures the overall atmosphere of the prompt, particularly the rain and cinematic lighting. However, Qwen Image 2.0 provides a more realistic and detailed depiction of the man himself, with superior skin textures and a more convincing candid composition, despite some messy elements on the edge of the frame.
Fantasy Warrior
Text-to-Image“Close portrait of a battle-worn paladin in ornate engraved plate armor, hair braided with small beads, faint scars and dirt on the skin, warm torchlight reflecting off metal, shallow depth of field, bokeh sparks, lifelike eyes, highly detailed texture on leather straps and cloth underlayer.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Excellent anatomical realism in the facials features and skin texture.
- + Ornate engraving on the plate armor is clean and well-defined.
- + Subtle and realistic battle-worn effects like small scratches and grime.
- − The braids are very simple and lack the complex 'small beads' detail expected.
- − Lighting feels a bit flat and resembles a studio set rather than an environment.
Qwen Image 2.0
- + Strong adherence to the 'beads in braids' and 'leather straps' prompts.
- + Excellent cinematic atmosphere with dramatic warm lighting and bokeh sparks.
- + Captures the 'battle-worn' essence more effectively through dirt and scars.
- − Anatomy in the hand is slightly distorted and lacks polish.
- − The orange eyes look somewhat unnatural/supernatural rather than just 'lifelike'.
Verdict: While FLUX.2 [klein] 4B produces a much cleaner and more anatomically correct portrait, Qwen Image 2.0 captures the specific details and atmosphere of the prompt far better. Qwen includes more complex braiding, visible beads, and a superior sense of battle-worn grit, despite some minor issues with hand rendering.
Modern Clean Menu
Text-to-Image“Modern minimalist restaurant menu design, white background with colorful food photos in grid, sections for appetizers/pizza/mains, bold sans-serif fonts, vibrant accents, clean professional layout for casual dining.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Excellent photographic quality with vibrant, realistic food textures.
- + Logical menu layout with clear sections and varied price points.
- + High-quality graphic design elements like the logo and color accents.
- − Several text hallucinations and misspellings in the headings.
- − The grid is a bit chaotic with varying image sizes.
Qwen Image 2.0
- + Perfectly adhered to the request for sections: Appetizers, Pizza, and Mains.
- + Uniform grid layout creates a very clean, professional aesthetic.
- + High consistency in font usage and price placement.
- − The food photos look slightly more generic/stock-like compared to Model A.
- − Text contains significant gibberish and artifacts under the images.
Verdict: Qwen Image 2.0 followed the structural instructions of the prompt much better, providing the specific sections (Appetizers, Pizza, Mains) requested in a clean grid. While FLUX.2 [klein] 4B produced more appetizing and high-fidelity food photography, its layout was less organized and it failed to include the specific requested headers accurately.
Magic Burger Explosion: Fiery Photorealism Challenge
Text-to-Image“Ad for 'Magic Burger'. Dynamic, exploded burger with all components (bun, patty, cheese, lettuce, tomato, sauce) suspended in mid-air. Emphasize photorealistic detail and a sense of motion. Dark, fiery background with glowing embers. Integrate text: 'MAGIC BURGER' as a prominent title, 'LIMITED TIME ONLY' as a secondary message, and '€6.99' in a starburst, all rendered with a fiery, glowing effect.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + High resolution and realistic texture on the burger bun and patties.
- + Correct starburst implementation for the price.
- + Effective use of glowing ember effects and dynamic sparks.
- − Major text rendering failure with 'MAGIC BURGER' becoming gibberish.
- − The burger is not 'exploded' or separated as requested; it is mostly assembled.
Qwen Image 2.0
- + Perfect text rendering for all requested strings with the specified fiery effect.
- + Successful 'exploded' composition with the top bun and ingredients suspended separately.
- + Excellent lighting and steam/smoke effects enhance the 'magic' theme.
- − The price text is slightly off-center within its starburst.
- − The starburst graphic is a bit more illustrative than photorealistic.
Verdict: Qwen Image 2.0 followed the prompt instructions much more accurately, successfully rendering all text correctly and capturing the 'exploded' motion of the burger. FLUX.2 [klein] 4B failed significantly on the primary text title and kept the burger mostly intact, ignoring the core layout request.
Chalkboard Menu
Text-to-Image“Handwritten-style chalkboard menu in a cozy café, all text rendered in the exact same realistic chalk handwriting style with natural variations in letter size, slight slant, and chalk texture — no printed or digital fonts anywhere on the board. Title at the top in elegant cursive chalk handwriting: ‘TODAY’S SPECIALS – APRIL 30, 2026’. Below it, three menu items also in the same handwritten chalk style: ‘Truffle Mushroom Risotto – $24’, ‘Grilled Octopus with Lemon & Herbs – $28’, ‘Brown Butter Chocolate Chip Cookies – $9’. At the very bottom, smaller text in the identical handwritten chalk style (slightly smaller but still clearly legible with the same handwriting characteristics): ‘All items made fresh daily • Ask about our gluten-free options’. Warm ambient café lighting, visible chalk dust and smudges, realistic handwriting imperfections, no clean printed text anywhere.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Excellent realistic chalk texture with dusty smudges
- + Consistent handwriting style across all lines of text
- + Centered composition that focuses entirely on the menu
- − Multiple significant spelling errors including 'Truffel', 'Musheram', 'Ootrpous', and 'Brawn Buter'
- − Added extra characters in the title such as 'S SPECIALS'
Qwen Image 2.0
- + Perfect spelling on nearly all requested menu items
- + Beautiful environmental lighting and background context of a cafe
- + Very clear and legible handwriting while maintaining the chalk aesthetic
- − Missed the request for 'elegant cursive' in the title, using standard print instead
- − Slightly less realistic chalk texture compared to the other model
Verdict: Qwen Image 2.0 is the clear winner due to its superior text rendering accuracy and correct spelling of complex words like 'Risotto' and 'Octopus'. While FLUX.2 [klein] 4B has a slightly more convincing chalk smudge texture, its frequent and distracting spelling errors make it less usable for the specific prompt requirements.
The Reversed Rodeo
Text-to-Image“Horse riding astronaut in space — horse on top, not vice versa. Surreal, highly detailed, cinematic.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Excellent anatomical realism for both the horse and the astronaut's suit.
- + High resolution with cinematic lighting and a clear background of Earth.
- + Natural-looking dust/mist effect trailing from the hooves which adds to the surrealism.
- − The composition is a bit more static and less adventurous in its surreal elements compared to the other model.
- − The transition between the horse's hooves and the 'ground' above Earth is somewhat blurry.
Qwen Image 2.0
- + Creative use of textures, such as the scale-like pattern on the horse's neck.
- + Dynamic composition with floating liquid droplets enhancing the surreal theme.
- + Stronger sense of motion in the horse's mane and tail.
- − Anatomical issues with the horse's front right leg which appears disjointed.
- − The astronaut's hands and the integration with the reins are less defined than its competitor.
- − Slightly less clarity in the fine details of the space suit.
Verdict: Both models successfully followed the prompt, placing the astronaut on top of the horse (reversing the common 'astronaut riding a horse' trope's potential confusion). FLUX.2 [klein] 4B is the winner due to significantly better anatomical accuracy and overall image clarity, whereas Qwen Image 2.0 has notable structural errors in the horse's legs despite having a more imaginative texture style.
The Capybara Taxi Driver
Text-to-Image“Photorealistic scene inside a yellow New York taxi at night. A capybara is driving, wearing a yellow taxi driver cap and a dark jacket. It has a calm, professional expression and both front paws on the steering wheel. In the back seat sits a human businesswoman in a coat, looking at her phone with a completely normal, bored expression (as if this is just another normal ride). Through the windows you can see the streets of Manhattan at night with blurred lights. Realistic taxi interior, photorealistic, detailed fur and fabric, 35mm lens, night lighting with reflections, shallow depth of field.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Excellent adherence to the 'bored' expression of the passenger
- + High-quality fur texture and realistic anatomy for the capybara
- + Clear, high-resolution interior details and a believable NYC bokeh background
- − The steering wheel looks a bit small and simple in design
Qwen Image 2.0
- + Dynamic composition and good use of colors for a night scene
- + Creative 'paws on the wheel' interaction
- + Effective reflection of city lights on the car window
- − The passenger is sitting in the front seat next to the driver instead of the back seat as requested
- − The passenger's head is strangely cropped by the car door frame
- − Poor hand/paw rendering on the steering wheel
Verdict: FLUX.2 [klein] 4B followed all prompt instructions perfectly, including the specific requirement for the passenger to be in the back seat with a bored expression. Qwen Image 2.0 failed the spatial arrangement by placing the passenger in the front seat and suffered from awkward anatomical distortions.
The Halloween Invitation
Text-to-Image“Vintage gothic Halloween party invitation. Dark parchment poster, spooky border with webs and thorns, central glowing jack-o-lantern, bats, twisted trees, moody night sky. Add elegant gothic title text saying "Halloween Party Invitation", a small scroll banner saying "You are invited to a night of frights", and event details at the bottom: Date: 30.10.2026 Time: 7pm Location: The Arches, NYC Spooky but polished, cinematic lighting, square format.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Atmospheric cinematic lighting on the pumpkin
- + Intricate spiderweb corner details
- − Significant spelling errors in the main title
- − Incomplete event details (missing time)
Qwen Image 2.0
- + Perfect text rendering for all requested strings
- + Stronger parchment aesthetic for the background
- − The jack-o-lantern contains slight internal geometry artifacts
- − The transition between the central scene and border is a bit stark
Verdict: Qwen Image 2.0 is the clear winner due to its ability to accurately render all the requested text without spelling errors. While FLUX.2 [klein] 4B has excellent lighting, its garbled typography and missing information make it unsuccessful as an invitation.
Isometric Miniature Diorama Scenes
Text-to-Image“Create a clear, 45° top-down isometric miniature 3D cartoon scene of Japan's signature dish: sushi, with soft refined textures, realistic PBR materials, gentle lighting, on a small raised diorama base with minimal garnish and plate. Solid light blue background. At top-center: 'JAPAN' in large bold text, 'SUSHI' below it, small flag icon. Perfectly centered, ultra-clean, high-clarity, square format.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Excellent 3D miniature/cartoon aesthetic that perfectly matches the 'soft refined textures' prompt.
- + Superior lighting and material rendering for a stylized look.
- + Solid isometric composition.
- − Spelling error in text ('SUSH' instead of 'SUSHI').
- − Flag icon is incorrect and unrecognizable as the Japanese flag.
Qwen Image 2.0
- + Perfect text rendering for both 'JAPAN' and 'SUSHI'.
- + Includes an accurate Japanese flag icon as requested.
- + Good interpretation of the 'diorama base' using a wooden block.
- − Visual style is more photorealistic than the requested '3D cartoon scene'.
- − The 45° isometric angle is slightly off compared to the request.
Verdict: Qwen Image 2.0 is the overall winner because it successfully followed all text and icon instructions, whereas FLUX.2 [klein] 4B failed to spell 'SUSHI' correctly and provided an incorrect flag. While FLUX.2 [klein] 4B better captured the specific '3D cartoon' aesthetic, the accuracy of the informational elements in Qwen Image 2.0 makes it a more functional image.
Adorable Baby Animals in Sunny Meadow
Text-to-Image“Hyper-photorealistic scene of fluffy baby animals—a golden retriever puppy, tabby kitten, baby bunny, and red fox kit—with big expressive eyes and ultra-detailed soft fur, playfully chasing butterflies and tumbling together in a lush wildflower meadow, warm golden sunrise light with god rays and dew sparkles, joyful wholesome vibe, 8K masterpiece.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Strong god rays and vibrant lighting
- + Excellent fur texture on the fox and puppy
- + Balanced composition with clear focal points
- − Failed to include the baby bunny
- − Included two kittens instead of one
- − Lighting feels slightly more digital/artificial
Qwen Image 2.0
- + Included all four requested animals correctly
- + Better dynamic interaction with animals tumbling together
- + More naturalistic lighting and background atmosphere
- − The fox's facial structure looks a bit distorted in its playful pose
- − Minor blurriness on the kittens face compared to Model A
Verdict: Qwen Image 2.0 is the winner because it successfully followed the prompt's requirement to include all four specific animals (puppy, kitten, bunny, and fox), whereas FLUX.2 [klein] 4B missed the bunny entirely and duplicated the kitten. Qwen Image 2.0 also captured the 'tumbling together' aspect of the prompt more effectively with its dynamic central grouping.
Vintage Cafe Logo
Text-to-Image“Vintage minimalist restaurant logo for "Caffè Florian", retro cloche dome with steam and "Est. 1720" banner, classic typography, warm brown and cream tones, subtle texture on light background, vector emblem style.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Clean vector emblem style that matches the minimalist request.
- + Balanced layout with a professional centered composition.
- + Accurate date inclusion in a banner as requested.
- − Major spelling error in the primary brand name ('FLAXTION' instead of 'Florian').
- − Redundant repetition of 'Est. 1720' both in and below the banner.
Qwen Image 2.0
- + Perfect adherence to text prompts including the correct spelling of 'Caffè Florian'.
- + High-quality illustration with attractive lighting and shading on the cloche.
- + Strong implementation of the requested banner and steam elements.
- − Slightly less 'minimalist' than Model A due to the shading complexity.
- − The steam is rendered more like a flame icon inside the cloche rather than rising from it.
Verdict: Qwen Image 2.0 is the clear winner because it correctly spelled the brand name 'Caffè Florian', whereas FLUX.2 [klein] 4B hallucinated the name as 'FLAXTION'. While FLUX.2 captured a more minimalist vector style, Qwen Image 2.0 followed every specific text prompt accurately and provided a more polished final graphic.
Apollo 11: Journey to Tranquility
Text-to-Image“Create a clean, modern vector infographic poster about the Apollo 11 mission. NASA-inspired palette (navy, white, muted red, light gray). Flat-vector style, crisp lines, consistent iconography, subtle gradients only. Steps (stop at landing): 1. Launch (Saturn Vicon) 2. Earth Orbit (Earth + orbit ring icon) 3. Translunar (trajectory arc icon) 4. Lunar Orbit (Moon + orbit ring icon) 5. Descent (lunar module descending icon) 6. Landing (lunar module on the surface icon) Small supporting elements (minimal text): • Crew strip: three silhouette icons with only last names: Armstrong, Aldrin, Collins. • Landing site marker: Moon pin labeled "Tranquility" only. Layout constraints: generous margins, large readable labels, clean background with subtle stars. Vector-only, print-poster look, high resolution.”
AI Judge Analysis
FLUX.2 [klein] 4B
- + Clean vector aesthetic for the character silhouettes and the lunar module.
- + Follows the NASA-inspired color palette effectively.
- − Text rendering is unintelligible with numerous spelling errors.
- − Layout is disorganized and lacks a clear chronological flow for an infographic.
- − Includes duplicate Earth icons that do not represent separate mission stages.
Qwen Image 2.0
- + Excellent text legibility and nearly perfect spelling of key terms.
- + Strong vertical layout that logically follows the mission progression steps.
- + Excellent adherence to the requested iconography for each specific stage.
- − Small typo in 'Translunjar' (extra 'j').
- − The Saturn V icon for the 'Launch' stage is very small compared to other elements.
Verdict: Qwen Image 2.0 is the clear winner as it successfully creates a functional infographic with a logical flow and highly legible text. While FLUX.2 [klein] 4B has a nice artistic vector style, its failure to generate readable text or a coherent sequence makes it unusable for the requested infographic purpose.
Explore each model
Alibaba's Qwen Image 2.0 model with enhanced text rendering, supporting both Chinese and English prompts with up to 6 images per request